Privacy-Preserving Intrusion Detection in Industrial IoT Ecosystems via Decentralized Federated Learning with Proximal Regularization

Privacy-Preserving Intrusion Detection in Industrial IoT Ecosystems via Decentralized Federated Learning with Proximal Regularization

 

Zainab H. Mohammad

 

Department of Mathematics, Najaf Center, The Open Educational College, Ministry of Education, Najaf, Iraq, zainabaldulame@gmail.com.

 

 

 Abstract

The rapid development of the Industrial Internet of Things (IIoT) has transformed the current industrial control systems (ICS), but in the process has revealed operational technology (OT) to advanced cyber-threats. Conventional centralized intrusion detection systems (IDS) demand the pooling of massive telemetry traffic, a requirement that contributes to prohibitive communication latencies and breaches strict corporate data privacy requirements. Although Federated Learning (FL) provides a more decentralized alternative by training the models locally on edge gateways, it performs badly in non-Identically and Independently Distributed (non-IID) network traffic a condition ubiquitous in heterogeneous factory setting. The paper suggests a privacy-preserving, robust, localized deep learning-based IDS, which utilizes Federated Proximal (FedProx) optimization. The framework is based on multi-layer feed-forward neural network architecture, only that it incorporates parameterized proximal regularization term (mu = 0.5) as a part of local loss function to penalize client parameters deviation and eliminate client drift phenomenon. The suggested framework is experimented in benchmarking, with the simulated variants of the realistic Edge-IIoTset dataset, in the multi-dimensional instrumentation pipeline that monitors the validation accuracy, convergence on the cross-round loss and confusion matrix profiling. Our regularized framework, as experimental results indicate, exhibits monotonic convergence, and reaches the highest cyber-threat classification accuracy of 91.4 percent and massive reduction in false-positive rates, whereas the baseline Vanilla Federated Averaging (FedAvg) displays violent oscillation (between 58% and 74%). The complete MATLAB simulation testbed is open-sourced so as to give a verified benchmark with respect to secure OT collaborative intelligence.

Keywords: Industrial Internet of Things (IIoT); privacy-preserving machine learning; non-IID network traffic; edge computing; operational technology security.

 

1. Introduction

The radical digital transformation of the modern industrial industry is occurring, where the paradigm of the Industrial Internet of Things (IIoT) is incorporated into the operational technology (OT) and the information technology (IT) [1]. The architectures of IIoT have developed unprecedented manufacturing efficiency, predictive maintenance, and visibility of the supply chain by integrating pervasive sensing, cloud-edge computing, and smart actuation layers into physical workflows [2], [3]. This excessive interconnectivity has however fundamentally changed the threat landscape [4].

Previously air-gapped insulated, industrial control infrastructures including Supervisory Control and Data Acquisition (SCADA) systems and programmable logic controllers (PLCs) are now susceptible to high-consequence cyber-attacks, including distributed denial-of-service (DDoS), ransomware, industrial espionage, and protocol-specific injections [5], [6], [7]. The combination of smart Intrusion Detection Systems (IDS) is necessary so that such vital facilities could be safeguarded against disruption [4], [8].

Conventional network-wide IDS designs are based on a centralized topology, in which tremendous streams of high-frequency raw network packet logs and telemetry information are continuously sent out from localized manufacturing cells to a monolithic cloud repository for batch machine learning analysis [8], [9]. Computationally simple but with devastating operational and structural vulnerabilities, this centralized data-pooling paradigm is:

 1. Network Bandwidth Satiation: Continuous ingestion of raw data over (low-power), lossy networks (LLNs), or remote industrial communication links quickly depletes available bandwidth, causing unacceptable telemetry ingestion latencies [10], [11].

 2. Corporate and Compliance Bottlenecks: The conservativeness of industrial stakeholders as far as data sovereignty is concerned is intense [12]. Proprietary operational recipes, production rates, and health profiles of assets are often encapsulated in network logs; these logs being exposed to centralized third-party cloud engines breach strict cooperate compliance directives and international regulations including the General Data Protection Regulation GDPR [13], [14], [15].

 3. Single Point of Failure Vulnerabilities: Centralized storage repositories are high value targets for malicious state-sponsored adversaries, in which one breach can result in the systemic intelligence of the entire industrial matrix being compromised [16], [17], [18].

In order to resolve the structural conflict between the cooperative threat intelligence and the local data confidentiality, Decentralized Collaborative Machine Learning particularly, Federated Learning (FL) is an innovative structural change [19], [20]. The basic idea of FL is that the training process of the model is no longer tied to the need to access data directly [20], [21] . In an FL orchestration, isolated industrial edge nodes make use of their local compute resources to independently train private intrusion detection models over their localized network traffic streams [12], [21]. These nodes do not send raw packet data, but instead, only the parameters of their abstracted mathematical model (i.e., weights and biases) are uploaded to a coordinating central server [21], [23].

The central server then runs an aggregation protocol to keep these localized updates together in a very generalized global defensive model that is then redistributed to edge nodes [19], [24], [25], [26]. Although this is theoretically a promising approach, when applied in practice to real-world industrial environments, the implementation of standard FL frameworks (such as Vanilla Federated Averaging - FedAvg) reveals a critical weakness: Statistical Heterogeneity (Non-IID Data) [27], [28]. In practice applications, an industrial water purification facility, a car assembly line, and a smart power grid can be configured to use completely different communication protocols (e.g., Modbus, MQTT, OPC UA), different traffic rates [29], [30].

As a result, their local network logs are non-Identically and Independently Distributed (non-IID) [27], [31]. In the case of vanilla aggregation of highly stratified non-IID data, the local updates computed by independent nodes are susceptible to pull the global model towards isolated, divergent local optima, a pathological vulnerability referred to as client drift [28], [32], [33]. This structural drift causes disastrous convergence of models with unstable classification oscillations and intolerably large false-alarm rates to the extent of disruption of normal factory operation [31], [33], [34].

1.1  Motivation and Contributions

The nature of the study is to develop a highly robust, privacy-protecting, decentralized intrusion detection system that is capable of sustaining a stable and high-accuracy defensive running in the face of extreme statistical non-IID data skewness between heterogeneous industrial edge nodes. To fill the gap between pragmatic deployment considerations and theoretical knowledge, we suggest a regularized federated learning architecture that acts to proactively resolve the client drift issue without undermining the privacy of the edge nodes. Explicit, high-impact contributions of the current study might be as follows:

 Architecting a Non-IID Resilient Edge IDS: We propose and implement a localized deep-learning-based IDS, and a parameterized proximal regularization method to stabilize global models’ convergence in the presence of extreme statistical divergence environments at multi-tenant industrial nodes.

 Multi-Dimensional Performance Instrumentation: Unlike the traditional literature where the accuracy is only evaluated, we do construct a complete academic benchmarking suite which evaluates the co-evolution of Classification Accuracy, Global Objective Loss Convergence and Confusion Matrix Profiling.

 Empirical Validation on Modern Protocols: Our simulation is based on a very structured telemetry distribution in accordance with the modern Edge-IIoT set benchmark to capture the multi-protocol dynamics of cyber-attacks (including malicious MQTT and Modbus injections).

 A Complete Open-Source MATLAB Simulation Engine: Our implementation is a fully instrumented structurally transparent MATLAB simulation of isolated edge client optimization loops and centralized server weight replacements, which become an open-source baseline of future OT security research.

 

2. Proposed Methodology

The following section describes the architectural mechanics and mathematical basis of our privacy-preserving decentralized intrusion detection system that has been developed to survive strong non-IID skewness of an industrial grid [19], [20].

Fig1. Illustrates the architectural mechanisms and mathematical model of a decentralized intrusion detection system designed to maintain privacy and withstand extreme heterogeneous deviation.

2.1 System and Network Architecture

The decentralized framework models an enterprise industrial ecosystem consisting of K distinct, geographically isolated or structurally siloed IIoT edge gateways, coordinated by a central orchestration server located either in a private corporate cloud or a high-availability node [21], [35]. Each local edge gatewayacts as a defensive sentry for its specific manufacturing cell, collecting real-time network traffic logs into a highly localized private dataset[12], [22].

The aggregate global dataset across the entire industrial matrix is denoted as, with the total sample volume quantified as [19], [21].

To strictly enforce data privacy as well as zero-trust corporate compliance, raw packet captures within never exit the boundary regarding edge node k [11], [36]. The unified optimization objective of collective network is to minimize the global empirical risk function over shared vector of deep neural network weights[19], [15], [37]:

where is the local loss function evaluated over the private data silo of client [19], [38], [39]:

Here, represents the cross-entropy loss tracking the classification errors of the network model mapping input telemetry features to cyber-attack labels.

2.2 Mathematical Formulation of Proposed Proximal Regularization (FedProx)

In a vanilla federated execution (FedAvg), each client performs a fixed number of local stochastic gradient descent (SGD) iterations based on the global state distributed at the beginning of round t, updating its local parameters to[19], [28]. In the case when underlying datasets  are intensely non-IID, local gradients point toward wildly different geometric directions [27], [32]. Under simple averaging of standard FedAvg, such unconstrained updates cancel each other out or destabilize global parameter path, showing severe accuracy drop-offs and optimization divergence [28], [33].

For resolving such problem, our approach replaces vanilla optimization loop with structured proximal regularization algorithm (FedProx) [38]. At each global communication round, rather than minimizing standard empirical risk each local IIoT node k is forced to optimize regularized local objective function mathematically specified as [38], [32]:

where:

  is the original empirical local classification loss over the node's distinct traffic data [19].

 is the static global model parameter vector received from the central server at the start of the current round [21].

 represents the evolving local model parameter vector during the current edge training epochs [23].

is a meticulously calibrated proximal regularization hyperparameter (configured to in our experimental pipeline) [38].

The Role of the Proximal Penalty Term: The quadratic expansion component functions as a dynamic gravitational constraint [38], [40]. It continuously calculates the  norm distance between the local model weights w and the global reference model[41], [42].

The proximal term penalizes extreme change when an edge node experiences an extreme skew, of a certain type of attack (e.g., a giant localized DDoS burst), such that the local gradients of the model would need to change in response [38], [43]. The frameworks were mathematical constraints that bounded the effect on client drift by limiting how far individual node could wander relative to common global baseline [31]. This ensures that the consolidated model is not only high-accuracy generalized to all the participating factory nodes, but it is also not brought down by localized statistical anomalies 38], [33].

2.3 Deep Neural Network Structural Design

The classification engine utilized is the same on the server as on the edge nodes constituting optimized, deep feed-forward neural network specially constructed to run fast inference and high discrimination with tabular network logs [9], [44]. The model projects an 8-dimensional input feature space to a multi-class probability distribution. The architecture of the layer graph is [45], [46]:

1.     Feature Input Layer: Accepts normalized 8-dimensional continuous vector streams representing core industrial packet indicators:

 

  

 2. Fully Connected Layer 1: Expands the input space to 64 hidden neurons via a matrix dot product and bias addition: , where .

 3. ReLU Activation Layer 1: Introduces non-linear thresholding:, eliminating gradient saturation.

 4. Fully Connected Layer 2: Compresses information into 32 hidden neurons:, isolating complex cross-protocol features.

 5. ReLU Activation Layer 2: Applies identical element-wise rectifying activation:.

 6. Output Bridge Fully Connected Layer: Maps the hidden features directly to the target classification dimensions:  where ( being the target categorical class count).

 7. Softmax Transformation Layer: Converts raw logits into continuous probabilistic scores spanning a bounded range [0, 1]:

 8. Classification Output Layer: Computes cross-entropy loss against ground-truth categorical assignments during training and outputs the maximum-likelihood binary or multi-class assignment during real-time edge monitoring.

3. Experimental Implementation and Testbed Instrumentation

In order to test the theoretical arguments of the proposed proximal regularized IDS framework rigorously, a highly advanced distributed simulation platform was completely executed in MATLAB (R2024a/2026b). In this part, the setup of the experimental environment, the strategies used for dataset partitioning, and the structural verification code that were executed to produce the empirical results are outlined.

3.1 Dataset Preprocessing and Realistic Non-IID Partitioning

The empirical pipeline operates on data structures that are directly matching to the Edge-IIoTset benchmark dataset that represents a modern IoT/IIoT security testing corpus capturing practical industrial protocol settings (e.g., MQTT, Modbus, HTTP) with targeted malicious exploits [30], [47]. To test the system with extreme statistical heterogeneity, 15,000 various telemetry records are ingested, 8 critical structural network features, and categorical attack vectors [29], [30].

In order to simulate a fragmented, adversarial industrial network, in which the distribution of the data across the factories is extremely unbalanced, the data are subjected to a strict Non-IID Stratification Sequence [48], [49], [50]. The entire corpus is ordered according to its attack label distribution (sortIdx), the legitimate records and the different logs of cyber-exploits are grouped into isolated blocks. The sorted samples are then cut into separate unequal segments and distributed to 3 separate edge client cell matrices (X_Local, Y_Local). This ensures that Client 1 is inundated heavily with benign operational logs, and that Client 2 and Client 3 have their own, very skewed segments of malicious injections. This forms a severe mismatch on an extreme scale, which presents an extreme convergence difficulty to conventional decentralized learning algorithms.

 

4. Empirical Results and Detailed Discussion

This part presents an analytical, in-depth exploration of the 15-dimensional outputs that are yielded during the 15 rounds of the global federation.

4.1 Convergence and Accuracy Analysis under Non-IID Skewness

The accuracy and cross-entropy loss trajectories of the global model are co-evolved in real-time and plotted in Fig. 2.

Fig 2. The real-time co-evolution of the global model’s accuracy and cross-entropy loss trajectories.

An introductory analysis of the empirical patterns shows that there is a sharp performance divergence. Vanilla FedAvg (Baseline) is highly mathematically unstable and its classification performance is violently oscillating with a minimum bound of 58.1% and a high threshold of only 74.3%. It does not show steady upward learning curve. This disastrous pattern is the actual real-life manifestation of client drift. Since the data partitioning sequence maps highly separated, polarized subsets of Edge-IIoTset packet dynamics to the various edge clients, each local gateway updates its own model parameters to vastly distinct geographic coordinates in the optimization space. When the cloud server averages the plain average regarding such divergent vectors, the parameters are in conflict, which weakens predictive abilities of the model.

On the contrary, the suggested FedProx Framework establishes significantly smooth, high-stability monotonic convergence profile. Through activating proximal regularization modifier, the system penalizes gradient adjustments that venture outside global consensus boundaries. Thus, local optimization paths are constrained, suppressing individual weight divergence. The suggested global model surpasses the baseline by round 3, reduces erratic fluctuations, and achieves stable, definitive classification accuracy of 91.4% at round 15.

At the same time, the global objective loss trajectory (Fig. 2, Right Subplot) is mathematically validated by inspection. Vanilla FedAvg training profile has an erratic plateau because it has constant gradient conflict. The Proposed FedProx loss curve, on the other hand, shows a healthy continuous exponential decay, and indicates that the structural inclusion of the proximal regularization term can indeed stabilize the multi-tenant optimization trajectory with extreme statistical heterogeneity.

 

4.2 Statistical Robustness via Confusion Matrix Profiling

In an effort to comprehensively check the multi-class categorization integrity of the frameworks, Fig. 3 plots the cross-normalized confusion charts obtained at the final cycle of communication.

 

Fig 3. The cross-normalized confusion charts extracted at the final communication cycle.

 

The critical operational performance can be studied by analyzing the row summaries. The baseline Vanilla FedAvg model is highly statistically vulnerable with a high False Positive (FP) inflation rate. To be more exact, it misclassifies a significant amount of legitimate baseline control packets as malicious anomalies. Such unstable false alarm profile would not be sustainable in operational manufacturing facility, where it would lead to frequent and unnecessary line shutdowns and delays in operations.

On the other hand, the Proposed FedProx framework creates a definite statistical advantage. True Positive Rate (TPR / Recall) of the normal operation logs and of the active network injections reaches maximum efficiency and the False Positive Rate (FPR) is kept to the minimum. This model effectively classifies small deviations in protocols without confusing normal industrial control pattern and malicious payloads. This shows that the localized model drift can be directly constrained to optimize the statistical accuracy of the decentralized defensive models.

 5. Conclusion and Future Work

The presented work has efficiently provided a privacy-preserving, robust decentralized intrusion detection system that is particularly engineered for withstanding intense statistical non-IID data heterogeneity across multi-protocol Industrial IoT environments. Through transforming conventional unconstrained local optimization loop into regularized objective through parameterized proximal constraint, our framework efficiently mitigates the destructive client drift phenomenon inherent in fragmented industrial data silos.

Large-scale multi-dimensional empirical benchmarking showed that whereas conventional Vanilla Federated Averaging (FedAvg) shows violent optimization oscillations and suboptimal generalization behavior with respect to skewed data distributions, the regularized architecture proposed an architecture offers stable, monotonic convergence, and achieves optimal security classification rates of 91.4 percent with low false alarms. This makes it viable in the protection of high-security OT. We will extend this work in two main directions in future research:

 1. Cryptographic Privacy Upgrades: Combining localized Differential Privacy (DP) and lightweight Secure Multi-Party Computation (sMPC) protocols to protect the parameter exchange channel against advanced gradient leakage and model inversion attacks.

 2. Resource-Constrained Optimization: Testing the performance of the framework on the lossy and bandwidth-constrained dynamic wireless networks to define low-latency, resilience defensive intelligence on the highly volatile physical edge infrastructures.

 

REFERENCES

[1] P. Kairouz et al., “Advances and open problems in federated learning,” Found. Trends Mach. Learn., vol. 14, nos. 1-2, pp. 1-210, 2021.

[2] D. C. Nguyen et al., “Federated learning for Internet of Things: A comprehensive survey,” IEEE Commun. Surveys Tuts., vol. 23, no. 3, pp. 1622-1658, 3rd Quart., 2021.

[3] A. Brecko, E. Kajati, J. Koziorek, and I. Zolotova, “Federated learning for edge computing: A survey,” Appl. Sci., vol. 12, no. 18, Art. no. 9124, Sep. 2022.

[4] J. Lansky, S. Ali, M. Mohammadi, M. K. Majeed, S. H. T. Karim, S. Rashidi, M. Hosseinzadeh, and A. M. Rahmani, “Deep learning-based intrusion detection systems: A systematic review,” IEEE Access, vol. 9, pp. 101574-101599, 2021.

[5] M. A. Ferrag, O. Friha, L. Maglaras, H. Janicke, and L. Shu, “Federated deep learning for cybersecurity in the Internet of Things: Concepts, applications, and experimental analysis,” IEEE Access, vol. 9, pp. 138509-138542, 2021.

[6] K. S. Pokkuluri, A. Kumar, K. K. Singh Gautam, P. Deshmukh, and L. Abualigah, “Collaborative intelligence for IoT: Decentralized network security and confidentiality,” J. Intell. Syst. Internet Things, vol. 13, no. 2, 2024.

[7] E. Gyamfi and A. Jurcut, “Intrusion detection in Internet of Things systems: A review on design approaches leveraging multi-access edge computing, machine learning, and datasets,” Sensors, vol. 22, no. 10, Art. no. 3744, 2022.

[8] A. Aldweesh, A. Derhab, and A. Z. Emam, “Deep learning approaches for anomaly-based intrusion detection systems: A survey, taxonomy, and open issues,” Knowl.-Based Syst., vol. 189, Art. no. 105124, 2020.

[9] Elmobark, N., Abdzaid, A. Y., Alaa, F., & Saad, A. (2026). Experimental Evaluation of Automated and Manual Data Cleaning Systems: A Case Study Using Organizational Data. International Journal of Theoretical & Applied Computational Intelligence, 2026, 76-103.

[10] M. Chen, N. Shlezinger, H. V. Poor, Y. C. Eldar, and S. Cui, “Communication-efficient federated learning,” Proc. Natl. Acad. Sci. USA, vol. 118, no. 17, Art. no. e2024789118, 2021.

[11] A. Reisizadeh et al., “FedPAQ: A communication-efficient federated learning framework with quantized updates,” in Proc. 23rd Int. Conf. Artif. Intell. Stat. (AISTATS), 2020, pp. 2421-2431.

[12] Y. Chen, L. Liu, Y. Ping, M. Atiquzzaman, S. Mumtaz, Z. Zhang, M. Guizani, and Z. Tian, “A lightweight and fair privacy-preserving federated learning framework for IoT,” IEEE Trans. Netw. Service Manag., vol. 21, no. 5, pp. 5843-5858, 2024.

[13] R. C. Geyer, T. Klein, and B. Mair, “Differentially private federated learning: A client-level perspective,” arXiv:1712.07557, 2017.

[14] S. Truex, N. Baracaldo, A. Anwar, T. Steinke, H. Ludwig, R. Zhang, and Y. Zhou, “A hybrid approach to privacy-preserving federated learning,” in Proc. 12th ACM Workshop Artif. Intell. Security, Nov. 2019, pp. 1-11.

[15] M. Mohri, G. Sivek, and A. T. Suresh, “Agnostic federated learning,” in Proc. 36th Int. Conf. Mach. Learn. (ICML), 2019, pp. 4615-4625.

[16] H. HaddadPajouh et al., “A survey on multi-access edge computing security in IIoT environments,” Digit. Commun. Netw., vol. 7, no. 4, pp. 512-526, 2021.

[17] V. Mothukuri et al., “A survey on security and privacy of federated learning,” Comput. Sci. Rev., vol. 39, Art. no. 100340, 2021.

[18] L. Lyu, H. Yu, and Q. Yang, “Threats to federated learning: A survey,” arXiv:2003.02133, 2020.

[19] Elmobark, N., Saad, A., Hasan, S. H., & Badouch, M. (2026). The Hadoop Ecosystem: An Open-Source Framework for Enterprise-Scale Big Data Processing and Analytics. Engineering Headway, 35, 210-226.

[20] Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated machine learning: Concept and applications,” ACM Trans. Intell. Syst. Technol., vol. 10, no. 2, pp. 1-19, 2019.

[21] Y. Baseri, A. S. Hafid, D. Makrakis, and H. Fereidouni, “Privacy-preserving federated learning framework for risk-based adaptive authentication,” arXiv:2508.18453, 2025.

[22] K. Bonawitz et al., “Towards federated learning at scale: System design,” Proc. Mach. Learn. Syst. (MLSys), vol. 1, pp. 374-388, 2019.

[23] Z. Zhang, S. Rath, J. Xu, and T. Xiao, “Federated learning for smart grid: A survey on applications and potential vulnerabilities,” ACM Trans. Cyber-Phys. Syst., vol. 10, no. 1, pp. 1-26, 2026.

[24] V. Smith, C. K. Chiang, M. Sanjabi, and A. Talwalkar, “Federated multi-task learning,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 4424-4434, 2017.

[25] S. Caldas et al., “LEAF: A benchmark for federated settings,” arXiv:1812.01097, 2018.

[26] C. He et al., “FedML: A research library and benchmark for federated machine learning,” in Proc. NeurIPS Workshop, 2020.

[27] Y. Zhao et al., “Federated learning with non-IID data,” arXiv:1806.00582, 2018.

[28] X. Li, K. Huang, W. Yang, S. Wang, and Z. Zhang, “On the convergence of FedAvg on non-IID data,” in Proc. Int. Conf. Learn. Representations (ICLR), 2020.

[29] M. Alemayehu, M. C. Ghanem, H. Kheddar, D. Dunsin, and M. J. Lacerda, “Systematic analysis on the use of AI techniques in industrial IoT DDoS attack detection, mitigation, and prevention,” IoT, vol. 7, no. 51, pp. 1-56, 2026.

[30] Saad, A. (2026, January). Federated Learning-Based Intrusion Detection System Using Convolutional Neural Networks for IoT Networks. In 2026 5th International Conference on Electrical, Computer & Telecommunication Engineering (ICECTE) (pp. 1-6). IEEE.

[31] S. K. A. Hashim, Y. B. M. Yussoff, and S. B. Shahbudin, “Mitigating zero-day vulnerabilities in IIoT systems: Challenges and advances in AI-powered intrusion detection systems,” Mesopotamian J. CyberSecurity, vol. 5, no. 3, pp. 1184-1198, 2025.

[32] C. Y. Huang, K. Srinivas, X. Zhang, and X. Li, “Overcoming data and model heterogeneities in decentralized federated learning via synthetic anchors,” arXiv:2405.11525, 2024.

[33] L. Yuan, J. Zhang, M. Duan, G. Xiao, Z. Tang, and K. Li, “PRFL: Personalized and robust federated learning for non-IID data with malicious participants,” IEEE Trans. Mobile Comput., early access, 2025.

[34] A. Yazdinejad, A. Dehghantanha, H. Karimipour, G. Srivastava, and R. M. Parizi, “A robust privacy-preserving federated learning model against model poisoning attacks,” IEEE Trans. Inf. Forensics Security, vol. 19, pp. 6693-6708, 2024.

[35] W. Y. B. Lim et al., “Federated learning in mobile edge networks: A comprehensive survey,” IEEE Commun. Surveys Tuts., vol. 22, no. 3, pp. 2031-2063, 3rd Quart., 2020.

[36] Saad, A., Hasan, N. F., & Altaher, A. W. (2025). Deep Reinforcement Learning-Based Network Intrusion Prevention in Cloud-Edge Architectures.

[37] S. Reddi et al., “Adaptive federated optimization,” in Proc. Int. Conf. Learn. Representations (ICLR), 2021.

[38] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” Proc. Mach. Learn. Syst. (MLSys), vol. 2, pp. 429-450, 2020.

[39] H. Liu, X. Zhang, X. Shen, and H. Sun, “A federated learning framework for smart grids: Securing power traces in collaborative learning,” arXiv:2103.11870, 2021.

[40] L. Collins, H. Hassani, A. Mokhtari, and S. Shakkottai, “Exploiting shared representations for personalized federated learning,” in Proc. 38th Int. Conf. Mach. Learn. (ICML), 2021, pp. 2089-2099.

[41] A. Fallah, A. Mokhtari, and A. Ozdaglar, “Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 3557-3568, 2020.

[42] C. T. Dinh, N. H. Nguyen, N. H. Tran, and W. Zhang, “Personalized federated learning with Moreau envelopes,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 21396-21407, 2020.

[43] K. Pillutla, S. M. Kakade, and Z. Harchaoui, “Robust aggregation for federated learning using geometric medians,” IEEE Trans. Signal Process., vol. 70, pp. 1211-1225, 2022.

[44] M. F. L. Becerra-Suarez, V. A. Tuesta-Monteza, H. I. Mejia-Cabrera, and J. Arcila-Diaz, “Performance evaluation of deep learning models for classifying cybersecurity attacks in IoT networks,” Informatics, vol. 11, no. 2, Art. no. 32, May 2024.

[45] A. Hard et al., “Federated learning for mobile keyboard prediction,” arXiv:1811.03604, 2018.

[46] J. Xu et al., “Federated learning for healthcare informatics: Challenges and applications,” IEEE J. Biomed. Health Inform., vol. 25, no. 8, pp. 2821-2831, Aug. 2021.

[47] M. A. Ferrag, O. Friha, D. Hamouda, L. Maglaras, and H. Janicke, “Edge-IIoTset: A new comprehensive realistic cybersecurity dataset of IoT and IIoT applications for centralized and federated learning,” IEEE Access, vol. 10, pp. 40281-40306, 2022.

[48] P. Blanchard et al., “Machine learning with adversaries: Byzantine-tolerant gradient descent,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 119-129, 2017.

[49] C. Markarian and A. Panthakkan, “Algorithmic self-repair: Frontiers in fault-tolerant computation,” Front. Comput. Sci., vol. 8, Art. no. 1717711, 2026.

[50] J. Tu, L. Yang, and J. Cao, “Distributed machine learning in edge computing: Challenges, solutions, and future directions,” ACM Comput. Surv., vol. 57, no. 5, pp. 1-37, 2025.