Privacy-Preserving Intrusion Detection in Industrial IoT Ecosystems
via Decentralized Federated Learning with Proximal Regularization
Department of Mathematics, Najaf Center, The Open Educational
College, Ministry of Education, Najaf, Iraq, zainabaldulame@gmail.com.
Abstract
The rapid development of the Industrial Internet of Things (IIoT)
has transformed the current industrial control systems (ICS), but in the
process has revealed operational technology (OT) to advanced cyber-threats.
Conventional centralized intrusion detection systems (IDS) demand the pooling
of massive telemetry traffic, a requirement that contributes to prohibitive
communication latencies and breaches strict corporate data privacy
requirements. Although Federated Learning (FL) provides a more decentralized
alternative by training the models locally on edge gateways, it performs badly
in non-Identically and Independently Distributed (non-IID) network traffic a
condition ubiquitous in heterogeneous factory setting. The paper suggests a
privacy-preserving, robust, localized deep learning-based IDS, which utilizes
Federated Proximal (FedProx) optimization. The framework is based on
multi-layer feed-forward neural network architecture, only that it incorporates
parameterized proximal regularization term (mu = 0.5) as a part of local loss
function to penalize client parameters deviation and eliminate client drift
phenomenon. The suggested framework is experimented in benchmarking, with the
simulated variants of the realistic Edge-IIoTset dataset, in the
multi-dimensional instrumentation pipeline that monitors the validation
accuracy, convergence on the cross-round loss and confusion matrix profiling.
Our regularized framework, as experimental results indicate, exhibits monotonic
convergence, and reaches the highest cyber-threat classification accuracy of
91.4 percent and massive reduction in false-positive rates, whereas the
baseline Vanilla Federated Averaging (FedAvg) displays violent oscillation
(between 58% and 74%). The complete MATLAB simulation testbed is open-sourced
so as to give a verified benchmark with respect to secure OT collaborative
intelligence.
Keywords: Industrial
Internet of Things (IIoT); privacy-preserving machine learning; non-IID network
traffic; edge computing; operational technology security.
1. Introduction
The radical digital transformation of the modern industrial
industry is occurring, where the paradigm of the Industrial Internet of Things
(IIoT) is incorporated into the operational technology (OT) and the information
technology (IT) [1]. The architectures of IIoT have developed unprecedented
manufacturing efficiency, predictive maintenance, and visibility of the supply
chain by integrating pervasive sensing, cloud-edge computing, and smart
actuation layers into physical workflows [2], [3]. This excessive
interconnectivity has however fundamentally changed the threat landscape [4].
Previously air-gapped insulated, industrial control infrastructures
including Supervisory Control and Data Acquisition (SCADA) systems and
programmable logic controllers (PLCs) are now susceptible to high-consequence
cyber-attacks, including distributed denial-of-service (DDoS), ransomware,
industrial espionage, and protocol-specific injections [5], [6], [7]. The
combination of smart Intrusion Detection Systems (IDS) is necessary so that
such vital facilities could be safeguarded against disruption [4], [8].
Conventional network-wide IDS designs are based on a centralized
topology, in which tremendous streams of high-frequency raw network packet logs
and telemetry information are continuously sent out from localized
manufacturing cells to a monolithic cloud repository for batch machine learning
analysis [8], [9]. Computationally simple but with devastating operational and
structural vulnerabilities, this centralized data-pooling paradigm is:
1. Network Bandwidth
Satiation: Continuous ingestion of raw data over (low-power), lossy
networks (LLNs), or remote industrial communication links quickly depletes
available bandwidth, causing unacceptable telemetry ingestion latencies [10], [11].
2. Corporate and
Compliance Bottlenecks: The conservativeness of industrial stakeholders as
far as data sovereignty is concerned is intense [12]. Proprietary operational
recipes, production rates, and health profiles of assets are often encapsulated
in network logs; these logs being exposed to centralized third-party cloud
engines breach strict cooperate compliance directives and international regulations
including the General Data Protection Regulation GDPR [13], [14], [15].
3. Single Point of
Failure Vulnerabilities: Centralized storage repositories are high value
targets for malicious state-sponsored adversaries, in which one breach can
result in the systemic intelligence of the entire industrial matrix being
compromised [16], [17], [18].
In order to resolve the structural conflict between the cooperative
threat intelligence and the local data confidentiality, Decentralized
Collaborative Machine Learning particularly, Federated Learning (FL) is an
innovative structural change [19], [20]. The basic idea of FL is that the
training process of the model is no longer tied to the need to access data
directly [20], [21] . In an FL orchestration, isolated industrial edge nodes
make use of their local compute resources to independently train private
intrusion detection models over their localized network traffic streams [12],
[21]. These nodes do not send raw packet data, but instead, only the parameters
of their abstracted mathematical model (i.e., weights and biases) are uploaded
to a coordinating central server [21], [23].
The central server then runs an aggregation protocol to keep these
localized updates together in a very generalized global defensive model that is
then redistributed to edge nodes [19], [24], [25], [26]. Although this is
theoretically a promising approach, when applied in practice to real-world
industrial environments, the implementation of standard FL frameworks (such as
Vanilla Federated Averaging - FedAvg) reveals a critical weakness: Statistical
Heterogeneity (Non-IID Data) [27], [28]. In practice applications, an
industrial water purification facility, a car assembly line, and a smart power
grid can be configured to use completely different communication protocols
(e.g., Modbus, MQTT, OPC UA), different traffic rates [29], [30].
As a result, their local network logs are non-Identically and
Independently Distributed (non-IID) [27], [31]. In the case of vanilla
aggregation of highly stratified non-IID data, the local updates computed by
independent nodes are susceptible to pull the global model towards isolated,
divergent local optima, a pathological vulnerability referred to as client drift
[28], [32], [33]. This structural drift causes disastrous convergence of models
with unstable classification oscillations and intolerably large false-alarm
rates to the extent of disruption of normal factory operation [31], [33], [34].
1.1 Motivation and Contributions
The nature of the study is to develop a highly robust,
privacy-protecting, decentralized intrusion detection system that is capable of
sustaining a stable and high-accuracy defensive running in the face of extreme
statistical non-IID data skewness between heterogeneous industrial edge nodes.
To fill the gap between pragmatic deployment considerations and theoretical
knowledge, we suggest a regularized federated learning architecture that acts
to proactively resolve the client drift issue without undermining the privacy
of the edge nodes. Explicit, high-impact contributions of the current study
might be as follows:
Architecting a Non-IID Resilient Edge IDS:
We propose and implement a localized deep-learning-based IDS, and a
parameterized proximal regularization method to stabilize global models’
convergence in the presence of extreme statistical divergence environments at
multi-tenant industrial nodes.
Multi-Dimensional Performance
Instrumentation: Unlike the traditional literature where the accuracy is
only evaluated, we do construct a complete academic benchmarking suite which
evaluates the co-evolution of Classification Accuracy, Global Objective Loss
Convergence and Confusion Matrix Profiling.
Empirical Validation on Modern Protocols: Our
simulation is based on a very structured telemetry distribution in accordance
with the modern Edge-IIoT set benchmark to capture the multi-protocol dynamics
of cyber-attacks (including malicious MQTT and Modbus injections).
A Complete Open-Source MATLAB Simulation
Engine: Our implementation is a fully instrumented structurally transparent
MATLAB simulation of isolated edge client optimization loops and centralized
server weight replacements, which become an open-source baseline of future OT
security research.
2. Proposed Methodology
The following section describes the architectural mechanics and
mathematical basis of our privacy-preserving decentralized intrusion detection
system that has been developed to survive strong non-IID skewness of an
industrial grid [19], [20].
2.1 System and Network Architecture
The decentralized framework models
an enterprise industrial ecosystem consisting of K distinct, geographically
isolated or structurally siloed IIoT edge gateways, coordinated by a central
orchestration server located either in a private corporate cloud or a
high-availability node [21], [35]. Each local edge gatewayacts as a defensive sentry for its specific manufacturing
cell, collecting real-time network traffic logs into a highly localized private
dataset
[12], [22].
The aggregate global dataset across
the entire industrial matrix is denoted as, with the total sample volume quantified as
[19], [21].
To strictly enforce data privacy as well as zero-trust corporate
compliance, raw packet captures within never exit the boundary regarding edge node k [11], [36]. The
unified optimization objective of collective network is to minimize the global
empirical risk function
over shared vector of deep neural network weights
[19], [15], [37]:
where is the local loss function evaluated over the private data
silo of client
[19], [38], [39]:
Here, represents the cross-entropy loss tracking the classification
errors of the network model
mapping input telemetry features
to cyber-attack labels
.
2.2 Mathematical Formulation of Proposed Proximal Regularization
(FedProx)
In a vanilla federated execution (FedAvg), each client performs a
fixed number of local stochastic gradient descent (SGD) iterations based on the
global state distributed at the beginning of round t, updating its local
parameters to
[19], [28]. In the case when underlying datasets
are intensely non-IID,
local gradients point toward wildly different geometric directions [27], [32].
Under simple averaging of standard FedAvg, such unconstrained updates cancel
each other out or destabilize global parameter path, showing severe accuracy
drop-offs and optimization divergence [28], [33].
For resolving such problem, our approach replaces vanilla
optimization loop with structured proximal regularization algorithm (FedProx) [38].
At each global communication round, rather than minimizing standard empirical risk
each local IIoT node k is forced to optimize regularized
local objective function
mathematically specified as [38], [32]:
where:
is the original
empirical local classification loss over the node's distinct traffic data [19].
is the static global model parameter vector received from the
central server at the start of the current round [21].
represents the evolving local model parameter vector during
the current edge training epochs [23].
is a meticulously calibrated proximal regularization
hyperparameter (configured to
in our experimental pipeline) [38].
The Role of the Proximal Penalty Term: The quadratic expansion
component functions as a dynamic
gravitational constraint [38], [40]. It continuously calculates the
norm distance between
the local model weights w and the global reference model
[41], [42].
The proximal term penalizes extreme change when an edge node
experiences an extreme skew, of a certain type of attack (e.g., a giant
localized DDoS burst), such that the local gradients of the model would need to
change in response [38], [43]. The frameworks were mathematical constraints
that bounded the effect on client drift by limiting how far individual node
could wander relative to common global baseline [31]. This ensures that the
consolidated model is not only high-accuracy generalized to all the
participating factory nodes, but it is also not brought down by localized
statistical anomalies 38], [33].
2.3 Deep Neural Network Structural Design
The classification engine utilized is the same on the server as on
the edge nodes constituting optimized, deep feed-forward neural network
specially constructed to run fast inference and high discrimination with
tabular network logs [9], [44]. The model projects an 8-dimensional input
feature space to a multi-class probability distribution. The architecture of
the layer graph is [45], [46]:
1.
Feature
Input Layer: Accepts normalized 8-dimensional continuous vector streams
representing core industrial packet indicators:
2. Fully Connected Layer 1:
Expands the input space to 64 hidden neurons via a matrix dot product and bias
addition: , where
.
3. ReLU Activation Layer 1:
Introduces non-linear thresholding:, eliminating gradient saturation.
4. Fully Connected Layer 2:
Compresses information into 32 hidden neurons:, isolating complex cross-protocol features.
5. ReLU Activation Layer 2: Applies
identical element-wise rectifying activation:.
6. Output Bridge Fully
Connected Layer: Maps the hidden features directly to the target classification
dimensions: where
(
being the target
categorical class count).
7. Softmax Transformation
Layer: Converts raw logits into continuous probabilistic scores spanning a
bounded range [0, 1]:
8. Classification Output
Layer: Computes cross-entropy loss against ground-truth categorical assignments
during training and outputs the maximum-likelihood binary or multi-class
assignment during real-time edge monitoring.
3. Experimental Implementation and Testbed Instrumentation
In order to test the theoretical arguments of the proposed proximal
regularized IDS framework rigorously, a highly advanced distributed simulation
platform was completely executed in MATLAB (R2024a/2026b). In this part, the
setup of the experimental environment, the strategies used for dataset
partitioning, and the structural verification code that were executed to
produce the empirical results are outlined.
3.1 Dataset Preprocessing and Realistic Non-IID Partitioning
The empirical pipeline operates on data structures that are
directly matching to the Edge-IIoTset benchmark dataset that represents a
modern IoT/IIoT security testing corpus capturing practical industrial protocol
settings (e.g., MQTT, Modbus, HTTP) with targeted malicious exploits [30], [47].
To test the system with extreme statistical heterogeneity, 15,000 various
telemetry records are ingested, 8 critical structural network features, and
categorical attack vectors [29], [30].
In order to simulate a fragmented, adversarial industrial network,
in which the distribution of the data across the factories is extremely
unbalanced, the data are subjected to a strict Non-IID Stratification Sequence
[48], [49], [50]. The entire corpus is ordered according to its attack label
distribution (sortIdx), the legitimate records and the different logs of
cyber-exploits are grouped into isolated blocks. The sorted samples are then
cut into separate unequal segments and distributed to 3 separate edge client
cell matrices (X_Local, Y_Local). This ensures that Client 1 is inundated
heavily with benign operational logs, and that Client 2 and Client 3 have their
own, very skewed segments of malicious injections. This forms a severe mismatch
on an extreme scale, which presents an extreme convergence difficulty to
conventional decentralized learning algorithms.
4. Empirical Results and Detailed Discussion
This part presents an analytical, in-depth exploration of the
15-dimensional outputs that are yielded during the 15 rounds of the global
federation.
4.1 Convergence and Accuracy Analysis under Non-IID Skewness
The accuracy and cross-entropy loss trajectories of the global
model are co-evolved in real-time and plotted in Fig. 2.
Fig 2.
The real-time co-evolution of the global model’s accuracy and cross-entropy
loss trajectories.
An introductory analysis of the empirical patterns shows that there
is a sharp performance divergence. Vanilla FedAvg (Baseline) is highly
mathematically unstable and its classification performance is violently
oscillating with a minimum bound of 58.1% and a high threshold of only 74.3%.
It does not show steady upward learning curve. This disastrous pattern is the
actual real-life manifestation of client drift. Since the data partitioning
sequence maps highly separated, polarized subsets of Edge-IIoTset packet
dynamics to the various edge clients, each local gateway updates its own model
parameters to vastly distinct geographic coordinates in the optimization space.
When the cloud server averages the plain average regarding such divergent
vectors, the parameters are in conflict, which weakens predictive abilities of
the model.
On the contrary, the suggested FedProx Framework establishes significantly
smooth, high-stability monotonic convergence profile. Through activating
proximal regularization modifier, the system penalizes gradient adjustments that venture
outside global consensus boundaries. Thus, local optimization paths are
constrained, suppressing individual weight divergence. The suggested global
model surpasses the baseline by round 3, reduces erratic fluctuations, and
achieves stable, definitive classification accuracy of 91.4% at round 15.
At the same time, the global objective loss trajectory (Fig. 2,
Right Subplot) is mathematically validated by inspection. Vanilla FedAvg
training profile has an erratic plateau because it has constant gradient
conflict. The Proposed FedProx loss curve, on the other hand, shows a healthy
continuous exponential decay, and indicates that the structural inclusion of
the proximal regularization term can indeed stabilize the multi-tenant
optimization trajectory with extreme statistical heterogeneity.
4.2 Statistical Robustness via Confusion Matrix Profiling
In an effort to comprehensively check the multi-class
categorization integrity of the frameworks, Fig. 3 plots the cross-normalized
confusion charts obtained at the final cycle of communication.
Fig 3. The cross-normalized confusion charts
extracted at the final communication cycle.
The critical operational performance can be studied by analyzing
the row summaries. The baseline Vanilla FedAvg model is highly statistically
vulnerable with a high False Positive (FP) inflation rate. To be more exact, it
misclassifies a significant amount of legitimate baseline control packets as
malicious anomalies. Such unstable false alarm profile would not be sustainable
in operational manufacturing facility, where it would lead to frequent and
unnecessary line shutdowns and delays in operations.
On the other hand, the Proposed FedProx framework creates a
definite statistical advantage. True Positive Rate (TPR / Recall) of the normal
operation logs and of the active network injections reaches maximum efficiency
and the False Positive Rate (FPR) is kept to the minimum. This model effectively
classifies small deviations in protocols without confusing normal industrial
control pattern and malicious payloads. This shows that the localized model
drift can be directly constrained to optimize the statistical accuracy of the
decentralized defensive models.
5. Conclusion and Future
Work
The presented work has efficiently provided a privacy-preserving,
robust decentralized intrusion detection system that is particularly engineered
for withstanding intense statistical non-IID data heterogeneity across
multi-protocol Industrial IoT environments. Through transforming conventional
unconstrained local optimization loop into regularized objective through
parameterized proximal constraint, our framework efficiently mitigates the destructive client
drift phenomenon inherent in fragmented industrial data silos.
Large-scale multi-dimensional empirical benchmarking showed that
whereas conventional Vanilla Federated Averaging (FedAvg) shows violent
optimization oscillations and suboptimal generalization behavior with respect
to skewed data distributions, the regularized architecture proposed an
architecture offers stable, monotonic convergence, and achieves optimal
security classification rates of 91.4 percent with low false alarms. This makes
it viable in the protection of high-security OT. We will extend this work in
two main directions in future research:
1. Cryptographic
Privacy Upgrades: Combining localized Differential Privacy (DP) and
lightweight Secure Multi-Party Computation (sMPC) protocols to protect the
parameter exchange channel against advanced gradient leakage and model
inversion attacks.
2. Resource-Constrained
Optimization: Testing the performance of the framework on the lossy and
bandwidth-constrained dynamic wireless networks to define low-latency,
resilience defensive intelligence on the highly volatile physical edge
infrastructures.
REFERENCES
[1] P. Kairouz et al., “Advances and open
problems in federated learning,” Found. Trends Mach. Learn., vol. 14, nos. 1-2,
pp. 1-210, 2021.
[2] D. C. Nguyen et al., “Federated learning
for Internet of Things: A comprehensive survey,” IEEE Commun. Surveys Tuts.,
vol. 23, no. 3, pp. 1622-1658, 3rd Quart., 2021.
[3] A. Brecko, E. Kajati, J. Koziorek, and I.
Zolotova, “Federated learning for edge computing: A survey,” Appl. Sci., vol.
12, no. 18, Art. no. 9124, Sep. 2022.
[4] J. Lansky, S. Ali, M. Mohammadi, M. K.
Majeed, S. H. T. Karim, S. Rashidi, M. Hosseinzadeh, and A. M. Rahmani, “Deep
learning-based intrusion detection systems: A systematic review,” IEEE Access,
vol. 9, pp. 101574-101599, 2021.
[5] M. A. Ferrag, O. Friha, L. Maglaras, H.
Janicke, and L. Shu, “Federated deep learning for cybersecurity in the Internet
of Things: Concepts, applications, and experimental analysis,” IEEE Access,
vol. 9, pp. 138509-138542, 2021.
[6] K. S. Pokkuluri, A. Kumar, K. K. Singh
Gautam, P. Deshmukh, and L. Abualigah, “Collaborative intelligence for IoT:
Decentralized network security and confidentiality,” J. Intell. Syst. Internet
Things, vol. 13, no. 2, 2024.
[7] E. Gyamfi and A. Jurcut, “Intrusion
detection in Internet of Things systems: A review on design approaches
leveraging multi-access edge computing, machine learning, and datasets,”
Sensors, vol. 22, no. 10, Art. no. 3744, 2022.
[8] A. Aldweesh, A. Derhab, and A. Z. Emam,
“Deep learning approaches for anomaly-based intrusion detection systems: A
survey, taxonomy, and open issues,” Knowl.-Based Syst., vol. 189, Art. no.
105124, 2020.
[9] Elmobark, N., Abdzaid, A. Y., Alaa, F.,
& Saad, A. (2026). Experimental Evaluation of Automated and Manual Data
Cleaning Systems: A Case Study Using Organizational Data. International Journal
of Theoretical & Applied Computational Intelligence, 2026, 76-103.
[10] M. Chen, N. Shlezinger, H. V. Poor, Y. C.
Eldar, and S. Cui, “Communication-efficient federated learning,” Proc. Natl.
Acad. Sci. USA, vol. 118, no. 17, Art. no. e2024789118, 2021.
[11] A. Reisizadeh et al., “FedPAQ: A
communication-efficient federated learning framework with quantized updates,”
in Proc. 23rd Int. Conf. Artif. Intell. Stat. (AISTATS), 2020, pp. 2421-2431.
[12] Y. Chen, L. Liu, Y. Ping, M. Atiquzzaman,
S. Mumtaz, Z. Zhang, M. Guizani, and Z. Tian, “A lightweight and fair
privacy-preserving federated learning framework for IoT,” IEEE Trans. Netw.
Service Manag., vol. 21, no. 5, pp. 5843-5858, 2024.
[13] R. C. Geyer, T. Klein, and B. Mair,
“Differentially private federated learning: A client-level perspective,”
arXiv:1712.07557, 2017.
[14] S. Truex, N. Baracaldo, A. Anwar, T.
Steinke, H. Ludwig, R. Zhang, and Y. Zhou, “A hybrid approach to
privacy-preserving federated learning,” in Proc. 12th ACM Workshop Artif.
Intell. Security, Nov. 2019, pp. 1-11.
[15] M. Mohri, G. Sivek, and A. T. Suresh,
“Agnostic federated learning,” in Proc. 36th Int. Conf. Mach. Learn. (ICML),
2019, pp. 4615-4625.
[16] H. HaddadPajouh et al., “A survey on
multi-access edge computing security in IIoT environments,” Digit. Commun.
Netw., vol. 7, no. 4, pp. 512-526, 2021.
[17] V. Mothukuri et al., “A survey on security
and privacy of federated learning,” Comput. Sci. Rev., vol. 39, Art. no.
100340, 2021.
[18] L. Lyu, H. Yu, and Q. Yang, “Threats to
federated learning: A survey,” arXiv:2003.02133, 2020.
[19] Elmobark, N., Saad, A., Hasan, S. H.,
& Badouch, M. (2026). The Hadoop Ecosystem: An Open-Source Framework for
Enterprise-Scale Big Data Processing and Analytics. Engineering Headway, 35,
210-226.
[20] Q. Yang, Y. Liu, T. Chen, and Y. Tong,
“Federated machine learning: Concept and applications,” ACM Trans. Intell.
Syst. Technol., vol. 10, no. 2, pp. 1-19, 2019.
[21] Y. Baseri, A. S. Hafid, D. Makrakis, and
H. Fereidouni, “Privacy-preserving federated learning framework for risk-based
adaptive authentication,” arXiv:2508.18453, 2025.
[22] K. Bonawitz et al., “Towards federated
learning at scale: System design,” Proc. Mach. Learn. Syst. (MLSys), vol. 1,
pp. 374-388, 2019.
[23] Z. Zhang, S. Rath, J. Xu, and T. Xiao,
“Federated learning for smart grid: A survey on applications and potential
vulnerabilities,” ACM Trans. Cyber-Phys. Syst., vol. 10, no. 1, pp. 1-26, 2026.
[24] V. Smith, C. K. Chiang, M. Sanjabi, and A.
Talwalkar, “Federated multi-task learning,” in Advances in Neural Information
Processing Systems (NeurIPS), vol. 30, pp. 4424-4434, 2017.
[25] S. Caldas et al., “LEAF: A benchmark for
federated settings,” arXiv:1812.01097, 2018.
[26] C. He et al., “FedML: A research library
and benchmark for federated machine learning,” in Proc. NeurIPS Workshop, 2020.
[27] Y. Zhao et al., “Federated learning with
non-IID data,” arXiv:1806.00582, 2018.
[28] X. Li, K. Huang, W. Yang, S. Wang, and Z.
Zhang, “On the convergence of FedAvg on non-IID data,” in Proc. Int. Conf.
Learn. Representations (ICLR), 2020.
[29] M. Alemayehu, M. C. Ghanem, H. Kheddar, D.
Dunsin, and M. J. Lacerda, “Systematic analysis on the use of AI techniques in
industrial IoT DDoS attack detection, mitigation, and prevention,” IoT, vol. 7,
no. 51, pp. 1-56, 2026.
[30] Saad, A. (2026, January). Federated
Learning-Based Intrusion Detection System Using Convolutional Neural Networks
for IoT Networks. In 2026 5th International Conference on Electrical,
Computer & Telecommunication Engineering (ICECTE) (pp. 1-6). IEEE.
[31] S. K. A. Hashim, Y. B. M. Yussoff, and S.
B. Shahbudin, “Mitigating zero-day vulnerabilities in IIoT systems: Challenges
and advances in AI-powered intrusion detection systems,” Mesopotamian J.
CyberSecurity, vol. 5, no. 3, pp. 1184-1198, 2025.
[32] C. Y. Huang, K. Srinivas, X. Zhang, and X.
Li, “Overcoming data and model heterogeneities in decentralized federated
learning via synthetic anchors,” arXiv:2405.11525, 2024.
[33] L. Yuan, J. Zhang, M. Duan, G. Xiao, Z.
Tang, and K. Li, “PRFL: Personalized and robust federated learning for non-IID
data with malicious participants,” IEEE Trans. Mobile Comput., early access,
2025.
[34] A. Yazdinejad, A. Dehghantanha, H.
Karimipour, G. Srivastava, and R. M. Parizi, “A robust privacy-preserving
federated learning model against model poisoning attacks,” IEEE Trans. Inf.
Forensics Security, vol. 19, pp. 6693-6708, 2024.
[35] W. Y. B. Lim et al., “Federated learning
in mobile edge networks: A comprehensive survey,” IEEE Commun. Surveys Tuts.,
vol. 22, no. 3, pp. 2031-2063, 3rd Quart., 2020.
[36] Saad, A., Hasan, N. F., & Altaher, A.
W. (2025). Deep Reinforcement Learning-Based Network Intrusion Prevention in
Cloud-Edge Architectures.
[37] S. Reddi et al., “Adaptive federated
optimization,” in Proc. Int. Conf. Learn. Representations (ICLR), 2021.
[38] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi,
A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,”
Proc. Mach. Learn. Syst. (MLSys), vol. 2, pp. 429-450, 2020.
[39] H. Liu, X. Zhang, X. Shen, and H. Sun, “A
federated learning framework for smart grids: Securing power traces in
collaborative learning,” arXiv:2103.11870, 2021.
[40] L. Collins, H. Hassani, A. Mokhtari, and
S. Shakkottai, “Exploiting shared representations for personalized federated
learning,” in Proc. 38th Int. Conf. Mach. Learn. (ICML), 2021, pp. 2089-2099.
[41] A. Fallah, A. Mokhtari, and A. Ozdaglar,
“Personalized federated learning with theoretical guarantees: A model-agnostic
meta-learning approach,” in Advances in Neural Information Processing Systems
(NeurIPS), vol. 33, pp. 3557-3568, 2020.
[42] C. T. Dinh, N. H. Nguyen, N. H. Tran, and
W. Zhang, “Personalized federated learning with Moreau envelopes,” in Advances
in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 21396-21407,
2020.
[43] K. Pillutla, S. M. Kakade, and Z.
Harchaoui, “Robust aggregation for federated learning using geometric medians,”
IEEE Trans. Signal Process., vol. 70, pp. 1211-1225, 2022.
[44] M. F. L. Becerra-Suarez, V. A.
Tuesta-Monteza, H. I. Mejia-Cabrera, and J. Arcila-Diaz, “Performance
evaluation of deep learning models for classifying cybersecurity attacks in IoT
networks,” Informatics, vol. 11, no. 2, Art. no. 32, May 2024.
[45] A. Hard et al., “Federated learning for
mobile keyboard prediction,” arXiv:1811.03604, 2018.
[46] J. Xu et al., “Federated learning for
healthcare informatics: Challenges and applications,” IEEE J. Biomed. Health
Inform., vol. 25, no. 8, pp. 2821-2831, Aug. 2021.
[47] M. A. Ferrag, O. Friha, D. Hamouda, L.
Maglaras, and H. Janicke, “Edge-IIoTset: A new comprehensive realistic
cybersecurity dataset of IoT and IIoT applications for centralized and
federated learning,” IEEE Access, vol. 10, pp. 40281-40306, 2022.
[48] P. Blanchard et al., “Machine learning
with adversaries: Byzantine-tolerant gradient descent,” in Advances in Neural
Information Processing Systems (NeurIPS), vol. 30, pp. 119-129, 2017.
[49] C. Markarian and A. Panthakkan,
“Algorithmic self-repair: Frontiers in fault-tolerant computation,” Front.
Comput. Sci., vol. 8, Art. no. 1717711, 2026.
[50] J. Tu, L. Yang, and J. Cao, “Distributed
machine learning in edge computing: Challenges, solutions, and future
directions,” ACM Comput. Surv., vol. 57, no. 5, pp. 1-37, 2025.