Multi-source federal cross-domain and source-domain enhanced millimeter wave action recognition method and system

Through the multi-source federated cross-domain and source domain enhanced millimeter wave action recognition method, federated learning and weighted knowledge aggregation technology are used to solve the domain offset problem in cross-environment and cross-user scenarios, efficient human movement recognition is achieved, and user privacy is protected.

CN120180221APending Publication Date: 2025-06-20XI AN JIAOTONG UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510256386.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has domain offset problems in cross-environment and cross-user scenarios, resulting in a significant decline in human body movement recognition performance, and traditional domain adaptation methods are difficult to protect user privacy and data security.

Method used

The millimeter wave action recognition method with multi-source federated cross-domain and source domain enhancement is adopted to dynamically evaluate and fuse knowledge of multiple source domains through federated learning and weighted knowledge aggregation, optimize the generalization ability of the target model, and realize unsupervised learning through a pseudo-label method based on voting.

Benefits of technology

It significantly improves the adaptability and performance of the millimeter wave human motion recognition system in the new environment, protects user privacy, reduces the dependence on labeled data, and effectively avoids negative migration, achieving an accuracy rate comparable to supervised learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180221A_ABST
    Figure CN120180221A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source federal cross-domain and source-domain enhanced millimeter wave action recognition method and system, and the method comprises the steps: generating a micro-Doppler spectrogram through a millimeter wave radar, and extracting the motion features of human body actions through a signal processing module; in the federal multi-source domain adaptation module, dynamically evaluating and fusing knowledge of a plurality of source domains by adopting a voting-based pseudo-tag method and a weighted knowledge aggregation mechanism, and optimizing the generalization ability of a target model; and through a generalization gap optimization method, the performance of the source domain model is improved, and the robustness of the system in different environments is ensured. Through combination of a federated learning framework and a multi-source domain adaptation technology, unsupervised learning under the condition that a target domain has no annotated data is realized, only a single set of millimeter wave equipment is needed, a millimeter wave communication protocol is compatible, and the method has the characteristics of privacy protection, unsupervised learning, multi-source knowledge fusion and strong generalization ability. The method is suitable for application scenes of smart home, health monitoring, man-machine interaction and the like, and has wide practical application value and research prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Non-contact human activity recognition (HAR) is crucial for creating a natural user experience in various applications. For example, users can automatically adjust lighting, temperature, and music according to predefined activities, enabling seamless interaction with devices. In addition, caregivers can enhance the safety of the elderly by monitoring their daily activities. Traditional HAR utilizes computer vision technology or wearable devices as activity sensors. However, these methods often result in a poor user experience.

[0003] In recent years, non-contact HAR has emerged as a viable alternative. Researchers have explored using WiFi, RFID, ultrasonic, and millimeter-wave signals to support various human-computer interaction (HCI) applications. Among them, millimeter-wave signals provide fine and robust sensing capabilities, supporting applications such as vibration detection, vital sign monitoring, and silent speech recognition. As millimeter-wave radio frequencies are expected to become widespread in 5G Internet of Things devices, they may become powerful and ubiquitous sensing tools, highly suitable for HAR. With the help of deep learning technology, non-contact human activity recognition (HAR) has achieved remarkable accuracy within a single environment (i.e., "domain"). However, due to the inherent characteristics of wireless signals and the existence of domain shift, directly applying such models to a new environment usually leads to a significant performance decline. Domain adaptation is a promising technique that addresses the problems of the target domain by leveraging the transferable features of the source domain. Some of the latest domain adaptation methods propose constructing source-target pairs by combining data from two domains and achieving knowledge transfer by minimizing the feature distance between them.

[0004] However, certain relevant data (such as personal health records or hospital activity information) are usually stored on local devices and are subject to different privacy protection policies, making the above methods difficult to apply in practice. Other methods pre-train the HAR model in the source domain and fine-tune it using a small number of labeled samples in the target domain. These methods perform well when labeled samples are available, but for most end-users, labeling radio frequency signals is both complex and time-consuming. In addition, significant domain shift may cause these methods to fail because a small number of labeled samples may not be sufficient to achieve effective adaptation and may even lead to negative transfer. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a millimeter-wave action recognition method and system with multi-source federated cross-domain and source-domain enhancement in view of the deficiencies in the above-mentioned prior art. Based on a single set of millimeter-wave devices, a federated learning method is used with multiple source-domain models to train a new model in a new environment (domain) without data labels, while improving the performance of the participating source-domain models, so as to solve the technical problem of domain shift in cross-environment and cross-user scenarios in the prior art.

[0006] The present invention adopts the following technical solutions: A millimeter-wave action recognition method with multi-source federated cross-domain and source-domain enhancement, comprising the following steps: Transmit multiple chirp signals with linearly increasing frequencies, and receive the echo signals reflected by the object, and generate a micro-Doppler spectrogram using the obtained millimeter-wave radar; Adopt a voting-based pseudo-label method and a weighted knowledge aggregation mechanism to dynamically evaluate and fuse the knowledge of multiple source domains, and optimize the generalization ability of the server-side target model; Use the obtained target model to analyze the micro-Doppler heat map, identify different human actions in the target domain, and realize unsupervised human action recognition.

[0007] Preferably, generating the micro-Doppler spectrogram specifically is: Combine the chirp signal and the echo signal to generate an intermediate-frequency signal of the millimeter-wave radar, perform a one-dimensional fast Fourier transform on the intermediate-frequency signal of the millimeter-wave radar to obtain distance information, then perform a fast Fourier transform on the distance information to generate a range-Doppler heat map, and then obtain a micro-Doppler spectrogram showing the speed change during a specific activity.

[0008] Preferably, the intermediate-frequency signal of the millimeter-wave radar is:

[0009] wherein, is the amplitude after signal attenuation, is the starting frequency, is the bandwidth, is the chirp duration, is the time delay of the received signal relative to the transmitted signal.

[0010] Preferably, adopting the voting-based pseudo-label method is specifically as follows: Suppose there are trained source models , for each sample in the target model, use to represent the prediction result of the corresponding classifier, regard the maximum prediction value as the label, introduce a confidence threshold , and only retain the predictions whose confidence exceeds the confidence threshold As a result, the prediction result of the remaining model determines the final label through a voting mechanism. A group of high-confidence source models that support the voting result are obtained through voting. The average prediction of the high-confidence source models is used as the recognized knowledge, and the KL divergence loss is used to train the target model.

[0011] Preferably, the KL divergence loss is as follows:

[0012] where is the reweighting factor of the loss, is the divergence calculation function, is the prediction result of the target model for the j-th sample, is the average result of the prediction of the j-th sample by the high-confidence source models.

[0013] Preferably, the weighted knowledge aggregation mechanism is specifically as follows: Define the total quality obtained through voting, determine the contribution degree of the source model , and based on the contribution degree calculate the weights of the target model and each source model respectively. In each training iteration, the parameters of all source models and the target model are aggregated according to the weights, and the updated parameters of the target model's feature extractor will be broadcast to all source models for the next iteration of training.

[0014] Preferably, the updated parameters of the target model's feature extractor are:

[0015] where and respectively represent the feature extractor parameters of the target model and the source model, represents the number of source models participating in the training, and respectively represent the weighting factors of the target model and the i-th source model.

[0016] Preferably, optimizing the generalization ability of the target model on the server side is specifically as follows: In the first stage, use the generalization optimization algorithm to optimize the model parameters, generate a global feature map using the broadcast by the target model after training, and calculate the generalization gap loss ; In the second stage, use the labeled dataset to optimize the parameters and train through the cross-entropy loss function.

[0017] Preferably, the generalization gap loss and respectively is:

[0018]

[0019] wherein, is the divergence calculation function, is the classifier of the target model, is the feature extractor of the target model, is the source model dataset, is the parameter of the feature extractor of the target model, is the pre-training result corresponding to the source model, is the feature extractor of the source model, is the parameter of the feature extractor of the source model, is the cross-entropy loss of the source model, is the number of samples included in the i-th source model dataset, is the class label of the j-th sample in the i-th source model dataset, is the predicted class of the j-th sample in the i-th source model dataset.

[0020] In a second aspect, an embodiment of the present invention provides a multi-source federated cross-domain and source domain enhanced millimeter-wave action recognition system, including: A generation module that emits multiple chirp signals with linearly increasing frequencies, receives the echo signals reflected by an object, and generates a micro-Doppler spectrogram using the obtained millimeter-wave radar; A generalization module that adopts a voting-based pseudo-label method and a weighted knowledge aggregation mechanism to dynamically evaluate and fuse the knowledge of multiple source domains and optimize the generalization ability of the target model on the server side; An identification module that analyzes the micro-Doppler heat map using the obtained target model to identify different human actions in the target domain and realizes unsupervised human action recognition.

[0021] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above multi-source federated cross-domain and source domain enhanced millimeter-wave action recognition method are implemented.

[0022] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium including a computer program. When the computer program is executed by a processor, the steps of the above multi-source federated cross-domain and source domain enhanced millimeter-wave action recognition method are implemented.

[0023] Fifth aspect, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned multi-source federated cross-domain and source-domain enhanced millimeter-wave action recognition method are implemented.

[0024] Sixth aspect, an embodiment of the present invention provides an electronic device including a computer program. When the computer program is executed by the electronic device, the steps of the above-mentioned multi-source federated cross-domain and source-domain enhanced millimeter-wave action recognition method are implemented.

[0025] Compared with the prior art, the present invention has at least the following beneficial effects: A multi-source federated cross-domain and source-domain enhanced millimeter-wave action recognition method effectively protects the privacy and data security. Traditional domain adaptation methods usually require training the data of the source domain and the target domain on a single server. When it comes to personal health records or sensitive information, this may lead to the leakage of user privacy data. By adopting the federated learning framework, the data of the source domain always remains local, and only the model parameters are uploaded, avoiding the transmission and centralized storage of the original data, thus effectively protecting user privacy.

[0026] Furthermore, unsupervised domain adaptation is adopted. Most existing domain adaptation methods require labeled data in the target domain for fine-tuning. Especially in a new environment, labeling radio frequency (RF) signals is usually time-consuming and complex. The present invention can perform unsupervised learning without labeled data in the target domain through a voting-based pseudo-label method, significantly reducing the need for manual labeling and improving the practicality of the system.

[0027] Furthermore, through multi-source knowledge fusion and weighted aggregation, a contribution evaluation mechanism is used to dynamically evaluate the contribution of each source domain to the target domain, and weighted knowledge aggregation is performed based on the contribution. Compared with traditional multi-source domain adaptation methods that usually use simple parameter averaging, ignoring the contribution differences of different source domains to the target domain, which may lead to negative transfer, the present invention effectively avoids the negative impact of low-quality source domains and improves the performance of the target model.

[0028] Furthermore, through an optimization method based on the generalization gap, not only the performance of the target model is improved, but also the generalization ability of the source domain model is enhanced. Compared with existing methods that usually only focus on improving the performance of the target domain and ignore the generalization ability of the source domain, resulting in the performance of the source domain model may decline in a new environment, the present invention enables the source domain models participating in the federated learning to also obtain performance improvement.

[0029] Furthermore, it has cross - domain adaptability and remarkable practical application effects. Through micro - Doppler heatmaps and multi - source knowledge transfer, it can effectively address the domain shift problems brought about by environmental changes and user differences. When facing environmental changes (such as multipath effects and user individual differences), the performance of existing methods drops significantly, especially in the wireless signal sensing domain where the domain shift problem is particularly prominent. The present invention achieves an accuracy comparable to supervised learning in the target domain. In addition, experimental results show that the present invention performs excellently in five different environments, with the average accuracy in the target domain reaching 92%, approaching the results of supervised learning. At the same time, the performance of the source - domain model has also been significantly improved. Compared with existing methods that require a large amount of labeled data and complex parameter - tuning processes in practical applications and show unstable performance in new environments, it has high efficiency and practicality in practical applications.

[0030] Furthermore, it reduces the risk of negative transfer. Through a voting mechanism and contribution evaluation, it can effectively filter out low - quality source domains, reduce the risk of negative transfer, and ensure that the target model can obtain beneficial knowledge from high - quality source domains, avoiding the negative transfer that may be caused by low - quality source domains in existing methods and affecting the performance of the target model.

[0031] Furthermore, it is applicable to dynamic and heterogeneous environments. Through multi - source knowledge transfer and generalization optimization, it can adapt to dynamic and heterogeneous environments and is applicable to various application scenarios such as smart homes, health monitoring, and human - computer interaction. Compared with existing methods that usually assume relatively stable environments in the source domain and target domain, the present invention solves the problem of domain shift in dynamic and heterogeneous environments that is difficult for existing methods to handle.

[0032] It can be understood that the beneficial effects of the second to sixth aspects above can refer to the relevant descriptions in the first aspect above and will not be elaborated here.

[0033] In summary, through technologies such as federated learning, multi - source knowledge transfer, unsupervised learning, and generalization optimization, the present invention, a millimeter - wave action recognition technology based on multi - source federated cross - domain and source - domain enhancement, significantly improves the adaptability and performance of the millimeter - wave human action recognition system in new environments, while protecting user privacy, reducing the dependence on labeled data, and having good application value, research prospects, and development potential.

[0034] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0035] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments of the present application will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0036] Figure 1 This is the flow chart of the present invention; Figure 2 This is the specific process diagram of the federated training of the present invention; Figure 3 This is the experimental scenario diagram of the present invention; Figure 4 This is the experimental gesture diagram; Figure 5 This is the self-supervised experimental result diagram of each environmental domain; Figure 6 This is the comparison experiment and ablation experiment accuracy comparison diagram of the present invention in the gesture recognition task; Figure 7 This is the accuracy comparison diagram of the present invention using different numbers of source domains in the gesture recognition task; Figure 8 This is the accuracy comparison diagram of the present invention using different proportions of sample numbers in the source domain in the gesture recognition task; Figure 9 This is the accuracy diagram of the present invention using different user samples in the source domain in the gesture recognition task; Figure 10 This is the experimental result diagram of each user of the present invention using different user samples in the source domain in the gesture recognition task; Figure 11 This is the experimental result diagram of each user of the present invention using different user samples in the source domain in different domains in the gesture recognition task; Figure 12 This is the accuracy diagram of whether there are new users participating in the source domain of the present invention in the gesture recognition task; Figure 13 This is the performance schematic diagram of FMDA at different distances and directions between users and the radar Figure 14 This is the schematic diagram of the computer device provided by an embodiment of the present invention; Figure 15 This is the block diagram of an electronic device provided by an embodiment of the present invention.

[0037] Among them, 60. Computer device; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access storage unit; 6202. Cache storage unit; 6203. Read-only storage unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] In the description of the present invention, it should be understood that the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0040] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0041] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the preceding and following related objects.

[0042] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0043] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0044] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the accompanying drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary. In practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0045] The present invention provides a multi-source federated cross-domain and source-domain enhanced millimeter-wave action recognition method to solve the limitations of existing methods under the assumption of environmental consistency, mainly including solving the domain shift problem in cross-environment and cross-user scenarios of existing technologies, the problem that existing multi-party participation domain adaptation technologies cannot feedback to improve the model performance of participants, the data privacy protection problem caused by centralized learning technologies, and the high data annotation cost involved in supervised learning. The present invention combines domain adaptation technologies and proposes a framework of Federated Multi-Source Domain Adaptation (FMDA), which is applicable to complex scenarios in millimeter-wave HAR tasks. FMDA realizes knowledge transfer by evaluating the contributions of each source domain and performing weighted parameter aggregation, so as to perform unsupervised training on the target domain model without accessing any source domain data. This method improves the learning effects of all participants and enhances the overall performance by optimizing the generalization gap between the source domain model and the target domain model. In addition, the present invention also adopts a series of optimization algorithms to make the target domain model comparable to supervised learning methods in terms of performance, while improving the effectiveness of the source domain model to varying degrees. Verified by a large number of experiments, the present invention demonstrates its excellent application value and good development potential in the target domain and the source domain, and is applicable to diverse application scenarios; at the same time, the introduction of federated learning effectively avoids the direct transmission of raw data, meets the policy requirements of privacy protection, and is especially applicable to data such as medical health records, user biometric information, and financial data, which have data diversity and heterogeneity, and improves the generalization performance of the target domain model without destroying the data independence.

[0046] Embodiment 1 Please refer to Figure 1 , a multi-source federated cross-domain and source-domain enhanced millimeter-wave action recognition method of the present invention, includes the following steps: S1. Design a human action perception method. Using a millimeter-wave radar that emits Frequency-Modulated Continuous Wave (FMCW), by emitting multiple "chirp" signals with linearly increasing frequencies and receiving the echo signals reflected by the object, the perception of human actions is realized; After the transmitted signal and the received signal are combined, an intermediate frequency (IF) signal of the millimeter-wave radar is generated, and its expression is:

[0047] where is the amplitude after signal attenuation, is the starting frequency, is the bandwidth, is the chirp duration, is the time delay of the received signal relative to the transmitted signal. By analyzing the IF signal, the millimeter-wave radar can obtain the distance, speed, and angle information of the target.

[0048] In order to more efficiently sense human movements, the present invention performs a one-dimensional fast Fourier transform (1D-FFT) on the IF signal to obtain a distance heat map , specifically as follows:

[0049] where is the speed of light, is the frequency of the IF signal.

[0050] Subsequently, a static reflection removal algorithm is applied to reduce environmental noise interference.

[0051] On this basis, a two-dimensional fast Fourier transform (2D-FFT) is performed on the distance heat map to generate a distance-Doppler heat map to provide information on the moving speed of the target at different distances.

[0052] In addition, to characterize the time-series speed changes of different actions, a micro-Doppler spectrogram is generated. The specific method is to sum the power along the same distance and stack the Doppler spectrum bars belonging to the same frame, and finally obtain a micro-Doppler spectrogram showing the speed changes during a specific activity.

[0053] The present invention characterizes various human actions through micro-Doppler heatmaps. To highlight the motion patterns of different activities, when generating the micro-Doppler heatmap, values near zero frequency (the 60th frequency position of the velocity frequency) are suppressed by multiplying by a constant close to zero, so as to more clearly display the activity patterns. Through experiments where volunteers performed different actions in front of the radar, it was found that the micro-Doppler heatmaps of similar actions (such as pushing and clapping) completed by the same user in the same environment showed similar patterns, while different activities showed significantly different patterns. However, by comparing the heatmaps of the same action completed in different environments or by different users, it was found that the multipath effect and individual user differences would lead to significant pattern changes, namely "domain shift". The existence of domain shift brings uncertainty to activity recognition, but the experiments show that the micro-Doppler heatmap can still effectively characterize various activity patterns, providing rich basic data support for domain adaptation methods.

[0054] S2. Based on the voting of pseudo-labels, the micro-Doppler spectrogram obtained in step S1 is used as the input of the model for training. Considering that both the source model and the target model are composed of a feature extractor (·) and an activity classifier c(·), after the source model completes local training, the parameters of the source model and c (·) are sent to the target side (as shown in step 1 of Figure 2 [MOU1]). Since the target domain data lacks annotation information, a scheme for training the target model through pseudo-labels is proposed; An intuitive method is to use the mean of the class distributions output by the classifiers c(·) of all source models as pseudo-labels. However, due to the existence of domain shift, simply combining all class distribution information usually cannot obtain ideal performance.

[0055] Therefore, a voting-based pseudo-label generation method is adopted (as shown in step 2 of Figure 2 ). Its basic principle is that when a certain sample is predicted as a certain class more confidently by more source models, this class is more likely to be the true class of the sample.

[0056] Suppose there are trained source models . For each sample in the target domain, let represent the prediction result of the corresponding classifier.

[0057] Usually, the maximum predicted value is regarded as the label, that is, . To improve the credibility of the prediction, a confidence threshold is introduced, and only the results with a prediction confidence exceeding this threshold are retained, filtering out the models with lower confidence. The prediction results of the remaining models determine the final label through a voting mechanism, and the source models inconsistent with the voting results will be excluded.

[0058] Obtain a set of high-confidence source models that support the voting results through voting, and take the average prediction of these source models as the recognized knowledge, denoted as , where is the number of source models retained after voting.

[0059] Based on this knowledge, use the KL divergence loss to train the target model (including (·) and c(·)) (as shown in step 3 of Figure 2 ):

[0060] where is the reweighting factor of the loss, with a default value of 1.

[0061] Generally, if no source model passes the confidence threshold, it means that these models cannot provide valuable knowledge for the target domain. At this time, take the average value of the predictions of all source models as the general knowledge, and set to 0 to avoid the influence of these predictions on the training of the target model.

[0062] This step effectively filters out unreliable source model predictions through the voting mechanism, significantly improves the quality of the pseudo-labels, and thus enhances the learning effect of the target domain model.

[0063] S3. Weighted knowledge aggregation: Use the source models and pseudo-labels filtered in step S2 for subsequent knowledge aggregation; by evaluating the contribution degree of each source model in each iteration of training, weighted aggregation is achieved based on the contribution degree, and the source domain model that can stably provide high-quality, high-confidence prediction results identical to the true knowledge labels should be assigned a higher weight; Generally, fusing data knowledge from multiple source domains is of great significance for improving the performance and generalization ability of the target model. In federated learning, the common method is to aggregate multiple client models into a global model using fixed weights, and allocate specific weights according to the size of each client's local dataset. However, directly applying this aggregation strategy to the system of the present invention has low efficiency. For example, in a usage scenario with a complex reflection environment, millimeter-wave signals are significantly affected by the multipath effect. If the source domain model continuously contributes network parameters with fixed weights during the aggregation process, it may lead to a negative transfer effect on the target model.

[0064] To solve the above problems, based on the target model trained with KL divergence loss, this invention evaluates the contribution degree of each source model in each iteration training and realizes weighted aggregation based on this. As described in step S2, the pseudo-labels based on voting are crucial for enhancing the final prediction performance of the target domain. Subsequent experiments show that in the aggregation process, the source domain models that can stably provide high-quality, highly confident prediction results identical to the true knowledge labels should be assigned higher weights.

[0065] The specific steps of knowledge weight aggregation are as follows: S301. Quantification of source model contribution degree To quantify the contribution of the source model of this invention, it is proposed to evaluate its knowledge quality . First, define the total quality obtained by voting as:

[0066] Among them, represents the number of target domain samples, is the number of source models retained by voting, is the recognized knowledge label.

[0067] Generally, when the knowledge label is supported by more source models, it is more likely to reflect the true label of the target sample, proving that the model quality is higher.

[0068] Based on the total quality , the contribution of the source model can be expressed as:

[0069] Among them, is the total quality after removing the source model .

[0070] Finally, use the elimination method to estimate whether each source domain performs significantly. If a certain source model performs poorly on the target samples, it means that the influence of this source domain on the general knowledge is not significant, and it is considered that its contribution degree is low.

[0071] S302. Weight calculation Based on the contribution degree, the weight calculation methods of the target model and each source model are as follows: The weight of the target model, where represents the number of target samples with recognized labels:

[0072] The weight of the source model:

[0073] S303. Parameter Aggregation and Update In each training iteration of the model, the parameters of all source models and the target model are aggregated according to the above weights (as shown in step 4 of Figure 2 ):

[0074] Among them, and represent the feature extractor parameters of the target model and the source model respectively.

[0075] In this step, the classifier of the target model remains fixed, and only the feature extractor parameters are updated; Subsequently, the updated feature extractor parameters of the target model will be sent to all source models (as shown in step 5 of Figure 2 ) for the next iteration of training.

[0076] The global optimization strategy based on weighted aggregation proposed by the present invention can effectively improve the generalization ability of the target model and significantly reduce the possibility of negative transfer under the consideration of the influence of domain shift, thereby further enhancing the overall performance and applicability of the system.

[0077] S4. Generalization Optimization Method. After receiving the updated parameters sent in step S3, the source model will be fine-tuned.

[0078] The objective of the present invention is to develop a HAR model that can identify various activities in a new environment with domain shift. Therefore, in order to further enhance the generalization ability of the target model and provide positive feedback to all source domain models, a generalization optimization method is proposed. The strategy design of model training will mainly focus on extracting domain-independent activity features to optimize the generalization performance of the target model.

[0079] Research shows that flatness is an effective technique for improving generalization ability, and this technique can be quantified by the generalization gap between the source model and the target model. In the FMDA framework of the present invention, the generalization gap is defined as:

[0080] Among them, is the distance evaluation function, is the target classifier updated by KL divergence in step S3, is the pre-trained source classifier.

[0081] Based on the above concepts, the present invention ensures better flatness of the model across all domains by constraining the generalization gap between source models. The generalization optimization algorithm is implemented at the source side while ensuring that the classifier of the source model remains frozen during training. Specifically, training the feature extractor is divided into two stages: The first stage: optimizing the generalization error In the first stage, the generalization optimization algorithm is used to optimize the model parameters, and the KL divergence is used as the distance function to evaluate the generalization gap. Specifically, the global feature map broadcast by the target model after the end of the previous round of training is utilized, and the generalization gap loss is calculated:

[0082] The second stage: cross-entropy optimization In the second stage, the labeled dataset is used to further optimize the parameters and train through the cross-entropy loss function:

[0083] The above training process can gradually reduce the average generalization gap, making each closer to . At the same time, the generalization gap loss at the source side dynamically adjusts the weights of different source domains, thus preventing the target model from being biased towards certain specific parties. In addition, as the feature extractors of each source model gradually extract domain-agnostic active features, the performance of the source model will also be improved over time. The generalization optimization strategy proposed by the present invention not only significantly improves the generalization ability of the target model in new domains but also enables knowledge sharing among source models, enhancing the performance of source models, thereby further enhancing the overall performance and reliability of the system.

[0084] In the embodiment of the present invention, a commercially available AWR1443BOOST millimeter-wave radar (operating frequency of 77 ∼ 81 GHz) is used and paired with DCA1000EVM for data acquisition. This radar is a millimeter-wave sensor integrating multiple functions, capable of providing accurate distance, speed, and angle information, and is suitable for complex environmental perception tasks. The AWR1443BOOST radar has a sensing field of view (FOV) of 120° in the E plane and 30° in the H plane, transmits 30 frames of data per second, and can achieve fast real-time data updates. Its maximum sensing range is 10 meters, which is suitable for most indoor scenarios and can capture the motion information of objects in the surrounding environment, especially being prominent in enclosed spaces such as indoor corridors and laboratories.

[0085] Please refer to Figure 3, five different scenarios were considered in the experiment to comprehensively evaluate the performance of the millimeter-wave radar in different environments. The selected scenarios included a hall (Env1), a laboratory (Env2), a corridor (Env3), a lounge (Env4), and a corner (Env5). Each scenario had different structural features, spatial layouts, and possible interference factors, thus providing a diverse test environment. In each scenario, the millimeter-wave radar was precisely placed at a height of approximately 1 meter to ensure that its field of view could cover the experimental area. The volunteers in the experiment maintained a distance of approximately 2 meters from the radar, thus simulating common practical application scenarios such as smart home, security monitoring, and health monitoring.

[0086] In addition, during the experiment, the impact of environmental changes in different scenarios on the performance of the millimeter-wave radar was also considered. For example, reflective objects in the environment, the layout of the walls, and the movement patterns of people might cause certain interference to the radar's perception effect. Therefore, the test data for each scenario was analyzed in detail to ensure the accuracy and stability of the experimental results. By conducting experiments on these scenarios, we can not only verify the radar's perception ability but also evaluate its adaptability in complex environments, providing a reference for subsequent algorithm optimization and practical applications.

[0087] Those skilled in the art of the present technology can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "platforms" here.

[0088] Embodiment 2 The present invention provides a millimeter-wave action recognition system with multi-source federated cross-domain and source-domain enhancement, which can be used to implement the above-mentioned millimeter-wave action recognition method with multi-source federated cross-domain and source-domain enhancement. Specifically, the millimeter-wave action recognition system with multi-source federated cross-domain and source-domain enhancement includes a generation module, a generalization module, and an identification module.

[0089] Among them, the generation module emits multiple chirp signals with linearly increasing frequencies and receives the echo signals reflected by the object, and uses the obtained millimeter-wave radar to generate a micro-Doppler spectrogram; The generalization module adopts a voting-based pseudo-label method and a weighted knowledge aggregation mechanism to dynamically evaluate and fuse the knowledge of multiple source domains, and optimize the generalization ability of the server-side target model; The identification module analyzes the micro-Doppler heat map using the obtained target model to identify different human actions in the target domain, and realizes unsupervised human action recognition.

[0090] Embodiment 3 The present invention provides a terminal device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Graphics Processing Unit (GPU), Tensor Processing Unit (TPU), Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention can be used for the operation of the multi-source federated cross-domain and source-domain enhanced millimeter-wave action recognition method, including: Transmit multiple chirp signals with linearly increasing frequencies, and receive the echo signals reflected by the object, and generate a micro-Doppler spectrogram using the obtained millimeter-wave radar; adopt a voting-based pseudo-label method and a weighted knowledge aggregation mechanism to dynamically evaluate and fuse the knowledge of multiple source domains, and optimize the generalization ability of the server-side target model; use the obtained target model to analyze the micro-Doppler heat map to identify different human actions in the target domain and achieve unsupervised human action recognition.

[0091] Please refer to Figure 14 , the terminal device is a computer device. The computer device 60 in this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, it implements the multi-source federated cross-domain and source-domain enhanced millimeter-wave action recognition method in the embodiment. To avoid repetition, it will not be elaborated here one by one. Alternatively, when the computer program 63 is executed by the processor 61, it implements the functions of each model / unit in the multi-source federated cross-domain and source-domain enhanced millimeter-wave action recognition system in the embodiment. To avoid repetition, it will not be elaborated here one by one.

[0092] The computer device 60 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 14 These are merely examples of the computer device 60 and do not constitute a limitation on the computer device 60. It may include more or fewer components than those shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, a bus, etc.

[0093] The so-called processor 61 may be a central processing unit (CPU), or may also be other general-purpose processors, a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0094] The memory 62 may be an internal storage unit of the computer device 60, such as the hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk equipped on the computer device 60, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0095] Furthermore, the memory 62 may also include both an internal storage unit of the computer device 60 and an external storage device. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is to be output.

[0096] Please refer to Figure 15 , the terminal device is the electronic device 600, and the electronic device 600 is presented in the form of a general computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0097] Among them, the storage unit stores program code, which can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the method part of this specification. For example, the processing unit 610 can execute the steps as shown in Figure 1 .

[0098] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0099] The storage unit 620 may also include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0100] The bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0101] The electronic device 600 may also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem). Such communication may be carried out through the input / output interface 650. Moreover, the electronic device 600 may also communicate with one or more networks (such as a local area network, a wide area network, and / or a public network, such as the Internet) through the network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0102] Example 4 The present invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. It can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that more specific examples of the computer-readable storage medium here include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0103] The computer-readable storage medium also includes data signals propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, radio frequency, etc., or any suitable combination of the above.

[0104] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network or a wide area network, or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0105] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the millimeter-wave action recognition method related to multi-source federated cross-domain and source domain enhancement in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps: Transmit multiple chirp signals with linearly increasing frequencies, and receive the echo signals reflected by the object, and generate a micro-Doppler spectrogram using the obtained millimeter-wave radar; adopt a voting-based pseudo-label method and a weighted knowledge aggregation mechanism to dynamically evaluate and fuse the knowledge of multiple source domains, and optimize the generalization ability of the server-side target model; analyze the micro-Doppler heat map using the obtained target model to identify different human actions in the target domain, and realize unsupervised human action recognition.

[0106] The databases involved in the embodiments provided in the present application may include at least one of a relational database and a non-relational database. The non-relational database may include a blockchain-based distributed database, etc., and is not limited thereto. The processors involved in the embodiments provided in the present application may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., and are not limited thereto.

[0107] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the present invention described and shown in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0108] Taking gesture recognition as an example, please refer to Figure 4 , for 10 common human-computer interaction gestures (push, pull, clockwise rotation of the hand, up / down movement of the hand, left / right waving of the hand, hands crossed in a cross, clapping of both hands, waving), 10 undergraduate students were invited as volunteers to conduct experiments, and a total of 2,000 experimental data were collected. The hardware settings are as shown above.

[0109] Please refer to Figure 6, in the gesture recognition application scenario, the performance of FMDA was compared with six baseline methods. Among them, the FedAvg method provides a feasible way to transfer source knowledge and achieve unsupervised training of the target model, but its performance is poor. This method aggregates the parameters of different source models by averaging, without considering the domain differences and negative transfer problems. During the training process, the target model does not utilize the target samples, resulting in the model deviating from the target domain, thus leading to a low human activity recognition (HAR) accuracy; The FedAvg+ method uses vote-based pseudo-labels on the basis of FedAvg to train the target feature extractor and classifier. However, the parameter averaging method is still adopted in the update of the feature extractor without considering the individual contributions of each source domain. The obtained accuracy, precision, recall, and F1-score are 90.8%, 91.3%, 90.8%, and 91.1% respectively, all lower than the performance of FMDA; The SL method trains the target model in a supervised manner, achieving an accuracy of 93.5%, a precision of 93.8%, a recall of 93.5%, and an F1-score of 93.6%; The FMDA-VGG16 method uses VGG16 instead of ResNet50 as the feature extraction backbone network. The experimental results show that the performance of FMDA-VGG16 is lower than that of the system of the present invention. Since ResNet50 is deeper than VGG16, the deeper network architecture and its skip connections contribute to more effective feature extraction; The mTransSee method uses the source domain data to pre-train the original recognition network, and then uses the target domain data to fine-tune this network. Its accuracy, precision, recall, and F1-score are 60.3%, 60.2%, 60.3%, and 59.6% respectively; The EI method adopts an adversarial training method and uses the same source / target domain settings as the mTransSee method, obtaining an accuracy of 66.9%, a precision of 66.7%, a recall of 66.5%, and an F1-score of 66.4%. At the same time, this is also the best result obtained after testing dozens of hyperparameters for the loss function of EI.

[0110] Notably, in the above method comparison experiment results, although the accuracy of FMDA is 1.5% lower than that of the SL method, considering that FMDA is an unsupervised method, we believe this difference is acceptable. In addition, after FMDA training, the performance of all participating source models has been improved, and the accuracy has increased by 3% to 6.8%. This improvement further proves the effectiveness and value of FMDA.

[0111] Please refer to Figure 6 , ablation experiments were conducted on each module of FMDA, and the following key findings were obtained: The generalization optimization module can effectively improve the accuracy of the source model by about 3%. By exchanging the parameters of the global target model with each source model, this module enhances the overall generalization ability. The weighted knowledge aggregation module can effectively improve the overall performance of the model, and the results of FedAvg+ further emphasize the effectiveness of evaluating the contributions of each source model and aggregating parameters accordingly. Lack of pre-training of the source model has little impact on the accuracy of the target model, but the overall training time will be reduced by 9.3% compared to FMDA. During the federated training process, if the source model is not updated with cross-entropy loss in each round, the performance of the target model will decrease significantly. This is in line with expectations because FMDA relies on the knowledge of different source models to train the target model.

[0112] During the training process, when the target feature extractor is broadcast and updated in each source model, the performance of the initial target model is not ideal. Therefore, if the source model is under-trained, the target model will not be able to obtain enough valuable knowledge, resulting in moderate performance.

[0113] Please refer to Figure 7 , for the present invention, the influence of the number of data sources on the performance of the target HAR is explored. Since FMDA involves multi-source domain adaptation, this experiment was tested under the conditions of 2, 3, and 4 data sources. In the two-source setting, Environments 3 and 5 were selected as the data source combination, which are both in the middle in supervised learning among all five environments and can better verify the effectiveness of the method of the present invention. Please refer to Figure 5 . In the three-source setting, two combinations were tested: (Environment 2, Environment 3, Environment 5) and (Environment 1, Environment 3, Environment 5). In this experiment, only the number of source domains was changed, and the number of samples and participants remained the same as the default settings.

[0114] The experimental results show that in the application scenario of the present invention, the influence of the number of source domains on the performance of the target HAR shows an incremental but small change, and the accuracy in all cases remains above 90%.

[0115] In particular, when comparing the results of (Environment 1, Environment 3, Environment 5) and (Environment 1, Environment 2, Environment 3, Environment 5), it can be found that the former has higher accuracy and recall. This may be because Environment 2 is full of desks and chairs, which leads to the lowest HAR accuracy in this environment, as Figure 3 shown, and may introduce a certain degree of negative transfer.

[0116] However, the present invention partially alleviates the influence of this negative transfer by evaluating the contributions of each source domain during knowledge aggregation, thus achieving performance comparable to that in the multi-source setting case.

[0117] Please refer to Figure 8, then the sample quantity ratio in each source domain was changed to study its impact on performance. Specifically, three different data ratio schemes were set up: (1) 1:1:0.5:0.5; (2) 1:0.75:0.5:0.25; (3) 1:1:1:1.

[0118] Among them, the source domain order was Environment 1, Environment 2, Environment 3, and Environment 5 in sequence.

[0119] The experimental results show that although the data ratio does have a certain impact on performance, overall the impact is relatively small. Generally, more balanced and larger numbers of samples will bring better performance. In all three schemes, the average accuracy rate exceeds 90%.

[0120] Please refer to Figure 9 , the domain adaptation performance was evaluated under the condition that each source domain has different and unevenly distributed users. The specific settings were as follows: Source domain 1 included Users #1 to #6, Source domain 2 included Users #3 to #8, Source domain 3 included Users #5 to #10, Source domain 4 included Users #2, #4, #6, #8, and #10.

[0121] The target domain included Users #1 to #10. Figure 9 The results show that the HAR accuracy rate of the target domain is only 0.2% lower than that of FMDA, indicating that the method of the present invention has strong robustness to the number of users.

[0122] In addition, the performance of each user was plotted in Figure 10 , and the performance of each user in the source domain and the target domain was further compared in Figure 11 . The results show that users in source domain 1 generally show better domain adaptation performance. This observation indicates that knowledge in relatively "clean" domains is more valuable for improving the adaptive performance.

[0123] Please refer to Figure 12, the performance of FMDA in the target domain when new users perform different activities was further evaluated. In the overall performance evaluation, there were differences in the HAR accuracy of different users during domain adaptation. Therefore, two experimental settings were designed in this experiment to compare the performance of users in the following two scenarios: one is that the users are in the source domain, and the other is that the users are completely new users. In the first experimental setting, the source domain included users #1 to #4 and #7 to #10, and the HAR accuracy of new users #5 and #6 was tested. In the second experimental setting, the source domain included users #3 to #10, and the new users were users #1 and #2. By observing the experimental results, it can be found that the performance of user #2 was always good, regardless of whether he was in the source domain or not. This is attributed to the fact that the movement pattern of user #2 is relatively general and similar to that of most users, enabling the system to effectively recognize his activities even in the absence of targeted training samples. In contrast, the movement patterns of other users are relatively unique, and their performance will decline to a certain extent when the system has not seen these patterns.

[0124] Please refer to Figure 13 , to further study the performance of FMDA at different distances and directions between users and the radar (default settings are 2 meters and 0°), an additional volunteer was invited in each source domain (Environment 1, Environment 2, Environment 3, and Environment 5) to perform ten predefined gestures. In addition, these four volunteers also repeated these gestures in the target domain (Environment 4). Specifically, each volunteer completed ten gestures in front of the radar at distances of 1 meter, 2 meters, and 3 meters respectively, and at directions of 15° (left and right) to the left and right of the radar at a distance of 2 meters. After integrating the newly collected data into FMDA for retraining, the system could achieve an accuracy, precision, recall, and F1 score of approximately 80% at different user-radar distances, and these metrics could also remain at approximately 75% at different directions. It is worth noting that without the design of FMDA, retraining each source domain separately and testing in the target domain, the average performance could only reach approximately 55% and 35% at different distances and directions. In addition, if all the data of the new volunteers were used for training in a supervised learning manner in Environment 4 and tested in Environments 1, 2, 3, and 5 respectively, the average performance could only be improved to approximately 65% and 45% at different distances and directions.

[0125] In summary, a millimeter-wave action recognition method and system with multi-source federated cross-domain and source domain enhancement according to the present invention can break through the limitations of existing human activity recognition methods in multi-domain data privacy protection and unsupervised learning. The present invention does not require access to source domain data, and realizes efficient knowledge transfer through weighted parameter aggregation and minimization of the generalization gap, while improving the performance of the target model and the effect of the source domain model. At the same time, the FMDA of the present invention is compatible with multi-party distributed data privacy protection policies, and is applicable to practical application scenarios such as personal health monitoring, biometric analysis, and smart home interaction. It has the characteristics of high efficiency, universality, privacy friendliness, and low cost, and demonstrates good application value, research prospects, and development potential.

[0126] The above content is only for explaining the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. A millimeter wave action recognition method with multi-source federation, cross-domain and source domain enhancement, characterized in that: The following steps are involved: Transmit multiple chirp signals with increasing linear frequency, receive echo signals reflected by objects, and use the obtained millimeter-wave radar to generate a micro-Doppler spectrum diagram; A voting-based pseudo-labeling method and a weighted knowledge aggregation mechanism are used to dynamically evaluate and fuse knowledge from multiple source domains to optimize the generalization ability of the server-side target model. The obtained target model is used to analyze the micro-Doppler heat map and identify different human motions in the target domain to achieve unsupervised human motion recognition.

2. The multi-source federated cross-domain and source domain enhanced millimeter wave motion recognition method according to claim 1 is characterized in that: Generate a micro-Doppler spectrum diagram as follows: The chirp signal and the echo signal are combined to generate a millimeter-wave radar intermediate frequency signal, and the millimeter-wave radar intermediate frequency signal is subjected to a one-dimensional fast Fourier transform to obtain the distance information. The distance information is then subjected to a fast Fourier transform to generate a range-Doppler heat map, and then a micro-Doppler spectrum diagram showing the speed changes during a specific activity is obtained.

3. The multi-source federated cross-domain and source domain enhanced millimeter wave motion recognition method according to claim 2 is characterized in that: Millimeter wave radar intermediate frequency signal for: in, is the amplitude of the signal after attenuation, is the starting frequency, is the bandwidth, is the chirp duration, is the time delay of the received signal relative to the transmitted signal.

4. The multi-source federated cross-domain and source domain enhanced millimeter wave motion recognition method according to claim 1 is characterized in that: A voting-based pseudo-labeling method is adopted, as follows: Features Trained source models , for each sample in the target model ,use Represents the prediction result of the corresponding classifier, regards the maximum prediction value as the label, and introduces a confidence threshold , only keep the predictions with confidence exceeding the confidence threshold The prediction results of the remaining models are used to determine the final labels through a voting mechanism. A set of high-confidence source models that support the voting results are obtained through voting. The predicted average of the high-confidence source models is used as the acknowledged knowledge, and the target model is trained using the KL divergence loss.

5. The multi-source federated cross-domain and source domain enhanced millimeter wave action recognition method according to claim 4 is characterized in that: KL divergence loss as follows: in, is the reweighting factor of the loss, is the divergence calculation function, is the target model prediction result of the jth sample, It is the average result predicted by the high confidence source model for the jth sample.

6. The multi-source federated cross-domain and source domain enhanced millimeter wave motion recognition method according to claim 1 is characterized in that: The specific weighted knowledge aggregation mechanism is: Define the total mass obtained by voting , determine the source model Contribution , based on contribution The weights of the target model and each source model are calculated separately. In each training iteration, the parameters of all source models and target models are aggregated according to the weights, and the updated target model feature extractor parameters are Will be broadcast to all source models for the next iteration of training.

7. The multi-source federated cross-domain and source domain enhanced millimeter wave motion recognition method according to claim 6 is characterized in that: Updated target model feature extractor parameters for: in, and denote the feature extractor parameters of the target model and the source model respectively, Indicates the number of source models involved in training, and Represent the weighting factors of the target model and the i-th source model respectively.

8. The multi-source federated cross-domain and source domain enhanced millimeter wave motion recognition method according to claim 1 is characterized in that: The specific steps to optimize the generalization capability of the server-side target model are: In the first stage, the generalization optimization algorithm is used to optimize the model parameters, using the target model broadcast after training. Generate global feature maps and calculate generalization gap loss ; In the second stage, using the labeled dataset Optimize parameters and train through cross entropy loss function.

9. The multi-source federated cross-domain and source domain enhanced millimeter wave motion recognition method according to claim 8, characterized in that: Generalization Gap Loss and respectively for: in, is the divergence calculation function, is the classifier of the target model, is the feature extractor of the target model, is the source model dataset, are the parameters of the target model feature extractor, is the pre-training result corresponding to the source model, is the feature extractor of the source model, are the parameters of the source model feature extractor, is the source model cross entropy loss, is the number of samples contained in the i-th source model dataset, is the class label of the jth sample in the i-th source model dataset, is the predicted class of the jth sample in the i-th source model dataset.

10. A multi-source federated cross-domain and source domain enhanced millimeter wave action recognition system, characterized in that: include: A generation module transmits multiple chirp signals with increasing linear frequencies, receives echo signals reflected by objects, and generates a micro-Doppler spectrum using the obtained millimeter-wave radar; The generalization module uses a voting-based pseudo-labeling method and a weighted knowledge aggregation mechanism to dynamically evaluate and fuse knowledge from multiple source domains to optimize the generalization ability of the server-side target model; The recognition module uses the obtained target model to analyze the micro-Doppler heat map and identify different human movements in the target domain to achieve unsupervised human movement recognition.

Citation Information

Cited By

  • AI-driven multi-wireless-signal candid shooting detection device

    CN121386018A

  • Radar image target identification method based on multi-source cross-angular domain adaptation

    CN121544873A

  • A radar image target recognition method based on multi-source cross-angle domain adaptation

    CN121544873B

  • Domain adaptation method and system based on active learning

    CN121600351A

  • Portable plant leaf nitrogen content nondestructive testing method, device and medium

    CN122042606A