Cross-domain object presence awareness model training method, object awareness method, and device

CN122047398BActive Publication Date: 2026-09-25UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610519625.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-20
Publication Date
2026-09-25
Estimated Expiration
2046-04-20

AI Technical Summary

Technical Problem

[0003]其中,基于摄像头和红外传感器的感知方案虽然精度较高,但存在安全隐患

Benefits of technology

[0012]根据本发明的实施例,通过多维度相位误差修正从信道状态样本信息中提取鲁棒的目标相位样本数据,并利用初始特征编码器提取富含多径散射信息的信道散射样本特征。进而,一方面基于不同域的信道散射样本特征的分布特性,最小化域间特征分布差异,实现对未知域的环境特征学习。另一方面,基于同类或异类样本的类别标签,引入对比损失,在特征空间中拉近同类样本特征距离、推远异类样本特征距离,增强特征对目标状态的判别性。并且将域对齐损失与对比损失联合优化,训练得到兼具环境鲁棒性和类别判别性的目标特征编码器,从而在零训练样本的未知域部署时仍能保持较高精度的对象存在感知。因此,至少部分地解决了相关技术中存在检测安全隐患以及检查精度较低的技术问题,提升了对象存在检测在实际应用中的适用性与灵活性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047398B_ABST
    Figure CN122047398B_ABST
Patent Text Reader

Abstract

The application provides a cross-domain object presence perception model training method, an object perception method and equipment, and can be applied to the field of artificial intelligence technology. The method comprises the following steps: performing multi-dimensional phase error correction on channel state sample information of multiple different domains respectively to obtain multiple target phase sample data; performing feature extraction on the target phase sample data by using an initial feature encoder of an initial presence perception model to obtain multiple channel scattering sample features; determining distribution differences between the channel scattering sample features of the multiple different domains based on the distribution characteristics of the multiple channel scattering sample features; determining a contrast loss based on the similarity between the channel scattering sample features and the category labels of the channel scattering sample features; and training the initial presence perception model to obtain a presence perception model by taking the minimization of the distribution differences and the contrast loss as a joint optimization target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a cross-domain object existence perception model training method, object perception method, and device. Background Technology

[0002] Object presence detection is one of the core foundational technologies for achieving spatial intelligence, playing an indispensable role in device linkage in smart homes and environmental status perception in the Internet of Things. Currently, this technology mainly relies on cameras, infrared sensors, and Wireless Fidelity (WiFi) signals.

[0003] While camera- and infrared sensor-based sensing solutions offer high accuracy, they also pose security risks. Conversely, WiFi-based sensing technologies, though providing improved security, suffer from a sharp decline in generalization ability when migrating between different scenarios, leading to a significant drop in detection accuracy. Summary of the Invention

[0004] In view of the above problems, the present invention provides a cross-domain object existence perception model training method, object perception method and device.

[0005] According to one aspect of the present invention, a method for training a cross-domain object presence-aware model is provided, comprising: performing multi-dimensional phase error correction on channel state sample information from multiple different domains to obtain multiple target phase sample data; using an initial feature encoder of an initial presence-aware model to extract features from each target phase sample data to obtain multiple channel scattering sample features; determining the distribution differences between each pair of channel scattering sample features from multiple different domains based on the distribution characteristics of each channel scattering sample feature; determining a contrast loss based on the similarity between each channel scattering sample feature and the category label of each channel scattering sample feature, wherein the contrast loss is used to guide the initial presence-aware model to increase the aggregation degree between channel scattering sample features with the same category label and the separation degree between channel scattering sample features with different category labels; and training the initial presence-aware model with minimizing the distribution difference and the contrast loss as a joint optimization objective to obtain a presence-aware model.

[0006] Another aspect of the present invention provides a cross-domain object presence-aware model training apparatus, comprising: a first correction module, configured to perform multi-dimensional phase error correction on channel state sample information from multiple different domains respectively, to obtain multiple target phase sample data; a feature extraction module, configured to extract features from each target phase sample data using an initial feature encoder of an initial presence-aware model, to obtain multiple channel scattering sample features; a first determination module, configured to determine the distribution differences between each pair of channel scattering sample features from multiple different domains based on the distribution characteristics of each channel scattering sample feature; a second determination module, configured to determine a contrast loss based on the similarity between each channel scattering sample feature and the category label of each channel scattering sample feature, wherein the contrast loss is used to guide the initial presence-aware model to increase the aggregation degree between channel scattering sample features with the same category label and the separation degree between channel scattering sample features with different category labels; and a training module, configured to train the initial presence-aware model with minimizing the distribution difference and the contrast loss as a joint optimization objective, to obtain a presence-aware model.

[0007] Another aspect of the present invention provides an object perception method, comprising: in response to receiving an existence perception task for a target domain, performing multi-dimensional phase error correction on channel state information collected by a wireless communication device in the target domain to obtain target phase data; using an existence perception model, performing object existence perception based on the target phase data to obtain a probability value corresponding to the object's existence state and an uncertainty of the probability value; and based on a fusion method corresponding to the uncertainty, fusing an estimated perception result determined by historical perception results and historical process noise with an initial perception result to obtain a target perception result indicating whether an object exists in the target domain.

[0008] Another aspect of the present invention provides an object sensing device, comprising: a second correction module, configured to, in response to receiving an existence sensing task for a target domain, perform multi-dimensional phase error correction on channel state information acquired by a wireless communication device in the target domain to obtain target phase data; a sensing module, configured to, using an existence sensing model, perform object existence sensing based on the target phase data to obtain an initial sensing result, the initial sensing result including a probability value corresponding to the object's existence state and an uncertainty of the probability value; and a smoothing module, configured to, based on a fusion method corresponding to the uncertainty, fuse an estimated sensing result determined by historical sensing results and historical process noise with the initial sensing result to obtain a target sensing result indicating whether an object exists in the target domain.

[0009] Another aspect of the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0010] Another aspect of the present invention provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0011] Another aspect of the present invention provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0012] According to embodiments of the present invention, robust target phase sample data is extracted from channel state sample information through multi-dimensional phase error correction, and channel scattering sample features rich in multipath scattering information are extracted using an initial feature encoder. Furthermore, on the one hand, based on the distribution characteristics of channel scattering sample features in different domains, the difference in feature distribution between domains is minimized, enabling the learning of environmental features in unknown domains. On the other hand, based on the category labels of similar or dissimilar samples, a contrastive loss is introduced to narrow the feature distance of similar samples and widen the feature distance of dissimilar samples in the feature space, enhancing the discriminative power of features for target states. Moreover, the domain alignment loss and contrastive loss are jointly optimized to train a target feature encoder that combines environmental robustness and category discriminative power, thus maintaining high accuracy in object presence detection even when deployed in unknown domains with zero training samples. Therefore, this at least partially solves the technical problems of detection security risks and low inspection accuracy in related technologies, improving the applicability and flexibility of object presence detection in practical applications. Attached Figure Description

[0013] The above-mentioned contents, as well as other objects, features and advantages of the present invention, will become clearer from the following description of embodiments of the present invention with reference to the accompanying drawings.

[0014] Figure 1 The diagram illustrates an application scenario of the cross-domain object existence perception model training method, object perception method, and apparatus according to embodiments of the present invention.

[0015] Figure 2 A flowchart of a cross-domain object existence awareness model training method according to an embodiment of the present invention is shown.

[0016] Figure 3 Waveforms of amplitude and phase under different states of presence according to embodiments of the present invention are shown.

[0017] Figure 4 A schematic diagram of phase data for each phase error correction stage according to an embodiment of the present invention is shown.

[0018] Figure 5 A schematic diagram of an initial feature encoder according to an embodiment of the present invention is shown.

[0019] Figure 6 A performance comparison chart of a feature encoder according to an embodiment of the present invention and a feature encoder in related technologies is shown.

[0020] Figure 7 A model architecture diagram of a cross-domain object existence awareness model training method according to an embodiment of the present invention is shown.

[0021] Figure 8 A schematic diagram of collaborative feature optimization based on domain alignment and supervised contrastive learning according to an embodiment of the present invention is shown.

[0022] Figure 9(a) shows the performance comparison results of different domain generalization methods according to embodiments of the present invention in cross-environment detection tasks.

[0023] Figure 9(b) shows a schematic diagram comparing the feature distribution of the training set with and without the introduction of domain alignment and supervised contrastive learning according to an embodiment of the present invention.

[0024] Figure 9(c) shows a schematic diagram comparing the feature distribution of the test set with and without the introduction of domain alignment and supervised contrastive learning according to an embodiment of the present invention.

[0025] Figure 10 A flowchart of an object-aware method according to an embodiment of the present invention is shown.

[0026] Figure 11 A schematic diagram is shown before and after smoothing the initial perception result using the estimated perception result according to an embodiment of the present invention.

[0027] Figure 12 A structural block diagram of a cross-domain object existence awareness model training device according to an embodiment of the present invention is shown.

[0028] Figure 13 A structural block diagram of an object sensing device according to an embodiment of the present invention is shown.

[0029] Figure 14 A block diagram of an electronic device suitable for implementing a cross-domain object presence awareness model training method and an object awareness method according to an embodiment of the present invention is shown. Detailed Implementation

[0030] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0031] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0032] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0033] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0034] In the technical solution of this invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, invention, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0035] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this invention offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.

[0036] The research revealed that object presence detection plays an indispensable role in smart homes, smart security, and IoT systems. Currently, the mainstream implementation of object presence detection relies primarily on cameras or infrared sensors. While these devices can achieve high detection accuracy, they pose security risks and their application in private scenarios such as homes and elderly care facilities is strictly limited, failing to meet users' core needs for privacy protection.

[0037] However, object presence sensing technologies based on WiFi signals suffer from a severe "domain dependency" problem. Because WiFi signals are highly sensitive to the environment and exhibit distribution shifts in different scenarios, their generalization ability deteriorates sharply when migrating across scenarios, leading to a significant drop in detection accuracy.

[0038] In view of this, embodiments of the present invention provide a method for training a cross-domain object presence awareness model, including: addressing the offset problem of WiFi signal presence under different environments, since Channel State Information (CSI) can characterize the multipath characteristics of wireless signals under specific propagation environments, the preprocessed CSI is used as input, robust time-frequency features are extracted through a time-frequency hybrid attention mechanism, cross-domain feature distribution is calibrated by domain alignment, category discrimination is enhanced by contrastive learning, and the initial presence awareness model is trained through multi-loss joint optimization, thereby achieving more accurate detection of cross-domain object presence.

[0039] Figure 1 The diagram illustrates an application scenario of the cross-domain object existence perception model training method, object perception method, and apparatus according to embodiments of the present invention.

[0040] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0041] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0042] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0043] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0044] It should be noted that the cross-domain object presence perception model training method and object perception method provided in the embodiments of the present invention can generally be executed by server 105. Correspondingly, the cross-domain object presence perception model training device and object perception device provided in the embodiments of the present invention can generally be located in server 105. The cross-domain object presence perception model training method and object perception method provided in the embodiments of the present invention can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the cross-domain object presence perception model training device and object perception device provided in the embodiments of the present invention can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0045] It should be understood that Figure 1 The number of first terminal devices, second terminal devices, third terminal devices, networks, and servers shown in the diagram is merely illustrative. Depending on implementation needs, any number of first terminal devices, second terminal devices, third terminal devices, networks, and servers can be included.

[0046] The following will be based on Figure 1 The described scene, through Figures 2-11 The cross-domain object existence awareness model training method of the embodiments of the invention is described in detail.

[0047] Figure 2 A flowchart of a cross-domain object existence awareness model training method according to an embodiment of the present invention is shown.

[0048] like Figure 2 As shown, the method includes operations S210 to S250.

[0049] In operation S210, multi-dimensional phase error correction is performed on channel state sample information from multiple different domains to obtain multiple target phase sample data.

[0050] In operation S220, the initial feature encoder of the initial existence-aware model is used to extract features from the phase sample data of each target to obtain multiple channel scattering sample features.

[0051] In operation S230, based on the distribution characteristics of multiple channel scattering sample features, the distribution differences between each pair of channel scattering sample features in multiple different domains are determined.

[0052] In operation S240, a contrast loss is determined based on the similarity between the features of each channel scattering sample and the category label of each channel scattering sample feature. The contrast loss is used to guide the initial existence-aware model to increase the degree of aggregation between channel scattering sample features with the same category label and the degree of separation between channel scattering sample features with different category labels.

[0053] In operation S250, the initial existence-aware model is trained with the joint optimization objective of minimizing distribution difference and contrast loss to obtain the existence-aware model.

[0054] There are no restrictions on the multiple channel state sample information used during training; it can include channel state sample information from multiple different domains and channel state sample information from multiple time windows within the same domain.

[0055] Channel state sample information can be obtained in the following way: using a wireless communication device or system with CSI acquisition capability, the channel state information formed during the propagation of WiFi signals is collected in an indoor environment. The receiving end acquires the channel frequency response corresponding to multiple subcarriers at each time step, thereby forming channel state sample information.

[0056] Channel state sample information: the channel frequency response of the w-th subcarrier corresponding to the m-th receiving antenna at time t. It can be modeled as shown in the following formula (1).

[0057] (1)

[0058] In the formula, This indicates the number of multipath components in the signal propagation path. and They represent Time of the first The amplitude attenuation coefficient and phase shift of each propagation path, The CSI matrix at time t can be represented as ,in Indicates the number of receiving antennas. represents the total number of subcarriers, and j represents the imaginary unit.

[0059] It can collect continuous data. CSI matrix at each time step or time point (i.e.) H[1] is the CSI matrix at the first time step, serving as channel state sample information. and Similarly.

[0060] A domain can refer to different detection environments, such as different rooms or different locations.

[0061] The object is not limited and can be people, animals, etc. within the domain.

[0062] By performing multi-dimensional phase error correction on channel state sample information, systematic interference can be suppressed, and the presence-related signal characteristics of the target object can be enhanced, providing stable and reliable input data for subsequent presence-aware models. The method of multi-dimensional phase error correction is not limited and can include linear error correction, common-mode phase interference correction, and environmental noise correction, among others.

[0063] The specific method for the initial feature encoder to extract features from the target phase sample data is not limited; it can extract time-domain features, frequency-domain features, or time-frequency hybrid features.

[0064] The network architecture for the initial feature encoder is not limited and can be a convolutional neural network, a Transformer feature extraction network, etc. The Transformer is a deep learning architecture based on the self-attention mechanism.

[0065] By performing feature mapping and nonlinear transformation through multi-layer convolution, pooling, or attention mechanisms in the initial feature encoder, a high-dimensional vector with discriminative power can be extracted from each target phase sample data, thereby obtaining multiple channel scattering sample features corresponding to the number of samples.

[0066] In calculating the distributional differences between channel scattering sample features in different domains, we can statistically analyze the distribution characteristics of each channel scattering sample feature, such as mean, variance, probability density distribution, and covariance. Then, we can select appropriate distributional difference measurement methods, such as maximum mean difference or Frobenius norm, to calculate the pairwise distributional difference values ​​between channel scattering sample features in different domains, thereby quantifying the degree of deviation in the distribution of features in different domains.

[0067] In determining the contrast loss, the similarity value between the features of any two channel scattering samples can be calculated, and the contrast loss can be constructed by combining the category labels corresponding to the features of each channel scattering sample, such as whether the object exists or does not exist. The contrast loss function can also be called the supervised contrast loss.

[0068] During model training, joint optimization training can be performed based on the aforementioned loss values. For example, the distribution difference value and the contrastive loss value can be weighted and fused to construct a joint optimized loss value. Then, by minimizing the joint optimized loss value as the training objective, the existence-aware model can be trained.

[0069] According to embodiments of the present invention, robust target phase sample data is extracted from channel state sample information through multi-dimensional phase error correction, and channel scattering sample features rich in multipath scattering information are extracted using an initial feature encoder. Furthermore, on the one hand, based on the distribution characteristics of channel scattering sample features in different domains, the difference in feature distribution between domains is minimized, enabling the learning of environmental features in unknown domains. On the other hand, based on the category labels of similar or dissimilar samples, a contrastive loss is introduced to narrow the feature distance of similar samples and widen the feature distance of dissimilar samples in the feature space, enhancing the discriminative power of features for target states. Moreover, the domain alignment loss and contrastive loss are jointly optimized to train a target feature encoder that combines environmental robustness and category discriminative power, thus maintaining high accuracy in object presence detection even when deployed in unknown domains with zero training samples. Therefore, this at least partially solves the technical problems of detection security vulnerabilities and low inspection accuracy in related technologies, improving the applicability and flexibility of object presence detection in practical applications.

[0070] According to an embodiment of the present invention, the target phase sample data is a phase difference sequence; the channel state sample information includes the original phase data; the original phase data includes the original phase at multiple time points; and the channel state sample information of multiple different domains is subjected to multi-dimensional phase error correction to obtain multiple target phase sample data, which may include the following operations.

[0071] By performing a linear transformation, the original phases at multiple time points are corrected for linear errors to obtain intermediate phase data. For the same time point, the difference between the original phases of adjacent receiving antennas is calculated to obtain the initial phase difference at that time point. The initial phase differences at each time point are arranged in chronological order to obtain the initial phase difference sequence. Environmental noise is filtered out from the initial phase difference sequence using a smoothing filter to obtain the phase difference sequence.

[0072] Figure 3 Waveforms of amplitude and phase under different states of presence according to embodiments of the present invention are shown.

[0073] like Figure 3As shown, in the case of a person, 310 shows the amplitude timing waveform 311, the phase difference timing waveform 312, and the phase difference frequency domain waveform 313 obtained by Fourier transform of the phase difference sequence in the channel state sample information under the unmanned state. 320 shows the amplitude timing waveform 321, the phase difference timing waveform 322, and the phase difference frequency domain waveform 323 under the manned state.

[0074] It can be seen that in the unmanned state, the overall changes of various signals are relatively stable, while in the manned state, compared with the amplitude characteristics, the phase difference time series exhibits more obvious dynamic fluctuations and abrupt changes. Further transformation of the phase difference data to the frequency domain reveals a more significant energy difference in the low-frequency region in the manned state, which is beneficial for characterizing the multipath structure changes caused by the presence of an object. Therefore, this invention uses phase difference data as the core input in the subsequent feature extraction stage and jointly utilizes its time-domain and frequency-domain information to improve robustness.

[0075] Furthermore, based on the aforementioned characteristics, it can be seen that among the two types of data included in the channel state sample information—raw amplitude data and raw phase data—raw amplitude data is more susceptible to factors such as transmit power fluctuations, automatic gain control, and environmental noise, exhibiting poor stability under different environmental conditions. In contrast, raw phase data is more sensitive to multipath variations caused by the presence of an object and can reflect more fine-grained channel disturbance characteristics.

[0076] However, the raw phase data is inevitably affected by hardware non-ideal factors during acquisition. To address linear errors such as clock synchronization deviation and carrier frequency offset in the raw phase data, a linear transformation can be used to eliminate phase offset across the entire frequency band. Then, the phase difference between adjacent receiving antennas can be used to further cancel the synchronization error. Next, a smoothing filter can be used to remove residual environmental burst noise, resulting in a phase difference sequence for subsequent modules.

[0077] There are no restrictions on the smoothing filter; it can be a Savitzky-Golay (SG) filter, a moving average filter, etc.

[0078] Figure 4 A schematic diagram of phase data for each phase error correction stage according to an embodiment of the present invention is shown.

[0079] like Figure 4As shown, the original phase data is first processed by calculating the phase difference to remove phase offset. During the correction process, global phase errors introduced by hardware non-ideals such as carrier frequency offset and clock asynchrony are eliminated. Next, the phase difference between adjacent receiving antennas is calculated to further remove phase offset and cancel common-mode phase interference. Finally, the phase difference sequence is subjected to SG filtering to obtain SG-filtered phase difference data, thereby suppressing sudden environmental noise and obtaining a stable CSI phase difference sequence. Observing the waveforms at each stage shows that after the above processing, systematic interference is effectively weakened, and phase changes caused by the presence of the target object are highlighted.

[0080] According to embodiments of the present invention, a phase difference sequence with higher signal-to-noise ratio and higher stability can be obtained by correcting hardware linearity errors through linear transformation, eliminating common-mode interference by utilizing the phase difference between adjacent antennas, and then suppressing environmental noise through smoothing filtering. This sequence can amplify the phase perturbation caused by the presence of objects, reduce the impact of environmental differences on the signal, and provide a high-quality data foundation for robust sensing of subsequent models in cross-domain scenarios.

[0081] According to an embodiment of the present invention, the target phase sample data is a phase difference sequence; the phase difference sequence includes: phase differences at multiple time points; the initial feature encoder includes: a time domain feature extraction module, a frequency domain feature extraction module, and a feature fusion module; by using the initial feature encoder, feature extraction is performed on each target phase sample data to obtain multiple channel scattering sample features, which may include the following operations.

[0082] For each target phase sample data, the temporal feature extraction module extracts temporal features from the phase differences at multiple time points to obtain temporal features. This temporal feature extraction module employs an attention network architecture, dynamically allocating attention weights to focus on the phase differences at time points relevant to the object's existence. The frequency domain feature extraction module performs Fourier transforms on the phase differences at multiple time points to obtain frequency domain representations of the phase differences, and then extracts frequency domain features from these representations to obtain frequency domain features. This frequency domain feature extraction module uses an attention network architecture isomorphic to the temporal feature extraction module. Finally, the feature fusion module fuses the temporal and frequency domain features to obtain the channel scattering sample features.

[0083] Figure 5 A schematic diagram of an initial feature encoder according to an embodiment of the present invention is shown.

[0084] like Figure 5 As shown, the initial feature encoder uses the phase difference sequence as a unified input and employs parallel temporal branch (temporal feature extraction module, i.e., temporal attention module in the figure) and frequency domain branch (frequency domain feature extraction module, i.e. frequency domain attention module in the figure) for feature extraction.

[0085] Both branches are based on Transformer Encoder Blocks for feature extraction. Each Transformer Encoder Block contains a multi-head self-attention mechanism, a feed-forward network, residual connections (Add), and layer normalization.

[0086] In the temporal branch, the phase difference sequence is first input into the embedding layer, which maps the phase difference at each time point into a high-dimensional feature vector, resulting in a temporal feature embedding representation. This is then fed into stacked Transformer encoding units, where a multi-head self-attention mechanism dynamically allocates attention weights, focusing on key time segments relevant to the object's existence, suppressing static redundancy, and outputting temporal dynamic features, such as a temporal feature map. These temporal dynamic features can then be compressed into a fixed-length temporal feature vector using average pooling.

[0087] In the frequency domain branch, the phase difference sequence is input into the embedding layer and transformed to the frequency domain via Fourier Transform to extract stable spectral components. These components are then fed into a Transformer coding unit whose structure is symmetrical to that of the time domain branch. A multi-head self-attention mechanism models the correlation between different frequency components, followed by an inverse Fourier Transform to output frequency domain features, such as a frequency domain feature map. Similarly, the frequency domain features are compressed into a fixed-length frequency domain feature vector using average pooling.

[0088] For example, the input phase difference sequence can be converted into a frequency domain representation using a Fast Fourier Transform (FFT) to highlight the effective frequency components and suppress high-frequency noise.

[0089] The time-domain feature vector and the frequency-domain feature vector are fused element-wise by adding them to the time-domain feature vector to obtain the output, which is a time-frequency hybrid feature representation, for subsequent classification and cross-domain optimization.

[0090] After retaining the low-frequency components and filtering out the high-frequency interference, the input can be linearly projected into the multi-head self-attention module. The feature association weights are calculated through the "query (Q)-key (K)-value (V)" mapping to capture the dependency relationship between different frequency components. The scaling dot product attention calculation formula can be shown in the following formula (2).

[0091] (2)

[0092] In the formula, Indicates attention, , , These represent the query matrix, key matrix, and value matrix, respectively. This indicates that the query matrix, key matrix, and value matrix all belong to d k The real space of dimension 1 and d k The dimensions of the query matrix, key matrix, and value matrix are represented by linear projection from the input matrix. The input features are projected onto multiple different query, key, and value spaces, and the outputs of multiple attention heads are computed in parallel. Then, the outputs of all attention heads are concatenated, and a linear transformation is performed through the output projection matrix to obtain the frequency domain features.

[0093] Figure 6 A performance comparison chart of a feature encoder according to an embodiment of the present invention and a feature encoder in related technologies is shown.

[0094] like Figure 6 As shown, the performance comparison results of different feature extraction networks in cross-environment detection tasks are presented under the same training conditions. The comparison models include Residual Network 50 (ResNet50), Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), a combination of Convolutional Neural Network (CNN) and GRU, and the feature encoder proposed in this invention. All models were trained on the same training data and tested on target environment data that were not used in the training. The results show that the feature encoder of this invention exhibits more stable cross-environment detection performance in both accuracy and F1 score, indicating that its joint modeling of the temporal dynamic features and frequency-domain stable features of CSI effectively improves cross-environment generalization ability.

[0095] According to an embodiment of the present invention, by utilizing the difference between the dynamic mutations in the time domain and the stable features in the frequency domain of CSI caused by the existence of an object, a time-frequency hybrid feature with both discriminative and robust characteristics is extracted through a time-frequency dual-branch collaborative architecture, thereby improving the generalization ability of the existence-aware model.

[0096] According to an embodiment of the present invention, determining the pairwise distribution differences between channel scattering sample features from multiple different domains based on the distribution characteristics of each of the multiple channel scattering sample features may include the following operations.

[0097] At least one channel scattering sample feature group is determined from multiple channel scattering sample features. The channel scattering sample feature group includes first channel scattering sample features and second channel scattering sample features from different domains. For each channel scattering sample feature group, the distribution features of the first channel scattering sample features and the second channel scattering sample features are extracted respectively to obtain the distribution characteristics of the first channel scattering sample features and the second channel scattering sample features. The cross-domain difference between the distribution characteristics of the first channel scattering sample features and the second channel scattering sample features is quantified using the Frobenius norm to obtain the distribution difference between the first channel scattering sample features and the second channel scattering sample features.

[0098] In determining the distributional differences between channel scattering sample characteristics in different domains, depth alignment can be performed, thereby reducing cross-domain differences without the need for target domain data.

[0099] During the calculation process, the channel scattering sample characteristics of any two domains are... and First, we can determine its distribution characteristics. For example, we can use the covariance matrix to characterize the distribution characteristics C, as shown in the following formula (3).

[0100] (3)

[0101] In the formula, This indicates the number of sub-sample features included in the channel scattering sample features. This represents a vector consisting entirely of 1s.

[0102] (4)

[0103] in, Indicates the distribution difference, where C represents the distribution characteristics of the first channel scattering sample features. This represents the distribution characteristics of the second channel scattering samples. Describe the Frobenius norm. Indicates the feature dimension.

[0104] According to an embodiment of the present invention, minimizing the distribution difference during training can make the feature distributions of different domains tend to be consistent.

[0105] According to embodiments of the present invention, by constructing cross-domain feature groups from channel scattering sample features of multiple domains and extracting the distribution characteristics of features in each domain, the modulation influence of environmental multipath structure on feature distribution can be effectively removed. By quantifying the differences between feature distributions in different domains using the Frobenius norm, alignment of cross-domain feature distributions can be achieved without target domain data. This alignment method suppresses the influence of environmental specificity on feature representation at the statistical distribution level, enabling the feature encoder to learn a common feature space decoupled from the specific environment, thereby improving the model's generalization ability in unknown environments and laying the foundation for zero-sample cross-domain object presence perception.

[0106] According to an embodiment of the present invention, determining the contrast loss based on the similarity between the features of each channel scattering sample and the category label of each channel scattering sample feature may include the following operations.

[0107] The features of each channel scattering sample are projected onto the contrast space to obtain multiple projected sample features. The feature dimension of the projected sample features is lower than that of the channel scattering sample features, and the information density is higher than that of the channel scattering sample features. Multiple target projected sample feature pairs randomly selected from the multiple projected sample features are fused to obtain multiple enhanced sample features. Anchor sample features are determined from the multiple enhanced sample features to serve as anchor points. The first similarity between the anchor sample features and enhanced sample features with the same category label in the contrast space, and the second similarity between the anchor sample features and multiple enhanced sample features in the contrast space are calculated. Based on the first similarity between the anchor sample features and enhanced sample features with the same category label in the contrast space, and the second similarity between the anchor sample features and multiple enhanced sample features in the contrast space, the contrast loss is determined.

[0108] The channel scattering sample features can be projected onto the contrast space using a nonlinear projection head G to obtain the initial projected sample features, and then normalized to distribute them on a unit hypersphere to obtain the final projected sample features z. The projection process can be shown in the following formula (5).

[0109] (5)

[0110] In the formula, Represents the ReLU activation function; , Both represent learnable weight matrices. Used to encode features Mapped to the hidden layer, Used to map hidden layer features to the projection space. express 3D real space, i.e., projective features It is 3D real-valued vector; This represents the dimension of the projection space, which can be equal to 128, and the number of hidden layer units can be set to 256.

[0111] Can be conduct Normalize it so that it is distributed on a unit hypersphere. Normalization refers to... The norm-based feature standardization operation maps features to a unit hypersphere, eliminating the influence of feature magnitude and retaining only orientation information, thereby improving the performance of subsequent classification tasks.

[0112] There are no restrictions on the enhancement strategy used to obtain enhanced sample features; it can be mixup enhancement, cutmix enhancement, etc.

[0113] In the implementation process, the features of two target projection samples included in the target projection sample pair can be mixed according to a preset mixing coefficient to obtain enhanced sample features. This method can expand the distribution range of similar samples in the contrast space, improving the model's tolerance to changes in object pose and position. Simultaneously, the introduction of soft labels allows the model to learn more fundamental category features, avoiding overfitting, thus making similar samples more closely aligned and dissimilar samples more distant, resulting in stronger generalization ability.

[0114] Anchor sample features can be determined from multiple augmented sample features, and there can be multiple anchor sample features.

[0115] For each anchor sample feature, other augmented sample features with the same category label as the anchor sample feature can be identified as the positive sample set. The cosine similarity between the anchor sample feature and each positive sample feature in the positive sample set is calculated in the contrast space as the similarity score; at the same time, the similarity between the anchor sample feature and all augmented sample features in the contrast space is calculated as the second similarity score.

[0116] Therefore, the contrast loss components of each anchor point sample feature are constructed using the first and second similarities mentioned above. When there is only one anchor point sample feature, this contrast loss component is used as the contrast loss. If there are multiple anchor point sample features, the average of the multiple contrast loss components is used as the contrast loss.

[0117] According to embodiments of the present invention, by projecting high-dimensional channel scattering sample features into a low-dimensional contrast space, a unified metric benchmark for feature similarity measurement is established while increasing information density. Furthermore, by fusing the projected sample features to generate enhanced samples, the distribution coverage of similar samples in the contrast space is effectively expanded. Based on this, by calculating the first similarity between anchor sample features and similar samples, and the second similarity with all samples, and constructing a contrast loss based on both, the model can simultaneously enhance the clustering degree of similar samples and the separation degree of dissimilar samples during the optimization process. This improves the discriminative ability of channel scattering sample features for target states, providing a more robust feature representation for subsequent cross-domain sensing tasks.

[0118] According to an embodiment of the present invention, feature fusion is performed on multiple target projection sample feature pairs randomly selected from multiple projection sample features to obtain multiple enhanced sample features, which may include the following operations.

[0119] Randomly select two target projection sample features from multiple projection sample features; perform linear interpolation on the two target projection sample features according to the mixing coefficient sampled from the preset distribution to generate enhanced sample features; perform linear interpolation on the category labels of the two target projection sample features according to the mixing coefficient to obtain the category labels corresponding to the enhanced sample features; repeat the above steps until a preset number of enhanced sample features are obtained.

[0120] Two target projection sample features can be randomly selected, and linear interpolation can be performed according to the mixing coefficient to generate a new enhanced sample feature. At the same time, the class labels of the two samples are linearly interpolated according to the same mixing coefficient to obtain the soft label corresponding to the enhanced sample feature.

[0121] The mixing coefficient can be ,in, express Follow the interval A uniform distribution on the surface, if the distribution parameter That is, the mixing coefficient can be obtained by uniform random sampling within the range of 0.9 to 1.0.

[0122] Repeat the above process until 2N enhanced sample features and their corresponding class labels are obtained, where N is the number of projected sample features.

[0123] For example: features of projected samples Corresponding category labels The N projected feature label pairs constitute 2N enhanced feature label pairs can be generated using a preset feature enhancement strategy. c represents the c-th projected feature label pair, and b represents the b-th augmented feature label pair.

[0124] According to embodiments of the present invention, by enhancing sample features as described above, on the one hand, the distribution coverage of samples of the same category in the contrast space can be expanded, making the model more tolerant to feature fluctuations caused by changes in object pose and position, thereby improving the generalization ability of features. On the other hand, by introducing soft labels, enhanced samples can slightly perturb feature positions while maintaining a clear category identity, enabling the model to learn more fundamental category discrimination boundaries, reducing overfitting to specific samples, and thus making the supervised contrastive loss more stably achieve the aggregation of similar categories and the separation of dissimilar categories during the optimization process.

[0125] According to an embodiment of the present invention, determining the contrast loss based on the first similarity between the anchor sample features and the enhanced sample features with the same category label in the contrast space and the second similarity with multiple enhanced sample features in the contrast space may include the following operations.

[0126] The first similarity contribution value is obtained by scaling each first similarity and taking its exponent. The second similarity contribution values ​​are then summed after scaling each second similarity value and taking its exponent. The first similarity contribution value is divided by the sum of the second similarity contribution values ​​to obtain the confidence of the anchor sample feature relative to the augmented sample feature with the same class label. The contrast loss component of the anchor sample feature is obtained based on the confidence and the number of augmented sample features with the same class label as the anchor sample feature. When there are multiple anchor sample features, the contrast loss components of the multiple anchor sample features are averaged to obtain the contrast loss.

[0127] Both the first and second similarity scores can be scaled using the temperature parameter.

[0128] The process of generating contrast loss using the first similarity and the second similarity can be shown in the following formulas (6) and (7).

[0129] (6)

[0130] (7)

[0131] In the formula, This is a temperature parameter used to adjust the weights of feature similarity. This indicates an enhancement of sample features. For anchor points and enhanced sample features The set of all positive samples of the same category has the number of elements equal to the number of positive samples. For indicator functions, when the condition inside the parentheses... (i.e., enhancing sample features) Enhanced sample features The value is 1 if the category is the same, otherwise the value is 0. , , These are the projected enhanced sample features l, j, k, respectively. Indicates comparative loss, This represents the contrastive loss component of anchor sample feature i.

[0132] According to embodiments of the present invention, by scaling, exponentially calculating, summing, and dividing the first and second similarities, the similarity between the anchor sample and samples of the same class can be placed in a global comparison with all samples, thereby obtaining the confidence level that the anchor point is correctly identified as its own category. Taking the negative logarithm of this confidence level and dividing it by the number of samples of the same class ensures that the loss function, during optimization, not only focuses on the clustering degree of samples of the same class but also implicitly increases the separation degree of samples of different classes through a global normalization mechanism.

[0133] The loss components of all anchor points are averaged to ensure that the model can optimize the contrast relationship of each sample in a balanced manner during training. The resulting contrast loss can effectively bring similar samples closer together and push dissimilar samples further apart in the contrast space, enhancing the ability of channel scattering sample features to discriminate the target state, and providing the model with more discriminative and robust feature representations in cross-domain perception tasks.

[0134] According to an embodiment of the present invention, the initial presence-aware model includes an initial classifier; the method may also include the following operations.

[0135] Using an initial classifier, the features of each channel scattering sample are mapped to test perception results. For each channel scattering sample feature, the classification loss corresponding to the channel scattering sample feature is determined based on the test perception results and category labels.

[0136] According to an embodiment of the present invention, the calculation of classification loss provides a more accurate supervision signal for the initial classifier. By quantifying the deviation between the model prediction result and the true label, the optimization direction of the classifier parameters can be clarified, thereby promoting the classifier to improve the class discrimination accuracy of the channel scattering sample features and making the test perception result more consistent with the true class label of the sample.

[0137] According to an embodiment of the present invention, the test perception result includes: a test probability value corresponding to the existence state of the object and a test uncertainty of the probability value; based on the test perception result and category label of the channel scattering sample features, determining the classification loss corresponding to the channel scattering sample features may include the following operations.

[0138] Determine the square of the difference between the test probability value and the class label; combine the constraint terms of the test uncertainty and the result of the division operation between the square of the difference and the test uncertainty to determine the classification loss. The classification loss is used to constrain the test probability value to gradually approach the class label, while also constraining the test probability value and the test uncertainty to match.

[0139] The model prediction can be modeled as a Gaussian distribution, and the test perception results output by the classifier can include the mean μ and variance of the predicted class distribution. That is, the mean can be used to represent the test probability value corresponding to the existence state of the object, and the variance can be used to represent the test uncertainty of the probability value.

[0140] Therefore, the classification loss L based on Gaussian negative log-likelihood (NLL) can be obtained. C Specifically, it can be shown in the following formula (8).

[0141] (8)

[0142] Where y represents the class label. The loss is derived from the prediction Gaussian distribution. With true labels (Dirac delta function) KL divergence between ).

[0143] According to embodiments of the present invention, by introducing a constraint term for test uncertainty, the situation where the classifier outputs excessive uncertainty to evade prediction error penalties is effectively reduced, making the model output meaningful. Simultaneously, dividing the square of the difference by the uncertainty allows the penalty weight for prediction bias to adaptively adjust with the uncertainty. When the uncertainty is large, the contribution of the current bias to the loss is weakened; when the uncertainty is small, the bias is strictly constrained. The combination of these two methods allows the classification loss to simultaneously constrain prediction accuracy and confidence calibration during the optimization process. The output test probability value not only more closely approximates the true label, but its accompanying uncertainty also truly reflects the reliability of the prediction results.

[0144] According to an embodiment of the present invention, the joint optimization objective further includes: minimizing the classification loss corresponding to the features of each channel scattering sample; training the initial existence-aware model with minimizing distribution difference and contrast loss as the joint optimization objective to obtain the existence-aware model may include the following operations.

[0145] Minimizing the distribution difference, the classification loss corresponding to the characteristics of each channel scattering sample, and the contrast loss are used as joint optimization objectives to train the initial existence-aware model, thus obtaining the existence-aware model.

[0146] During joint optimization, a joint loss value can be generated using distribution difference, classification loss, and contrastive loss, and model training can be achieved by minimizing this joint loss value. This joint loss value can be used not only for joint training of the initial feature encoder and initial classifier included in the model, but also for joint training of the initial feature encoder, initial classifier, and projection head G. The joint optimization objective can be shown in the following formula (9).

[0147] (9)

[0148] in, , , These represent the learnable parameters of the initial feature encoder, initial classifier, and projector head, respectively. for and The weights are used to optimize the model parameters using stochastic gradient descent, with the learning rate set to 0.001.

[0149] According to embodiments of the present invention, by jointly optimizing domain alignment loss, classification loss and contrast loss, the model achieves a balance between suppressing environmental distribution shift, improving feature discriminativeness and maintaining prediction accuracy, so that the trained presence-aware model can achieve high-precision object presence awareness across scenes and with zero samples without the need for target domain data.

[0150] Figure 7 A model architecture diagram of a cross-domain object existence awareness model training method according to an embodiment of the present invention is shown.

[0151] like Figure 7 As shown in the diagram, the model architecture can be divided into a feature encoder, a classifier, a domain alignment module, and a supervised contrastive learning module. During training, these modules are connected sequentially and work together to complete the entire perception process from CSI input to object presence determination output. The domain alignment module is used to calculate the distribution differences between different domains. The supervised contrastive learning module is used to calculate the contrastive loss.

[0152] During training, Channel State Sample Information (CSI) can be acquired in multiple different physical environments, each corresponding to a different source domain, as shown in the figure as Source Domain 1, Source Domain 2...Source Domain n. The acquired CSI is preprocessed to output a stable phase difference sequence, which is then input into the subsequent feature encoder, thus outputting a feature representation.

[0153] The feature representations are input into the classifier, the domain alignment module, and the supervised contrastive learning module, respectively, to obtain the classification loss L. C Differences in distribution across different domains and comparative loss The modules are then jointly trained based on the above loss. In the supervised contrastive learning module, the scattering samples from each channel are divided into negative and positive samples based on their category labels. Furthermore, the concept of anchor samples is introduced during feature enhancement to obtain a more accurate and effective contrastive loss.

[0154] Figure 8 A schematic diagram of collaborative feature optimization based on domain alignment and supervised contrastive learning according to an embodiment of the present invention is shown.

[0155] like Figure 8 As shown, the domain alignment module and supervised contrastive learning can constitute the feature optimization module. That is, the feature optimization module can take into account both the functions of inter-domain alignment and supervised contrastive learning. Input data from different source domains (i.e., different environments, such as environment one, environment two, and environment three) are processed by feature extraction and then input into this feature optimization module. Figure 8 The differences in feature distribution across various environments are illustrated. The feature optimization module aligns the statistical distributions of features from different source domains, mitigating the interference of environmental differences on feature representation. Furthermore, it leverages sample label information to guide the aggregation of similar features and the separation of dissimilar features. The loss value obtained from the feature optimization module is then used to train the initial presence-aware model. This training enables the presence-aware model to achieve more accurate classification under different domains and environmental conditions, effectively alleviating the problems of inconsistent CSI feature distributions and insufficient class discriminative power under different environmental conditions, thereby improving the model's accuracy in classifying room states.

[0156] Figure 9(a) shows the performance comparison results of different domain generalization methods according to embodiments of the present invention in cross-environment detection tasks.

[0157] As shown in Figure 9(a), the comparison methods include Maximum Mean Discrepancy (MMD), Model-Agnostic Meta-Learning (MAML), Domain-Adversarial Neural Network (DANN), and the cross-domain object presence awareness model training method proposed in this invention. All methods were trained under the same conditions and tested uniformly on unseen environmental data. The method of this invention shows more stable performance in terms of accuracy and F1 score, indicating that it achieves a better balance between generalization and discriminative abilities.

[0158] Figure 9(b) shows a schematic diagram comparing the feature distribution of the training set with and without the introduction of domain alignment and supervised contrastive learning according to an embodiment of the present invention.

[0159] Figure 9(c) shows a schematic diagram comparing the feature distribution of the test set with and without the introduction of domain alignment and supervised contrastive learning according to an embodiment of the present invention.

[0160] In Figure 9(b), 911 is the visualization effect of the training set feature distribution of the existence-aware model trained by the feature optimization module, and 912 is the visualization effect of the training set feature distribution of the existence-aware model without the feature optimization module.

[0161] In Figure 9(c), 921 is the visualization effect of the training set feature distribution of the existence-aware model trained by the feature optimization module, and 922 is the visualization effect of the training set feature distribution of the existence-aware model without the feature optimization module.

[0162] It can be seen that by introducing a feature optimization module to train the initial existence perception model, the clustering of similar features is more compact and the category boundaries are clearer in different environments, and this characteristic remains consistent in unseen environment data.

[0163] Figure 10 A flowchart of an object-aware method according to an embodiment of the present invention is shown.

[0164] like Figure 10 As shown, the object perception method includes operations S1010 to S1030.

[0165] In operation S1010, in response to receiving a presence awareness task for the target domain, the channel state information collected by the wireless communication device in the target domain is corrected for multi-dimensional phase error to obtain target phase data.

[0166] In operation S1020, the existence perception model trained by the above-mentioned cross-domain object existence perception model training method is used to perform object existence perception based on the target phase data to obtain the initial perception result, which includes the probability value corresponding to the object existence state and the uncertainty of the probability value.

[0167] In operation S1030, based on the fusion method corresponding to the uncertainty, the estimated perception result determined by the historical perception result and historical process noise is fused with the initial perception result to obtain the target perception result of whether an object exists in the target domain.

[0168] The multi-dimensional phase error correction process in the actual perception process and the reasoning process of the existence perception model are the same as the training process described above, and will not be repeated here.

[0169] After obtaining the probability value and uncertainty of the probability value corresponding to the existence state of the object, we can perform uncertainty modeling and Kalman filtering post-processing, that is, integrate the model output results into Kalman filtering to achieve temporal smoothing and suppress classification fluctuations caused by instantaneous noise.

[0170] For example: Suppose that at time t, the probability value output by the perception model is... The corresponding uncertainty is ,in This represents the probability that an object exists in the current domain (e.g., a room) at the current moment. Used to characterize the degree of uncertainty of the model regarding this probability value.

[0171] If the potential continuous representation of the actual occupancy status of a room (whether an object exists) is as follows: If its evolution over time satisfies a first-order Markov process, then the state transition model and the observation model can be represented by the following formulas (10) and (11).

[0172]

[0173]

[0174] In the formula, This represents the potential occupancy state at time t, i.e., the estimated perception result. This represents the potential or actual occupancy state at time t-1. Indicates that the observed value is taken = This is equal to the probability value output by the existence perception model. The process noise is used to describe the uncertainty of the natural change in room occupancy over time, and Q represents the process noise covariance parameter. The observation noise is used to describe the model prediction error, where the noise covariance is... Based on the dynamic setting of model prediction uncertainty, i.e. .

[0175] By using the above modeling method, the uncertainty of the model output can be directly involved in the filtering calculation process, thereby achieving adaptive adjustment of the observation reliability.

[0176] For example, at each time step, through the prediction and update process of Kalman filtering, the probability value corresponding to the object's existence state is fused with the estimated perception result obtained from the historical perception result of the previous time step and the historical process noise to obtain a smoothed estimate of the probability value corresponding to the object's existence state. .

[0177] The state covariance matrix is ​​updated simultaneously during the filtering process to characterize the magnitude of the estimation error. When the uncertainty... A larger value indicates lower confidence in the current prediction, and the filter automatically reduces the impact of the current model's output probability value on the state update. When the value is small, the weight of the current model output probability value is increased, thereby achieving dynamic weighted fusion based on uncertainty.

[0178] Set a preset decision threshold for the smoothed estimate of the dynamic weighted fusion. The room occupancy determination can be performed as shown in the following formula (12). After determining... Greater than or equal to In the case where the target perception result is determined to be an object existing in the room, otherwise, when it is less than... In this case, the target perception result is that there is no object in the room.

[0179] (12)

[0180] in, Indicates the preset decision threshold, and the preferred option is... , This indicates the result of target perception.

[0181] Through the above processing, while maintaining sensitivity to changes in the real state, isolated false alarms and short-term jitter can be effectively suppressed, thereby improving the stability and reliability of the detection results in the time dimension.

[0182] Figure 11 A schematic diagram is shown before and after smoothing the initial perception result using the estimated perception result according to an embodiment of the present invention.

[0183] like Figure 11 As shown, due to wireless channel noise and transient interference, the perception model is prone to short-term jumps in "manned / unmanned" state at the frame-level inference level. To address this, this invention uses the initial perception result output by the classifier as the observation value and uncertainty input to the Kalman filter, dynamically adjusting the observation reliability based on the uncertainty. Specifically, when the prediction uncertainty is high, the filter relies more on historical state estimates; when the prediction uncertainty is low, the influence of the current observation on state updates is enhanced. Figure 11 Actual occupancy status The probability value output by the existence-aware model The target perception result is obtained by smoothing the estimated perception result. Their respective waveforms and preset decision thresholds It is evident that the present invention can effectively suppress isolated false alarms and short-term jitter, so that the final detection result can maintain sensitivity to changes in the real state while significantly reducing unnecessary frequent switching.

[0184] Experiments show that this invention exhibits good stability and robustness under various scenario conditions. Extensive comparative verification results across multiple environmental conditions demonstrate that the overall detection accuracy of this invention in cross-domain object presence detection tasks is superior to related technologies, effectively suppressing performance degradation caused by environmental changes and layout differences. Compared to related domain adaptive schemes that rely on target domain samples for model adjustment, this invention achieves stable generalization to unknown scenarios without acquiring any target domain data during the training phase, significantly reducing system deployment and maintenance costs and improving the applicability and flexibility of the method in practical applications.

[0185] In summary, through the integrated design of time-frequency hybrid attention feature extraction, domain alignment and supervised contrastive learning collaborative optimization, multi-loss joint training and time-series post-processing, this invention solves the cross-domain generalization problem of WiFi object presence awareness, achieves accurate perception without target domain data, is suitable for practical deployment scenarios such as smart homes and smart security, and has high engineering application value.

[0186] Based on the aforementioned cross-domain object existence perception model training method and object perception method, this invention also provides a cross-domain object existence perception model training device and an object perception device. The following will be combined with... Figure 12 and Figure 13 The above-mentioned device will be described in detail.

[0187] Figure 12 A structural block diagram of a cross-domain object existence awareness model training device according to an embodiment of the present invention is shown.

[0188] like Figure 12 As shown, the cross-domain object existence perception model training device 1200 includes a first correction module 1210, a feature extraction module 1220, a first determination module 1230, a second determination module 1240, and a training module 1250.

[0189] The first correction module 1210 is used to perform multi-dimensional phase error correction on channel state sample information of multiple different domains to obtain multiple target phase sample data. In one embodiment, the first correction module 1210 can be used to perform the operation S210 described above, which will not be repeated here.

[0190] The feature extraction module 1220 is used to extract features from the target phase sample data using the initial feature encoder of the initial presence-aware model to obtain multiple channel scattering sample features. In one embodiment, the feature extraction module 1220 can be used to perform the operation S220 described above, which will not be repeated here.

[0191] The first determining module 1230 is used to determine the distribution differences between each pair of channel scattering sample features in multiple different domains based on the distribution characteristics of each channel scattering sample feature. In one embodiment, the first determining module 1230 can be used to perform the operation S230 described above, which will not be repeated here.

[0192] The second determining module 1240 is used to determine the contrast loss based on the similarity between the features of each channel scattering sample and the category label of each channel scattering sample feature. The contrast loss is used to guide the initial existence-aware model to increase the degree of aggregation between channel scattering sample features with the same category label and the degree of separation between channel scattering sample features with different category labels. In one embodiment, the second determining module 1240 can be used to perform the operation S240 described above, which will not be repeated here.

[0193] The training module 1250 is used to train the initial existence-aware model with the joint optimization objective of minimizing distribution difference and contrastive loss, thereby obtaining the existence-aware model. In one embodiment, the training module 1250 can be used to perform the operation S250 described above, which will not be repeated here.

[0194] Figure 13 A structural block diagram of an object sensing device according to an embodiment of the present invention is shown.

[0195] like Figure 13 As shown, the object sensing device 1300 includes a second correction module 1310, a sensing module 1320, and a smoothing module 1330.

[0196] The second correction module 1310 is used to perform multi-dimensional phase error correction on the channel state information collected by the wireless communication device in the target domain in response to receiving the presence awareness task for the target domain, so as to obtain the target phase data. In one embodiment, the second correction module 1310 can be used to perform the operation S1010 described above, which will not be repeated here.

[0197] The perception module 1320 is used to use the existence perception model to perceive the existence of an object based on the target phase data and obtain an initial perception result. The initial perception result includes the probability value corresponding to the existence state of the object and the uncertainty of the probability value. In one embodiment, the perception module 1320 can be used to perform the operation S1020 described above, which will not be repeated here.

[0198] The smoothing module 1330 is used to fuse the estimated perception result determined by historical perception results and historical process noise with the initial perception result based on the fusion method corresponding to the uncertainty, to obtain the target perception result of whether an object exists in the target domain. In one embodiment, the smoothing module 1330 can be used to perform the operation S1030 described above, which will not be repeated here.

[0199] According to embodiments of the present invention, any plurality of modules among the first correction module 1210, feature extraction module 1220, first determination module 1230, second determination module 1240, training module 1250, second correction module 1310, perception module 1320, and smoothing module 1330 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of the present invention, at least one of the first correction module 1210, feature extraction module 1220, first determination module 1230, second determination module 1240, training module 1250, second correction module 1310, perception module 1320, and smoothing module 1330 can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or any other reasonable means of integrating or packaging circuits, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the first correction module 1210, feature extraction module 1220, first determination module 1230, second determination module 1240, training module 1250, second correction module 1310, perception module 1320, and smoothing module 1330 can be at least partially implemented as computer program modules, which can perform corresponding functions when the computer program module is run.

[0200] Figure 14 A block diagram of an electronic device suitable for implementing a cross-domain object presence awareness model training method and an object awareness method according to an embodiment of the present invention is shown.

[0201] like Figure 14 As shown, an electronic device 1400 according to an embodiment of the present invention includes a processor 1401, which can perform various appropriate actions and processes according to a program stored in a read-only memory ROM 1402 or a program loaded from a storage portion 1408 into a random access memory RAM 1403. The processor 1401 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1401 may also include onboard memory for caching purposes. The processor 1401 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0202] RAM 1403 stores various programs and data required for the operation of electronic device 1400. Processor 1401, ROM 1402, and RAM 1403 are interconnected via bus 1404. Processor 1401 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 1402 and / or RAM 1403. It should be noted that the programs may also be stored in one or more memories other than ROM 1402 and RAM 1403. Processor 1401 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in said one or more memories.

[0203] According to an embodiment of the present invention, the electronic device 1400 may further include an input / output (I / O) interface 1405, which is also connected to a bus 1404. The electronic device 1400 may also include one or more of the following components connected to the input / output (I / O) interface 1405: an input section 1406 including a keyboard, mouse, etc.; an output section 1407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1408 including a hard disk, etc.; and a communication section 1409 including a network interface card such as a LAN card, modem, etc. The communication section 1409 performs communication processing via a network such as the Internet. A drive 1410 is also connected to the input / output (I / O) interface 1405 as needed. A removable medium 1411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1410 as needed so that computer programs read from it can be installed into the storage section 1408 as needed.

[0204] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0205] According to embodiments of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of the present invention, a computer-readable storage medium may include ROM 1402 and / or RAM 1403 and / or one or more memories other than ROM 1402 and RAM 1403 described above.

[0206] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the cross-domain object presence awareness model training method and object awareness method provided in the embodiments of the present invention.

[0207] When the computer program is executed by the processor 1401, it performs the functions defined in the system / apparatus of this embodiment of the invention. According to embodiments of the invention, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0208] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1409, and / or installed from the removable medium 1411. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0209] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1409, and / or installed from the removable medium 1411. When the computer program is executed by the processor 1401, it performs the functions defined in the system of this embodiment of the invention. According to embodiments of the invention, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0210] According to embodiments of the present invention, program code for executing the computer programs provided in the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0211] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0212] Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.

[0213] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.

Claims

1. A method for training a cross-domain object existence perception model, characterized in that, The method includes: Multi-dimensional phase error correction is performed on channel state sample information from multiple different domains to obtain multiple target phase sample data; Using the initial feature encoder of the initial presence-aware model, feature extraction is performed on the phase sample data of each target to obtain multiple channel scattering sample features; Based on the distribution characteristics of each of the multiple channel scattering sample features, the distribution differences between each pair of channel scattering sample features from multiple different domains are determined; Based on the similarity between the channel scattering sample features and the category label of each channel scattering sample feature, a contrast loss is determined. The contrast loss is used to guide the initial presence-aware model to increase the aggregation degree between channel scattering sample features with the same category label and the separation degree between channel scattering sample features with different category labels. Using the initial classifier of the initial existence perception model, the features of each channel scattering sample are mapped to test perception results, the test perception results including test probability values ​​corresponding to the existence state of the object and test uncertainty of the probability values; For each channel scattering sample feature, determine the square of the difference between the test probability value corresponding to the channel scattering sample feature and the category label; The classification loss is determined by combining the constraint term of the test uncertainty corresponding to the channel scattering sample characteristics and the result of the division operation between the square of the difference and the uncertainty. The classification loss is used to constrain the test probability value to gradually approach the category label, and at the same time constrain the test probability value and the test uncertainty to match. The initial existence-aware model is trained with the common optimization objective of minimizing the distribution difference, the contrast loss, and the classification loss corresponding to the features of each channel scattering sample, to obtain the existence-aware model.

2. The method according to claim 1, characterized in that, The determination of pairwise distribution differences between channel scattering sample features from multiple different domains based on their respective distribution characteristics includes: At least one channel scattering sample feature group is determined from a plurality of channel scattering sample features, the channel scattering sample feature group including first channel scattering sample features and second channel scattering sample features from different domains; For each group of channel scattering sample features, the distribution features of the first channel scattering sample features and the second channel scattering sample features are extracted respectively to obtain the distribution characteristics of the first channel scattering sample features and the second channel scattering sample features respectively; By using the Frobenius norm, the cross-domain differences between the distribution characteristics of the first-channel scattering sample features and the second-channel scattering sample features are quantified, thus obtaining the distribution differences between the first-channel scattering sample features and the second-channel scattering sample features.

3. The method according to claim 1 or 2, characterized in that, The determination of contrast loss based on the similarity between the features of each channel scattering sample and the category label of each channel scattering sample feature includes: Each of the channel scattering sample features is projected onto the contrast space to obtain multiple projected sample features. The feature dimension of the projected sample features is lower than that of the channel scattering sample features, and the information density is higher than that of the channel scattering sample features. Multiple target projection sample feature pairs randomly selected from multiple projection sample features are respectively fused to obtain multiple enhanced sample features; Anchor sample features are determined from the plurality of augmented sample features to serve as anchor points, and the first similarity between the anchor sample features and augmented sample features with the same category label in the comparison space and the second similarity between the anchor sample features and the plurality of augmented sample features in the comparison space are calculated. The contrast loss is determined based on the first similarity in the contrast space between the anchor sample features and the enhanced sample features with the same category label, and the second similarity in the contrast space between the anchor sample features and multiple enhanced sample features.

4. The method according to claim 3, characterized in that, The determination of the contrast loss based on the first similarity in the contrast space between the anchor sample features and the enhanced sample features with the same category label, and the second similarity in the contrast space between the anchor sample features and multiple enhanced sample features, includes: The first similarity score is scaled and then exponentially calculated to obtain the first similarity contribution value. The sum of the contributions of the second similarity scores is obtained by scaling each score individually, taking the exponent, and summing the results. Divide each of the first similarity contribution values ​​and the sum of the second similarity contribution values ​​to obtain the confidence of the anchor sample feature relative to the enhanced sample feature with the same category label. Based on the confidence level and the number of augmented sample features with the same class label as the anchor sample features, the contrastive loss component of the anchor sample features is obtained. When there are multiple anchor point sample features, the contrast loss is obtained by averaging the contrast loss components of the multiple anchor point sample features.

5. The method according to claim 3, characterized in that, The feature fusion of multiple target projection sample feature pairs randomly selected from multiple projection sample features to obtain multiple enhanced sample features includes: Two target projection sample features are randomly selected from the plurality of projection sample features; Linear interpolation is performed on the features of the two target projection samples according to the mixing coefficients sampled from the preset distribution to generate enhanced sample features; The category labels of the two target projection sample features are linearly interpolated according to the mixing coefficient to obtain the category labels corresponding to the enhanced sample features; Repeat the above steps until the preset number of enhanced sample features are obtained.

6. The method according to claim 1, characterized in that, The target phase sample data is a phase difference sequence; The phase difference sequence includes: phase differences at multiple time points; the initial feature encoder includes: a time-domain feature extraction module, a frequency-domain feature extraction module, and a feature fusion module; Specifically, the step of using the initial feature encoder to extract features from each of the target phase sample data to obtain multiple channel scattering sample features includes: for each of the target phase sample data, Using the time-domain feature extraction module, the phase difference at multiple time points is extracted to obtain time-domain features. The time-domain feature extraction module adopts an attention network architecture and dynamically allocates attention weights to focus on the phase difference at time points related to the existence of the object. Using the frequency domain feature extraction module, the phase differences at multiple time points are Fourier transformed to obtain the frequency domain representations of multiple phase differences, and frequency domain features are extracted from the frequency domain representations of the multiple phase differences to obtain frequency domain features. The frequency domain feature extraction module adopts an attention network architecture that is isomorphic to the time domain feature extraction module. The feature fusion module is used to fuse the time-domain features and the frequency-domain features to obtain the channel scattering sample features.

7. The method according to claim 1, characterized in that, The target phase sample data is a phase difference sequence; the channel state sample information includes raw phase data; the raw phase data includes raw phase at multiple time points; The step of performing multi-dimensional phase error correction on channel state sample information from multiple different domains to obtain multiple target phase sample data includes: By performing linear transformation, the original phases at multiple time points are corrected for linear errors to obtain intermediate phase data; For the same time point, calculate the difference in the intermediate phase between adjacent receiving antennas to obtain the initial phase difference at that time point; The initial phase differences at each of the aforementioned time points are arranged in chronological order to obtain an initial phase difference sequence; An environmental noise is filtered out from the initial phase difference sequence using a smoothing filter to obtain the phase difference sequence.

8. An object-aware method, characterized in that, The method includes: In response to receiving a presence awareness task for a target domain, the channel state information collected by the wireless communication device in the target domain is corrected for multi-dimensional phase error to obtain target phase data. Using the existence perception model trained by the cross-domain object existence perception model training method according to any one of claims 1 to 7, object existence perception is performed based on the target phase data to obtain an initial perception result, wherein the initial perception result includes a probability value corresponding to the object existence state and the uncertainty of the probability value; Based on the fusion method corresponding to the uncertainty, the estimated perception result determined by the historical perception result and historical process noise is fused with the initial perception result to obtain the target perception result of whether an object exists in the target domain.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.