Communication state identification method and identification model training method

By combining the feature extraction and fusion of channel impulse response data and image data in ultra-wideband communication, and using the recognition model to identify the communication status, the problem of reduced positioning reliability caused by non-line-of-sight propagation is solved, and the accuracy of communication status recognition is improved.

CN120602028APending Publication Date: 2025-09-05ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511005310.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In ultra-wideband communications, non-line-of-sight propagation causes signals to reflect, diffract, or penetrate obstacles, resulting in reduced positioning reliability. The accuracy of communication status recognition in existing technologies is low.

Method used

By acquiring channel impulse response data and image data, the recognition model is used for feature extraction and fusion, and the communication status is identified by combining the multi-layer perceptron and modal attention mechanism.

Benefits of technology

The accuracy of communication status recognition is improved, especially in the presence of obstacles, and the line-of-sight or non-line-of-sight communication status can be determined more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602028A_ABST
    Figure CN120602028A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a communication state identification method and an identification model training method. The method comprises the following steps: acquiring a to-be-identified data set, wherein the to-be-identified data set comprises channel impact response data and image data; the channel impact response data represents channel characteristics when the receiving end equipment communicates with the transmitting end equipment, and the image data represents environment information when the receiving end equipment communicates with the transmitting end equipment; performing identification processing on the channel impact response data and the image data in the to-be-identified data group according to the identification model to obtain a communication state; wherein the communication state represents the communication state when the receiving end device communicates with the sending end device. The method is used for achieving the effect of improving the communication state recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wireless communication technology, and in particular to a method for identifying a communication state and a method for training an identification model. Background Art

[0002] With the rapid development of science and technology, the application of ultra-wideband technology is becoming increasingly widespread. Ultra-wideband communications can potentially experience non-line-of-sight (NLOS) transmission. This can cause signals to reflect, diffract, or penetrate obstacles, reducing positioning reliability. For example, when vehicles communicate with each other, NLOS can occur if a large building blocks the signal transmission path, or if a large truck blocks the path between communicating vehicles in traffic.

[0003] Some technologies extract features from Channel Impulse Response (CIR) data to identify the communication state, that is, whether the communication state is currently non-line-of-sight. However, these technologies have low accuracy in identifying the communication state.

[0004] Therefore, there is an urgent need for a solution that can improve the accuracy of communication status recognition. Summary of the Invention

[0005] The communication status recognition method and recognition model training method provided in the embodiments of the present application are used to achieve the effect of improving the accuracy of communication status recognition.

[0006] In a first aspect, an embodiment of the present application provides a method for identifying a communication status, including:

[0007] Acquire a data group to be identified, the data group to be identified including channel impulse response data and image data; the channel impulse response data represents channel characteristics when the receiving end device communicates with the transmitting end device, and the image data represents environmental information when the receiving end device communicates with the transmitting end device;

[0008] The channel impulse response data and image data in the data group to be identified are identified and processed according to the identification model to obtain a communication state; wherein the communication state represents the communication state when the receiving device and the transmitting device communicate.

[0009] In a possible implementation, performing recognition processing on the channel impulse response data and the image data in the data group to be recognized according to the recognition model to obtain the communication status includes:

[0010] Performing feature extraction processing on the channel impulse response data based on the recognition model to obtain channel impulse response features; and performing feature extraction processing on the image data based on the recognition model to obtain image features, wherein the image features represent whether the receiving device and the transmitting device are in the same environment when communicating;

[0011] Based on the recognition model, the communication status is determined according to the image features and channel impulse response characteristics.

[0012] In one possible implementation, the image data includes first image data and second image data; wherein the first image data is an environmental image captured by the sending device when the receiving device communicates with the sending device; and the second image data is an environmental image captured by the receiving device when the receiving device communicates with the sending device.

[0013] In one possible implementation, performing feature extraction processing on image data based on a recognition model to obtain image features includes:

[0014] Performing feature extraction processing on the first image data based on the recognition model to obtain a first feature; wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device;

[0015] Performing feature extraction processing on the second image data based on the recognition model to obtain a second feature; wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the transmitting device;

[0016] The first feature and the second feature are subjected to feature fusion processing to obtain image features.

[0017] In a possible implementation, performing feature extraction processing on the channel impulse response data based on the recognition model to obtain the channel impulse response features includes:

[0018] Based on the recognition model, the channel impulse response data is subjected to convolution and pooling processing to obtain local spatial features; wherein the local spatial features represent the data characteristics of the channel impulse response data;

[0019] Based on the gating mechanism of the recognition model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response characteristics.

[0020] In one possible implementation, determining the communication state based on the recognition model and according to image features and channel impulse response features includes:

[0021] Based on the recognition model, the image features and channel impulse response features are fused to obtain fused features;

[0022] The multi-layer perceptron based on the recognition model determines the communication status according to the fusion features.

[0023] In a possible implementation, based on the recognition model, feature fusion processing is performed on the image features and the channel impulse response features to obtain fused features, including:

[0024] A multi-layer perceptron based on the recognition model determines, based on the image features and the channel impulse response features, a third feature corresponding to the image features and a fourth feature corresponding to the channel impulse response features; wherein the third feature and the fourth feature belong to the same semantic space;

[0025] Based on the modal attention mechanism of the recognition model, according to the third feature and the fourth feature, a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature are determined respectively;

[0026] A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

[0027] In one possible implementation, obtaining a data group to be identified includes:

[0028] In response to receiving the polling data packet, image data is acquired; and signal analysis processing is performed on the polling data packet to obtain original impulse response sequence data;

[0029] The original impulse response sequence data is truncated to obtain channel impulse response data;

[0030] A data group to be identified is constructed based on the image data and the channel impulse response data.

[0031] In a possible implementation, the image data and the channel impulse response data have the same round identifier; wherein the round identifier is used to indicate the rounds of the polling data packets corresponding to the image data and the channel impulse response data.

[0032] In a second aspect, an embodiment of the present application provides a method for training a recognition model, comprising:

[0033] Acquire a training data set; wherein the training data set includes at least one data group, the data group including channel impulse response data and image data; the channel impulse response data represents channel characteristics when the receiving end device communicates with the transmitting end device, and the image data represents environmental information when the receiving end device communicates with the transmitting end device;

[0034] The initial model is trained according to the channel impulse response data and image data in the training data set to obtain a recognition model;

[0035] The identification model is used to identify the communication state when the receiving device communicates with the sending device.

[0036] In a possible implementation, training an initial model based on channel impulse response data and image data in a training data set to obtain a recognition model includes:

[0037] Performing feature extraction processing on the channel impulse response data based on the initial model to obtain channel impulse response features; and performing feature extraction processing on the image data based on the initial model to obtain image features, wherein the image features represent whether the receiving device and the transmitting device are in the same environment when communicating;

[0038] According to the image features and the channel impulse response features, the predicted communication state is obtained;

[0039] According to the predicted communication state and the actual communication state corresponding to the data group, the initial model is trained to obtain the recognition model.

[0040] In one possible implementation, the image data includes first image data and second image data; wherein the first image data is an environmental image captured by the sending device when the receiving device communicates with the sending device; and the second image data is an environmental image captured by the receiving device when the receiving device communicates with the sending device.

[0041] In one possible implementation, performing feature extraction processing on image data based on the initial model to obtain image features includes:

[0042] Performing feature extraction processing on the first image data based on the initial model to obtain a first feature; wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device;

[0043] Performing feature extraction processing on the second image data based on the initial model to obtain a second feature; wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the transmitting device;

[0044] The first feature and the second feature are subjected to feature fusion processing to obtain image features.

[0045] In a possible implementation, performing feature extraction processing on the channel impulse response data based on the initial model to obtain channel impulse response features includes:

[0046] Based on the initial model, convolution pooling is performed on the channel impulse response data to obtain local spatial features; wherein the local spatial features represent the data characteristics of the channel impulse response data;

[0047] Based on the gating mechanism of the initial model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response features.

[0048] In one possible implementation, obtaining a predicted communication state based on image features and channel impulse response features includes:

[0049] Based on the initial model, the image features and channel impulse response features are fused to obtain fused features;

[0050] The multi-layer perceptron based on the initial model determines the predicted communication status according to the fusion features.

[0051] In a possible implementation, based on the initial model, feature fusion processing is performed on the image features and the channel impulse response features to obtain fused features, including:

[0052] A multilayer perceptron based on the initial model determines, according to the image features and the channel impulse response features, a third feature corresponding to the image features and a fourth feature corresponding to the channel impulse response features; wherein the third feature and the fourth feature belong to the same semantic space;

[0053] Based on the modal attention mechanism of the initial model, according to the third feature and the fourth feature, a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature are determined respectively;

[0054] A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

[0055] In one possible implementation, the initial model is trained based on the predicted communication state and the actual communication state corresponding to the data set to obtain a recognition model, including:

[0056] Perform iterative training on the initial model based on the predicted communication state, the actual communication state corresponding to the data group, and a preset composite loss function;

[0057] Among them, when the number of iterative training reaches a preset number, the recognition model is obtained; the composite loss function includes: fusion classification result deviation term, image classification result deviation term, channel impulse response classification result deviation term, divergence minimization term and cosine similarity maximization term.

[0058] In a third aspect, an embodiment of the present application provides a communication status identification device, including:

[0059] A first acquisition module is configured to acquire a data group to be identified, the data group to be identified including channel impulse response data and image data; the channel impulse response data represents channel characteristics when a receiving device communicates with a transmitting device, and the image data represents environmental information when the receiving device communicates with the transmitting device;

[0060] The processing module is used to identify and process the channel impulse response data and image data in the data group to be identified according to the recognition model to obtain the communication status; wherein the communication status represents the communication status when the receiving device communicates with the sending device.

[0061] In a possible implementation, the channel impulse response data and the image data in the data group to be identified are identified and processed according to the identification model to obtain the communication status. The processing module is configured to:

[0062] Performing feature extraction processing on the channel impulse response data based on the recognition model to obtain channel impulse response features; and performing feature extraction processing on the image data based on the recognition model to obtain image features, wherein the image features represent whether the receiving device and the transmitting device are in the same environment when communicating;

[0063] Based on the recognition model, the communication status is determined according to the image features and channel impulse response characteristics.

[0064] In one possible implementation, the image data includes first image data and second image data; wherein the first image data is an environmental image captured by the sending device when the receiving device communicates with the sending device; and the second image data is an environmental image captured by the receiving device when the receiving device communicates with the sending device.

[0065] In one possible implementation, feature extraction processing is performed on image data based on a recognition model to obtain image features, and the processing module is configured to:

[0066] Performing feature extraction processing on the first image data based on the recognition model to obtain a first feature; wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device;

[0067] Performing feature extraction processing on the second image data based on the recognition model to obtain a second feature; wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the transmitting device;

[0068] The first feature and the second feature are subjected to feature fusion processing to obtain image features.

[0069] In a possible implementation, feature extraction processing is performed on the channel impulse response data based on the recognition model to obtain channel impulse response features. The processing module is configured to:

[0070] Based on the recognition model, the channel impulse response data is subjected to convolution pooling processing to obtain local spatial features; wherein the local spatial features represent the data characteristics of the channel impulse response data;

[0071] Based on the gating mechanism of the recognition model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response characteristics.

[0072] In one possible implementation, based on the recognition model, the communication state is determined according to image features and channel impulse response features, and the processing module is configured to:

[0073] Based on the recognition model, the image features and channel impulse response features are fused to obtain fused features;

[0074] The multi-layer perceptron based on the recognition model determines the communication status according to the fusion features.

[0075] In a possible implementation, based on the recognition model, feature fusion processing is performed on the image features and the channel impulse response features to obtain fused features. The processing module is configured to:

[0076] A multi-layer perceptron based on the recognition model determines, based on the image features and the channel impulse response features, a third feature corresponding to the image features and a fourth feature corresponding to the channel impulse response features; wherein the third feature and the fourth feature belong to the same semantic space;

[0077] Based on the modal attention mechanism of the recognition model, according to the third feature and the fourth feature, a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature are determined respectively;

[0078] A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

[0079] In a possible implementation, to obtain a data group to be identified, the first obtaining module is configured to:

[0080] In response to receiving the polling data packet, image data is acquired; and signal analysis processing is performed on the polling data packet to obtain original impulse response sequence data;

[0081] The original impulse response sequence data is truncated to obtain channel impulse response data;

[0082] A data group to be identified is constructed based on the image data and the channel impulse response data.

[0083] In a possible implementation, the image data and the channel impulse response data have the same round identifier; wherein the round identifier is used to indicate the rounds of the polling data packets corresponding to the image data and the channel impulse response data.

[0084] In a fourth aspect, an embodiment of the present application provides a training device for a recognition model, comprising:

[0085] A second acquisition module is configured to acquire a training data set; wherein the training data set includes at least one data group, the data group including channel impulse response data and image data; the channel impulse response data represents channel characteristics when the receiving device communicates with the transmitting device, and the image data represents environmental information when the receiving device communicates with the transmitting device;

[0086] A training module is used to train the initial model based on the channel impulse response data and image data in the training data set to obtain a recognition model;

[0087] The identification model is used to identify the communication state when the receiving device communicates with the sending device.

[0088] In a possible implementation, the initial model is trained based on the channel impulse response data and image data in the training data set to obtain a recognition model. The training module is used to:

[0089] Performing feature extraction processing on the channel impulse response data based on the initial model to obtain channel impulse response features; and performing feature extraction processing on the image data based on the initial model to obtain image features, wherein the image features represent whether the receiving device and the transmitting device are in the same environment when communicating;

[0090] According to the image features and the channel impulse response features, the predicted communication state is obtained;

[0091] According to the predicted communication state and the actual communication state corresponding to the data group, the initial model is trained to obtain the recognition model.

[0092] In one possible implementation, the image data includes first image data and second image data; wherein the first image data is an environmental image captured by the sending device when the receiving device communicates with the sending device; and the second image data is an environmental image captured by the receiving device when the receiving device communicates with the sending device.

[0093] In one possible implementation, feature extraction processing is performed on image data based on the initial model to obtain image features, and the training module is used to:

[0094] Performing feature extraction processing on the first image data based on the initial model to obtain a first feature; wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device;

[0095] Performing feature extraction processing on the second image data based on the initial model to obtain a second feature; wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the transmitting device;

[0096] The first feature and the second feature are subjected to feature fusion processing to obtain image features.

[0097] In one possible implementation, feature extraction processing is performed on the channel impulse response data based on the initial model to obtain channel impulse response features. The training module is used to:

[0098] Based on the initial model, convolution pooling is performed on the channel impulse response data to obtain local spatial features; wherein the local spatial features represent the data characteristics of the channel impulse response data;

[0099] Based on the gating mechanism of the initial model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response features.

[0100] In one possible implementation, the predicted communication state is obtained based on the image features and the channel impulse response features, and the training module is used to:

[0101] Based on the initial model, the image features and channel impulse response features are fused to obtain fused features;

[0102] The multi-layer perceptron based on the initial model determines the predicted communication status according to the fusion features.

[0103] In one possible implementation, based on the initial model, feature fusion processing is performed on the image features and the channel impulse response features to obtain fused features. The training module is used to:

[0104] A multilayer perceptron based on the initial model determines, according to the image features and the channel impulse response features, a third feature corresponding to the image features and a fourth feature corresponding to the channel impulse response features; wherein the third feature and the fourth feature belong to the same semantic space;

[0105] Based on the modal attention mechanism of the initial model, according to the third feature and the fourth feature, a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature are determined respectively;

[0106] A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

[0107] In one possible implementation, the initial model is trained based on the predicted communication state and the actual communication state corresponding to the data set to obtain a recognition model. The training module is used to:

[0108] Perform iterative training on the initial model based on the predicted communication state, the actual communication state corresponding to the data group, and a preset composite loss function;

[0109] Among them, when the number of iterative training reaches a preset number, the recognition model is obtained; the composite loss function includes: fusion classification result deviation term, image classification result deviation term, channel impulse response classification result deviation term, divergence minimization term and cosine similarity maximization term.

[0110] In a fifth aspect, an embodiment of the present application provides a receiving device, including: a memory, a processor;

[0111] Memory stores computer-executable instructions;

[0112] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0113] In a sixth aspect, an embodiment of the present application provides a host computer device, including: a memory, a processor;

[0114] Memory stores computer-executable instructions;

[0115] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above second aspect and / or various possible implementations of the second aspect.

[0116] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method provided in the first or second aspect above.

[0117] In an eighth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method provided in the first or second aspect above.

[0118] The communication status identification method and recognition model training method provided in the embodiments of the present application obtain channel impulse response data and image data from the transmitter and receiver, input the data to be identified containing these two data into a trained recognition model for processing, and obtain the communication status of the transmitter and receiver during communication. By introducing image data, the recognition accuracy of the communication status of the transmitter and receiver during communication is improved. Based on the initial model, iterative training is performed in combination with the channel impulse response data and image data to obtain a recognition model for identifying the communication status, which can accurately identify the communication status. BRIEF DESCRIPTION OF THE DRAWINGS

[0119] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0120] Figure 1 A schematic diagram of a scenario for the communication status identification method provided in this application;

[0121] Figure 2 Schematic diagram of the process of identifying the communication status provided by this application Figure 1 ;

[0122] Figure 3 Schematic diagram of the process of identifying the communication status provided by this application Figure 2 ;

[0123] Figure 4 Schematic diagram of the process of identifying the communication status provided by this application Figure 3 ;

[0124] Figure 5 is a schematic diagram of an exemplary image feature extraction process;

[0125] Figure 6 Schematic diagram of the process of identifying the communication status provided by this application Figure 4 ;

[0126] Figure 7 is a schematic diagram of an exemplary channel impulse response feature extraction process;

[0127] Figure 8 Schematic diagram of the process of identifying the communication status provided by this application Figure 5 ;

[0128] Figure 9 A flowchart of the training method for the recognition model provided in this application;

[0129] Figure 10 A schematic diagram of the structure of the communication status identification device provided by this application;

[0130] Figure 11 A schematic diagram of the structure of the training device for the recognition model provided in this application;

[0131] Figure 12 A schematic diagram of the structure of the receiving device provided in this application;

[0132] Figure 13 This is a schematic diagram of the structure of the host computer device provided in this application.

[0133] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0134] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0135] First, let’s explain the terms involved in this application:

[0136] Ultra Wide Band (UWB) technology refers to a wireless communication technology that achieves high temporal resolution (sub-nanosecond level) by emitting extremely short pulses (nanosecond level). The operating frequency band is usually 3.1GHz-10.6GHz, and the bandwidth exceeds 500MHz.

[0137] Line of Sight (LOS) communication state: refers to the communication state in which there is a clear, unobstructed straight electromagnetic wave propagation path between the transmitter and the receiver.

[0138] Non-line of Sight (NLOS) communication refers to a communication state in which there is no clear, unobstructed, straight electromagnetic wave propagation path between the transmitter and receiver. Understandably, in NLOS, the signal propagation path is blocked by obstacles (such as walls, people, and furniture), rendering the direct path unusable. Signals must propagate through reflection, diffraction, or penetration.

[0139] Channel Impulse Response (CIR): This data represents the time-domain response of a wireless channel to an impulse signal, characterizing the amplitude, delay, and attenuation characteristics of multipath propagation. CIR data can be used in areas such as ranging, non-line-of-sight communication status identification, and railway communications.

[0140] With the rapid development of science and technology, the application of ultra-wideband technology is becoming increasingly widespread. In ultra-wideband communications, non-line-of-sight (NLOS) transmission often occurs. This can cause signals to reflect, diffract, or penetrate obstacles, reducing positioning reliability. For example, when vehicles communicate with each other, NLOS transmission can occur if a large building blocks the signal transmission path, or if a large truck blocks the path between communicating vehicles in traffic.

[0141] For example, Figure 1 This is a scene diagram of the communication status identification method provided by this application. The specific application scenario of this application can be applied to the communication scenario of the vehicle. Figure 1 As shown in FIG. 1 , a possible application scenario includes a vehicle 101 and an obstacle 102. The two vehicles 101 are communicating, but the obstacle 102 exists along the signal transmission path between the two vehicles 101. Therefore, the communication between the two vehicles 101 is in a non-line-of-sight (NLOS) state.

[0142] It should be noted that the obstacle 102 may be another vehicle or another object. For example, the obstacle may be another object such as a building or a fence.

[0143] In some embodiments, during the communication process, the receiving end and the transmitting end identify the communication status by extracting features of Channel Impulse Response (CIR) data.

[0144] In the above embodiment, the communication status is identified by extracting CIR data, which has a technical problem of low communication status identification accuracy.

[0145] The communication status identification method provided in this application obtains channel impulse response data and image data from the transmitter and receiver, then feeds the data to be identified, including these two data types, into a trained recognition model for processing. The method then determines the communication status of the transmitter and receiver during communication. By introducing image data, the accuracy of identifying the communication status of the transmitter and receiver during communication is improved.

[0146] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0147] Figure 2 Schematic diagram of the process of identifying the communication status provided by this application Figure 1 ,like Figure 2 As shown, the method includes:

[0148] Step 201: Obtain a data group to be identified.

[0149] The data group to be identified includes channel impulse response data and image data; the channel impulse response data represents the channel characteristics when the receiving device communicates with the transmitting device, and the image data represents the environmental information when the receiving device communicates with the transmitting device.

[0150] For example, the communication status identification method provided in this embodiment can be applied to a receiving device. The receiving device can communicate data with the transmitting device. In a vehicle application scenario, it can be applied to a vehicle receiving information. The receiving vehicle can communicate data with the transmitting vehicle.

[0151] The data group to be identified is obtained, and the data to be identified is analyzed and processed to obtain the communication status when the receiving device and the sending device are communicating data.

[0152] The data group to be identified may include multiple types of data. For example, the data group to be identified may include channel impulse response data (CIR data) and image data.

[0153] The CIR data is used to characterize the channel characteristics of the wireless channel through which data is communicated between a receiving device and a transmitting device.

[0154] The image data is used to represent the environment image in which the receiving device and the sending device are located when performing data communication, and can represent environmental information.

[0155] The CIR data is obtained by parsing the receiving device. Optionally, the image data may be acquired by the receiving device and / or the transmitting device. For example, the receiving device is provided with an image acquisition device, which acquires environmental information about the receiving device's environment; and / or the transmitting device is provided with an image acquisition device, which acquires environmental information about the transmitting device's environment.

[0156] Optionally, the image acquisition device on the receiving device and the image acquisition device on the transmitting device can be a camera or a webcam. The camera can be a fisheye camera. Alternatively, the camera can be a thermal imaging camera or a color depth camera (Red-Green-Blue-Depth, RGB-D). An RGB-D camera is a visual sensor that can simultaneously capture color images (RGB images) and depth information.

[0157] Optionally, the image acquisition device on the receiving device and the image acquisition device on the transmitting device may also be a radar. The radar can scan and acquire high-precision three-dimensional point cloud data of environmental information, and can determine whether there are obstacles when the receiving and transmitting devices are communicating, as well as the geometric topological characteristics of the obstacles (such as surface curvature, occlusion angle, etc.).

[0158] Step 202: Perform recognition processing on the channel impulse response data and image data in the data group to be recognized according to the recognition model to obtain the communication status.

[0159] The communication state represents the communication state when the receiving device communicates with the sending device.

[0160] Exemplarily, the CIR data and image data in the data group to be identified are input into a trained recognition model. After the recognition model identifies and processes the CIR data and image data in the data group to be identified, the communication status between the receiving device and the sending device during data communication can be obtained.

[0161] The communication state represents the communication state between the receiving device and the transmitting device. During data communication, the receiving device and the transmitting device may be in line-of-sight (LOS) or non-line-of-sight (NLOS) communication.

[0162] It can be understood that by performing recognition processing on the CIR data and image data in the data group to be recognized according to the recognition model, it can be determined whether the receiving device and the transmitting device are in the LOS state or the NLOS state during data communication.

[0163] Specifically, the recognition model can identify the CIR data in the data group to be identified and determine the channel characteristics of the wireless channel when the receiving device and the transmitting device communicate. The recognition model can also identify the image data in the data group to be identified and determine the environmental information of the environment in which the receiving device and the transmitting device communicate. Combining the channel characteristics and environmental information, the communication state of the receiving device and the transmitting device at the time corresponding to the current data group to be identified can be determined.

[0164] Optionally, the recognition model may be constructed based on a convolutional neural network (CNN); and / or, the recognition model may be constructed based on a long short-term memory network (LSTM).

[0165] The communication status identification method provided in the embodiments of the present application obtains channel impulse response data and image data from the transmitter and receiver, then inputs the data set to be identified containing these two data sets into a trained recognition model for recognition processing, thereby determining the communication status of the transmitter and receiver during communication. By introducing image data, rather than solely relying on channel impulse response data for communication status identification, the accuracy of identifying the communication status of the transmitter and receiver during communication can be improved.

[0166] Figure 3 Schematic diagram of the process of identifying the communication status provided by this application Figure 2 ,like Figure 3 As shown, this embodiment Figure 2 Based on the embodiment, step 202 of the communication status identification method is described in detail. The method includes:

[0167] Step 301: Perform feature extraction processing on the channel impulse response data based on the recognition model to obtain channel impulse response features; and perform feature extraction processing on the image data based on the recognition model to obtain image features.

[0168] The image feature indicates whether the receiving device and the transmitting device are in the same environment when communicating.

[0169] As can be seen from the preceding examples, the recognition model can include multiple networks. Specifically, the recognition model can include a CNN-LSTM network, a hybrid deep learning network that combines a convolutional neural network (CNN) and a long short-term memory network (LSTM). The CNN-LSTM network can be used to process sequence data with spatial and temporal dimensions. By extracting features from CIR data based on the CNN-LSTM network in the recognition model, channel impulse response features can be obtained.

[0170] The recognition model may also include a CNN network. Feature extraction of image data based on the CNN network in the recognition model can yield image features. These features are used to characterize whether the receiving device and the transmitting device are in the same environment during data communication.

[0171] Step 302: Based on the recognition model, the communication status is determined according to the image features and the channel impulse response features.

[0172] Exemplarily, based on the recognition model, recognition processing can be performed according to the extracted image features and channel impulse response features, thereby determining the communication state of the receiving device and the transmitting device corresponding to the image features and channel impulse response features during communication.

[0173] Specifically, the recognition model performs identification processing based on the channel impulse response characteristics, determining the characteristics of the current wireless channel's time-domain response to the pulse signal. If the channel impulse response characteristics indicate a significant increase in the number of multipath signals and a dispersed energy distribution, the current communication state may be NLOS. If the channel impulse response characteristics indicate a decrease in the main path energy, the current communication state may be NLOS. If the channel impulse response characteristics indicate increased signal attenuation, the current communication state may be NLOS.

[0174] Specifically, the recognition model performs recognition processing based on image features to determine the current environment information of the receiving and transmitting devices. If the environment information indicates that the receiving and transmitting devices are in different environments, the current communication state may be NLOS. If the environment information indicates that there is an obstacle between the receiving and transmitting devices, the current communication state may be NLOS.

[0175] In combination with the above example, the recognition model determines the communication status of the receiving device and the transmitting device based on the recognition of the channel impulse response characteristics and image characteristics.

[0176] In the above embodiment, the recognition model performs feature extraction processing on different types of data to obtain image features and channel impulse response features, respectively. This can avoid the problem of low accuracy when identifying the communication status based on a single channel impulse response data. This is because in some special scenarios, such as when blocked by a wooden door or when the obstacle is a thin metal plate, the characteristics of the channel impulse response data in the LOS state are similar to those in the NLOS state. By introducing image data that can represent environmental information, the accuracy of communication status identification can be improved. Based on the two features, identification processing is performed separately to obtain the communication status of the receiving and transmitting devices during communication.

[0177] In one example, the image data includes first image data and second image data. The first image data is an image of the environment captured by the transmitting device when the receiving device communicates with the transmitting device; the second image data is an image of the environment captured by the receiving device when the receiving device communicates with the transmitting device.

[0178] Exemplarily, the receiving device and the transmitting device are each configured with an image capture device. The image capture device on the transmitting device is used to capture an image of the environment in which the transmitting device is located when the receiving device and the transmitting device are communicating, i.e., the first image data; the image capture device on the receiving device is used to capture an image of the environment in which the receiving device is located when the receiving device and the transmitting device are communicating, i.e., the second image data.

[0179] The image acquisition device may be a fisheye camera, and the parameters of the fisheye camera may be set to: a shooting resolution of 1920×1080 and a field of view of 170°.

[0180] In the above example, by installing image acquisition devices on both the receiving and transmitting devices to capture the first and second image data, respectively, information about the environments of the receiving and transmitting devices can be obtained, avoiding the drawback of incomplete environmental information represented by a single image. This further improves the accuracy of image feature recognition, and based on accurate image features, the accuracy of communication status recognition is improved.

[0181] Figure 4 Schematic diagram of the process of identifying the communication status provided by this application Figure 3 In one example, Figure 4 As shown, in combination with the above example, in step 301, feature extraction processing is performed on the image data based on the recognition model to obtain image features, which may specifically include the following steps:

[0182] Step 401: Perform feature extraction processing on the first image data based on the recognition model to obtain a first feature.

[0183] The first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device.

[0184] Exemplarily, a feature extraction process is performed on the first image data based on a CNN network in the recognition model to obtain a first feature, wherein the first feature is used to characterize environmental features in the first image data collected by the sending device when the receiving device communicates with the sending device.

[0185] Specifically, in the process of the CNN network performing feature extraction processing on the first image data, a quantized convolution block can be used for processing, and each convolution block is composed of a 3×3 convolution layer, a ReLU activation function, and a 2×2 maximum pooling operation. For example, the process of using the quantized convolution block for processing can be expressed by the following formula (1):

[0186]

[0187] In formula (1), represents the maximum pooling process, represents the ReLU activation function, represents convolution processing, represents the first image data input, Represents the learnable convolution kernel parameters.

[0188] Through the above operations, the first feature corresponding to the first image data can be obtained.

[0189] Step 402: Perform feature extraction processing on the second image data based on the recognition model to obtain a second feature.

[0190] The second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the sending device.

[0191] Exemplarily, feature extraction processing is performed on the second image data based on the CNN network in the recognition model to obtain a second feature, wherein the second feature is used to characterize environmental features in the second image data collected by the receiving device when the receiving device communicates with the transmitting device.

[0192] It should be noted that the specific process of the CNN network performing feature extraction processing on the second image data can be found in the description of step 402 above and will not be expanded here.

[0193] Step 403: Perform feature fusion processing on the first feature and the second feature to obtain image features.

[0194] Exemplarily, feature fusion processing is performed on the extracted first feature and the second feature to obtain image features.

[0195] Specifically, the first feature is a 16-dimensional vector extracted from the first image data, and the second feature is a 16-dimensional vector extracted from the second image feature. The first and second features are concatenated to obtain the image feature. It can be understood that the image feature is a 32-dimensional vector after concatenation.

[0196] Figure 5 is a schematic diagram of an exemplary image feature extraction process. Figure 5As shown in the figure, the first image data is processed through a 2DCONV convolution layer and a MAXPOOL maximum pooling layer, followed by a 2DCONV convolution layer and a MAXPOOL maximum pooling layer, to obtain the first feature. The second image data is processed through a 2DCONV convolution layer and a MAXPOOL maximum pooling layer, followed by a 2DCONV convolution layer and a MAXPOOL maximum pooling layer, to obtain the second feature. The first and second features are concatenated to obtain the image feature. "2DCONV" refers to the convolution layer used to process two-dimensional data.

[0197] In the above example, the convolutional neural network in the recognition model can perform feature extraction processing on the first and second image data, respectively, to obtain first and second features. The first and second features can respectively represent the environmental information of the environment in which the transmitting device is located and the environmental information of the environment in which the receiving device is located. The first and second features are then fused to obtain image features. This allows accurate extraction of environmental information of the environment in which the transmitting device is located and the environment in which the receiving device is located. This allows the image features to more accurately reflect whether the transmitting and receiving devices are in the same environment, thereby improving the accuracy of communication status recognition.

[0198] Figure 6 Schematic diagram of the process of identifying the communication status provided by this application Figure 4 In one example, Figure 6 As shown, in combination with the above example, in step 301, feature extraction processing is performed on the channel impulse response data based on the recognition model to obtain the channel impulse response feature, which may specifically include the following steps:

[0199] Step 601: Based on the recognition model, perform convolution pooling processing on the channel impulse response data to obtain local spatial features.

[0200] The local spatial features represent the data characteristics of the channel impulse response data.

[0201] For example, the CNN network in the CNN-LSTM network of the recognition model performs convolutional pooling on the channel impulse response data to obtain local spatial features. These local spatial features are used to characterize the data characteristics of the channel impulse response data. For example, local spatial features refer to features related to CIR data, such as peak values, valley values, and maximum tap index. The maximum tap index (MTI) is the time index corresponding to the path with the highest energy (or maximum amplitude) in a multipath channel.

[0202] Step 602: Based on the gating mechanism of the recognition model, feature extraction processing is performed on the local spatial features to obtain channel impulse response features.

[0203] Exemplarily, based on the LSTM network in the CNN-LSTM network in the recognition model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response features.

[0204] Specifically, the LSTM network has a gating mechanism, through which the hidden state in the LSTM network is updated. and memory units , thereby capturing the time-dependent characteristics of path attenuation. For example, updating the hidden state in the LSTM network

[0205] and memory units The process can be expressed by the following formula (2):

[0206]

[0207] In formula (1), Represents the amplitude of the CIR data at time step t, represents the hidden state of the previous time step, Represents the trainable parameters in the LSTM network.

[0208] Repeat the process until the final hidden state is obtained, and concatenate the final hidden state of the forward propagation and the final hidden state of the backward propagation to obtain the channel impulse response characteristics.

[0209] Figure 7 FIG. 1 is a schematic diagram of an exemplary channel impulse response feature extraction process. Figure 7 As shown in the figure, the channel impulse response data (CIR data) is processed through the "1DCONV" convolution layer and the "1DCONV" convolution layer to obtain local spatial features. The local spatial features are then extracted using the LSTM network to obtain the channel impulse response features. The "1DCONV" refers to the convolution layer used to process one-dimensional data.

[0210] In the above example, the convolutional neural network in the recognition model performs convolution pooling on the channel impulse response data to obtain local spatial features. The long short-term memory network in the recognition model then extracts these local spatial features to obtain the channel impulse response features. This method preserves the multi-scale contextual information of the propagation path, accurately identifying features related to the communication state based on the channel impulse response data, thereby improving the accuracy of communication state recognition.

[0211] In conjunction with the above example, in one example, in step 302, determining the communication state based on the recognition model, image features and channel impulse response features can be implemented as follows:

[0212] Based on the recognition model, the image features and channel impulse response features are fused to obtain fused features.

[0213] Exemplarily, based on the image features and channel impulse response features extracted in the aforementioned example, feature fusion processing is performed to obtain fused features.

[0214] In one example, the feature fusion process may include: directly concatenating the vector of the image feature and the vector of the channel impulse response feature to obtain a fused feature.

[0215] In one example, the feature fusion process may include: performing a weighted summation between vectors according to the weight values ​​of the image features and the weight values ​​of the channel impulse response features to obtain a fused feature.

[0216] In one example, feature fusion processing may include: mapping the vector of image features and the vector of channel impulse response features to the same semantic space, and then performing weighted summation between the vectors according to corresponding weight values ​​to obtain fused features.

[0217] Specifically, in one example, the process of obtaining fused features by feature fusion processing includes:

[0218] A multi-layer perceptron based on the recognition model determines a third feature corresponding to the image feature and a fourth feature corresponding to the channel impulse response feature based on the image feature and the channel impulse response feature. The third feature and the fourth feature belong to the same semantic space.

[0219] Exemplarily, the image features and channel impulse response features are mapped to a unified semantic space based on the Multilayer Perceptron (MLP) in the recognition model.

[0220] Specifically, the third feature corresponding to the image feature is determined based on the image feature through MLP. For example, the third feature can be determined by the following formula (3):

[0221] (3)

[0222] In formula (3), v represents the image feature, v' represents the third feature, represents the trainable weight parameters of the MLP, represents the initial bias of the MLP.

[0223] Specifically, the fourth feature corresponding to the channel impulse response feature is determined by MLP based on the channel impulse response feature. Exemplarily, the fourth feature can be determined by the following formula (4):

[0224] (4)

[0225] In formula (4), c represents the channel impulse response characteristic, c' represents the fourth characteristic, represents the trainable weight parameters of the MLP, Represents the initial bias of the MLP. Through MLP, feature vectors representing different semantics can be unified in the semantic space, laying the foundation for subsequent feature fusion.

[0226] Based on the modal attention mechanism of the recognition model, a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature are determined according to the third feature and the fourth feature.

[0227] Exemplarily, based on the modal attention mechanism in the recognition model, the modality-specific weight is dynamically calculated. Specifically, the first weight value corresponding to the third feature is determined based on the third feature. Exemplarily, the first weight value can be determined by the following formula (5):

[0228] (5)

[0229] In formula (5), represents the first weight value, represents the Sigmoid function, tanh(·) represents the hyperbolic tangent function, represents the 256-dimensional attention parameter vector, and v' represents the third feature.

[0230] Specifically, the second weight value corresponding to the fourth feature is determined according to the fourth feature. For example, the second weight value can be determined by the following formula (6):

[0231] (6)

[0232] In formula (6), represents the second weight value, represents the Sigmoid function, tanh(·) represents the hyperbolic tangent function, represents the 256-dimensional attention parameter vector, and c' represents the fourth feature.

[0233] A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

[0234] Exemplarily, a weighted sum of vectors is performed based on the third feature, the fourth feature, the first weight value corresponding to the third feature, and the second weight value corresponding to the fourth feature in the same semantic space to obtain a fused feature.

[0235] Specifically, the fusion feature can be determined by the following formula (7):

[0236] (7)

[0237] In formula (7), represents the fusion feature, v' represents the third feature, represents the first weight value, c' represents the fourth feature, Indicates the second weight value.

[0238] In the above example, fused features are obtained by mapping image features and channel impulse response features to the same semantic space, determining the weights corresponding to the features in the same semantic space, and performing a weighted summation. This method preserves the characteristics of multimodal data, unifies the semantics, and then fuses them, resulting in features based on multimodal data. Recognizing communication status based on the fused features of multimodal data, rather than the data features of a single data source, can improve the accuracy of communication status recognition.

[0239] The multi-layer perceptron based on the recognition model determines the communication status according to the fusion features.

[0240] Exemplarily, the fused features are input into a multi-layer perceptron, and recognition processing is performed based on the fused features based on the multi-layer perceptron, so as to output the communication state corresponding to the fused features.

[0241] Specifically, the communication state output by the multilayer perceptron can be the probability of the LOS state and the probability of the NLOS state. If the probability of the LOS state is greater than the probability of the NLOS state, the communication state can be determined to be the LOS state, otherwise, the communication state can be determined to be the NLOS state.

[0242] In the above embodiment, image features and channel impulse response features are semantically unified based on a recognition model, and then weighted fusion is performed to obtain a fused feature. The fused feature is then identified and processed using a multi-layer perceptron within the recognition model to obtain the communication state corresponding to the fused feature. This allows for communication state identification based on multimodal data, thereby avoiding the low accuracy associated with feature identification based solely on channel impulse response data and improving communication state identification accuracy.

[0243] Based on any of the foregoing embodiments, in the process of acquiring the data group to be identified, image data and channel impulse response data may be acquired separately.

[0244] Figure 8 Schematic diagram of the process of identifying the communication status provided by this application Figure 5 In one example, Figure 8 As shown, the process of obtaining the data group to be identified includes:

[0245] Step 801: In response to receiving a polling data packet, image data is acquired; and signal analysis processing is performed on the polling data packet to obtain original impulse response sequence data.

[0246] Exemplarily, in response to receiving the polling data packet, the receiving device may respectively obtain the image data and the channel impulse response data.

[0247] In response to receiving the polling data packet, the receiving device acquires the first image data collected by the transmitting device and collects the second image data of the environment in which the receiving device is located.

[0248] In response to receiving a polling data packet, the receiving device performs signal analysis on the polling data packet to obtain raw impulse response sequence data. Specifically, the polling data packet is a reference signal pre-designed by the transmitting device. This reference signal may be attenuated during channel transmission. Consequently, the signal received by the receiving device may vary and can be recorded as the actual signal. Analysis is performed based on the actual signal and the reference signal to obtain raw impulse response sequence data. Raw impulse response sequence data is used to characterize the channel characteristics of the wireless channel through which data is communicated between the receiving device and the transmitting device.

[0249] Step 802: Truncate the original impulse response sequence data to obtain channel impulse response data.

[0250] Exemplarily, the original impulse response sequence data is truncated to obtain the channel impulse response data corresponding to the original impulse response sequence data.

[0251] Specifically, data from 100 sampling points near a main peak in the original impulse response sequence data is extracted as the channel impulse response data. The main peak refers to the peak with the highest peak value in the original impulse response sequence data; the data from the 100 sampling points near the main peak refers to data extracted from the original impulse response sequence data, including the main peak, using a truncated window of length 100. For example, the data from the 100 sampling points near the main peak may include data from 50 sampling points to the left of the sampling point corresponding to the main peak, data from the sampling point corresponding to the main peak, and data from 49 sampling points to the right of the sampling point corresponding to the main peak.

[0252] Furthermore, the image data and the channel impulse response data have the same round identifier, wherein the round identifier is used to indicate the round of the polling data packet corresponding to the image data and the channel impulse response data.

[0253] Exemplarily, when a transmitting device sends a polling packet to a receiving device, it updates the round identifier. For example, after sending the kth polling packet, the round identifier (Round ID) is incremented to k+1. When the transmitting device captures the first image data, the first image data also has the round identifier k+1. This round identifier indicates the round in which the polling packet corresponding to the image data (first image data, second image data) and channel impulse response data was sent.

[0254] After receiving the kth polling data packet, the receiving end device synchronously increments the round identifier to k+1. When the receiving end device collects the second image data, the second image data also has the round identifier with a value of k+1.

[0255] When the receiving device receives the kth polling data packet and parses the kth polling data packet to obtain the original impulse response sequence data, the original impulse response sequence data also has the round identifier having a value of k+1. Furthermore, because the channel impulse response data is obtained by truncating the original impulse response sequence data, the channel impulse response data also has the round identifier having a value of k+1.

[0256] In the above example, by setting the round identifier, the risk of cumulative error inherent in traditional timestamp matching methods can be eliminated, thereby avoiding the data misalignment problem. In other words, it can ensure that the image data and the channel impulse response data are synchronized in time.

[0257] Step 803: Construct a data group to be identified based on the image data and the channel impulse response data.

[0258] Exemplarily, a data set to be identified is constructed by combining the first image data acquired by the transmitting device, the second image data acquired by the receiving device, and the truncated channel impulse response data. In the data set to be identified, the first image data, the second image data, and the channel impulse response data have the same round identifier. Therefore, the first image data and the second image data are images of the environments of the transmitting device and the receiving device, acquired synchronously with the channel impulse response data.

[0259] In the above-described embodiment, truncation of the original impulse response sequence data can reduce data volume, improve computational speed, and reduce computational load. By setting a round identifier, temporal synchronization of image data and channel impulse response data can be ensured. Based on the image data synchronized with the channel impulse response data, an accurate representation of the environment in which the transmitting and receiving devices were located at the time corresponding to the channel impulse response data can be made, thereby accurately determining the communication state of the transmitting and receiving devices at the current moment.

[0260] The communication status identification method provided in the embodiments of the present application obtains channel impulse response data and image data from the transmitter and receiver, then inputs the data to be identified, including these two data types, into a trained recognition model for processing to determine the communication status of the transmitter and receiver during communication. By introducing image data, the accuracy of identifying the communication status of the transmitter and receiver during communication is improved.

[0261] This application also provides a method for training a recognition model. Figure 9 A flow chart of the training method for the recognition model provided in this application, such as Figure 9 As shown, the method includes:

[0262] Step 901: Obtain a training data set.

[0263] The training data set includes at least one data group, which includes channel impulse response data and image data; the channel impulse response data represents the channel characteristics when the receiving device communicates with the sending device, and the image data represents the environmental information when the receiving device communicates with the sending device.

[0264] For example, the method provided in this embodiment can be applied to a host computer device, wherein the host computer device is communicatively connected with a receiving device and a sending device respectively.

[0265] The host device obtains a training data set. The training data set includes at least one data set for model training. Each data set includes channel impulse response data (CIR data) and image data.

[0266] CIR data characterizes the characteristics of the wireless channel through which data is communicated between a receiving device and a transmitting device. CIR data is analyzed by the receiving device. Therefore, when acquiring a training dataset, the host device obtains CIR data from the receiving device.

[0267] Image data is used to represent the image of the environment in which the receiving device and the transmitting device are communicating data, and can represent environmental information. Image data can be collected by the receiving device and / or the transmitting device. Therefore, when obtaining a training dataset, the host device can obtain image data from the transmitting device and / or from the receiving device.

[0268] Step 902: Train the initial model based on the channel impulse response data and image data in the training data set to obtain a recognition model.

[0269] The identification model is used to identify the communication state when the receiving device communicates with the sending device.

[0270] Exemplarily, a training data set is input into an initial model, and the initial model is iteratively trained based on the channel impulse response data and image data in the training data set; after iterative training of the initial model, an identification model for identifying the communication status when a receiving device communicates with a transmitting device is obtained.

[0271] The initial model may be constructed based on a convolutional neural network (CNN); and / or the initial model may be constructed based on a long short-term memory network (LSTM).

[0272] The recognition model training method provided in an embodiment of the present application obtains channel impulse response data from a receiving device and image data from the receiving device and / or the transmitting device to construct a training dataset. An initial model is trained based on the training dataset to obtain a recognition model for identifying communication status. By introducing image data to train the initial model, the trained recognition model can be combined with the channel impulse response data and image data to jointly identify the communication status, thereby improving the recognition model's accuracy in identifying the communication status.

[0273] Based on the above embodiment, in one example, the process of training the initial model includes:

[0274] Based on the initial model, feature extraction processing is performed on the channel impulse response data to obtain channel impulse response features; and based on the initial model, feature extraction processing is performed on the image data to obtain image features.

[0275] The image feature indicates whether the receiving device and the transmitting device are in the same environment when communicating.

[0276] Exemplarily, feature extraction is performed on the CIR data based on the initial model to obtain the channel impulse response feature. Specifically, feature extraction is performed on the CIR data based on the CNN-LSTM network in the initial model to obtain the channel impulse response feature.

[0277] Exemplarily, feature extraction is performed on image data based on the initial model to obtain image features. The image features are used to indicate whether the receiving device and the transmitting device are in the same environment during data communication. Specifically, feature extraction is performed on image data based on a CNN network in the initial model to obtain image features.

[0278] It should be noted that the above specific process can refer to the description of step 301 in the above embodiment, which will not be expanded here.

[0279] According to the image features and channel impulse response features, the predicted communication state is obtained.

[0280] Exemplarily, based on the initial model, feature recognition is performed according to the extracted image features and channel impulse response features, thereby obtaining the predicted communication state predicted by the initial model.

[0281] It should be noted that the above specific process can refer to the description of step 302 in the above embodiment, which will not be expanded here.

[0282] According to the predicted communication state and the actual communication state corresponding to the data group, the initial model is trained to obtain the recognition model.

[0283] For example, in the training dataset, each data group corresponds to a real communication state. The real communication state can be understood as the label value corresponding to the data group, which represents the actual real communication state when the transmitting device and the receiving device communicate under the corresponding data group.

[0284] It can be understood that if the predicted communication state predicted by the initial model is consistent with the actual communication state, it means that the initial model can accurately predict the communication state; otherwise, the initial model cannot accurately predict the communication state and iterative training and parameter adjustment are required.

[0285] In the above example, iterative training is performed using the actual communication states corresponding to the data groups in the training dataset and the predicted communication states obtained by the initial model. Training is completed when the number of iterative training iterations reaches a preset number. The trained recognition model can accurately extract features from both image data and channel impulse response data, and accurately predict the communication state based on these two features.

[0286] Specifically, in one example, the process of training the initial model by combining the predicted communication state and the actual communication state may include:

[0287] The initial model is iteratively trained based on the predicted communication state, the actual communication state corresponding to the data set, and a preset composite loss function. When the number of iterative training reaches a preset number, a recognition model is obtained.

[0288] The composite loss function includes: fusion classification result deviation term, image classification result deviation term, channel impulse response classification result deviation term, divergence minimization term and cosine similarity maximization term.

[0289] Exemplarily, the initial model is iteratively trained based on the predicted communication state and the actual communication state corresponding to the data set. When the number of iterative training reaches a preset number, the training is completed, and the corresponding initial model is the trained recognition model.

[0290] During iterative training, a composite loss function is designed to balance different tasks, constraints, or optimization objectives during initial model training, enabling the trained recognition model to accurately identify communication status based on image and CIR data.

[0291] Exemplarily, the composite loss function may include: a fusion classification result deviation term, an image classification result deviation term, a channel impulse response classification result deviation term, a divergence minimization term, and a cosine similarity maximization term.

[0292] Among them, the meanings of the components of the composite loss function are as follows:

[0293] The deviation term of the fusion classification result represents the accuracy of the prediction result obtained by the initial model based on the image data and channel impulse response data. It is an indicator calculated based on the fusion classification result and the true label (the true communication state corresponding to the data group). It can avoid the sample sparsity problem in the training data set where the true communication state corresponding to the data group is the NLOS state.

[0294] The image classification result bias term characterizes the accuracy of the predictions obtained by the initial model based on image data. The channel impulse response classification result bias term characterizes the accuracy of the predictions obtained by the initial model based on channel impulse response data. These two terms impose independent classification constraints on the image and CIR modalities to prevent feature degradation in either modality.

[0295] The divergence minimization term characterizes the consistency between the prediction results of multiple data and the prediction results of a single data. It can be used to constrain the consistency of the fusion prediction distribution and the single-modal prediction distribution.

[0296] The cosine similarity maximization term characterizes the similarity between the features of the image data and the features of the channel impulse response data. It can be used to maximize the cosine similarity between visual features (image features) and CIR features (channel impulse response features).

[0297] For example, the expression of the composite loss function can be expressed by the following formula (8):

[0298]

[0299] In formula (8), represents the composite loss function, Represents the deviation term of the fusion classification result, Represents the image classification result deviation term, represents the deviation term of the channel impulse response classification result, represents the divergence minimization term, represents the cosine similarity maximization term.

[0300] Furthermore, the expression of the fusion classification result deviation term can be expressed by the following formula (9):

[0301] (9)

[0302] In formula (9), represents the automatically calculated category weights, represents the initial model's predicted probability for the true category. Compared to the standard cross-entropy calculation method, the above fusion classification result bias term imposes a stronger penalty on misclassified samples, enabling faster training convergence and improving the training efficiency of the initial model.

[0303] Furthermore, the expression of the image classification result deviation term can be expressed by the following formula (10):

[0304] (10)

[0305] In formula (10), represents the category weight calculated according to the distribution of the training data set, represents the label value (the actual communication status corresponding to the data group in the training dataset), Represents the unimodal prediction result, that is, the predicted communication state obtained by the initial model based on the image data. The image data can be processed separately by a lightweight multi-layer perceptron to obtain a predicted communication state based on the image data.

[0306] Furthermore, the expression of the channel impulse response classification result deviation term can be expressed by the following formula (11):

[0307] (11)

[0308] In formula (11), represents the category weight calculated according to the distribution of the training data set, represents the label value (the actual communication status corresponding to the data group in the training dataset), Represents the single-mode prediction result, that is, the predicted communication state obtained by the initial model based on the channel impulse response data. The channel impulse response data can be processed separately by a lightweight multi-layer perceptron to obtain a predicted communication state based on the channel impulse response data.

[0309] Furthermore, the expression of the divergence minimization term can be expressed by the following formula (12):

[0310] (12)

[0311] In formula (12), represents the KL divergence calculation, represents the predicted probability distribution of the initial model for the fusion features, represents the predicted probability distribution of the initial model for image features, Represents the predicted probability distribution of the initial model for the channel impulse response characteristics. The image data can be processed separately by a lightweight multi-layer perceptron to obtain the predicted probability distribution of the initial model for the image features; The channel impulse response data can be processed separately by a lightweight multi-layer perceptron to obtain the predicted probability distribution of the channel impulse response characteristics of the initial model.

[0312] Furthermore, the expression of the cosine similarity maximization term can be expressed by the following formula (13):

[0313] (13)

[0314] In formula (12), Indicates the number of data groups in the training data set, represents the image features corresponding to the i-th image data in multiple data groups, represents the channel impulse response feature corresponding to the i-th channel impulse response data in multiple data groups, Represents a modulo operation on a vector.

[0315] In the above example, a composite loss function was designed to optimize the efficiency of the initial model training process. By combining different terms in the composite loss function, different training objectives can be addressed during the initial model training process. This improves the efficiency of initial model training while also increasing the accuracy of the trained recognition model in identifying communication states during application.

[0316] In one example, the image data includes first image data and second image data. The first image data is an image of the environment captured by the transmitting device when the receiving device communicates with the transmitting device; the second image data is an image of the environment captured by the receiving device when the receiving device communicates with the transmitting device.

[0317] The first image data is obtained by the host device from the sending device; the second image data is obtained by the host device from the receiving device.

[0318] It should be noted that, regarding the explanation of the first image data and the second image data and the corresponding technical effects, please refer to the explanation of the aforementioned embodiment, which will not be elaborated here.

[0319] In one example, performing feature extraction processing on image data based on the initial model to obtain image features may specifically include the following steps:

[0320] The first image data is subjected to feature extraction processing based on the initial model to obtain a first feature, wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device.

[0321] Exemplarily, during the training process of the initial model, the initial model on the host computer device performs feature extraction based on the first image data of the data group in the training data set to obtain the first feature.

[0322] Specifically, feature extraction processing is performed on the first image data based on the CNN network in the initial model to obtain the first feature. Further, quantized convolution blocks can be used for processing, each of which is composed of a 3×3 convolution layer, a ReLU activation function, and a 2×2 maximum pooling operation.

[0323] It should be noted that the above specific process can refer to the description of step 401 in the above embodiment, which will not be expanded here.

[0324] The second image data is subjected to feature extraction processing based on the initial model to obtain a second feature, wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the transmitting device.

[0325] Exemplarily, during the training process of the initial model, the initial model on the host computer device performs feature extraction based on the second image data of the data group in the training data set to obtain the second feature.

[0326] Specifically, feature extraction processing is performed on the second image data based on the CNN network in the initial model to obtain the second feature.

[0327] It should be noted that the above specific process can refer to the description of step 401 and step 402 in the above embodiment, which will not be expanded here.

[0328] The first feature and the second feature are subjected to feature fusion processing to obtain image features.

[0329] Exemplarily, based on the initial model, feature fusion processing is performed on the first feature and the second feature to obtain image features.

[0330] Specifically, based on the initial model, the vectors corresponding to the first feature and the second feature are concatenated to obtain the image feature.

[0331] It should be noted that the above specific process can refer to the description of step 403 in the above embodiment, which will not be expanded here.

[0332] In the above example, during the initial model training process, the initial model performs feature extraction on the first and second image data, ensuring that the initial model can accurately extract environmental information about the environment in which the transmitting device resides, as well as the environment in which the receiving device resides. Image features are then fused to preserve the environmental information of both the transmitting and receiving devices, thereby characterizing whether the transmitting and receiving devices are in the same environment during communication. This allows the trained recognition model to accurately determine whether the transmitting and receiving devices are in the same environment during communication, thereby improving the accuracy of communication status recognition.

[0333] In one example, performing feature extraction processing on the channel impulse response data based on the initial model to obtain the channel impulse response features may specifically include the following steps:

[0334] Based on the initial model, convolution and pooling are performed on the channel impulse response data to obtain local spatial features, which represent the data characteristics of the channel impulse response data.

[0335] Exemplarily, during the training of the initial model, the initial model on the host device performs convolution pooling processing based on the channel impulse response data of the data group in the training data set to obtain local spatial features.

[0336] Specifically, based on the CNN network in the CNN-LSTM network in the initial model, convolution pooling is performed on the channel impulse response data to obtain local spatial features.

[0337] It should be noted that the above specific process can refer to the description of step 601 in the above embodiment, which will not be expanded here.

[0338] Based on the gating mechanism of the initial model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response features.

[0339] Exemplarily, during the training of the initial model, the initial model on the host device performs feature extraction processing on the local spatial features to obtain channel impulse response features.

[0340] Specifically, based on the LSTM network in the CNN-LSTM network in the initial model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response features.

[0341] It should be noted that the above specific process can refer to the description of step 602 in the above embodiment, which will not be expanded here.

[0342] In the above example, during initial model training, the initial model performs convolutional pooling on the channel impulse response data to obtain local spatial features. Feature extraction is then performed on these local spatial features to obtain channel impulse response features. This enables the trained recognition model to accurately identify features related to the communication state from the channel impulse response data, thereby improving the accuracy of communication state recognition.

[0343] In one example, obtaining a predicted communication state based on image features and channel impulse response features may specifically include the following steps:

[0344] Based on the initial model, feature fusion processing is performed on the image features and the channel impulse response features to obtain fused features.

[0345] Exemplarily, based on the initial model, feature fusion processing is performed on the image features and channel impulse response features extracted in the aforementioned example to obtain fused features.

[0346] Specifically, in an example, the following steps may be included:

[0347] The multilayer perceptron based on the initial model determines a third feature corresponding to the image feature and a fourth feature corresponding to the channel impulse response feature based on the image feature and the channel impulse response feature. The third feature and the fourth feature belong to the same semantic space.

[0348] Exemplarily, based on the multi-layer perceptron in the initial model, image features and channel impulse response features are mapped to a unified semantic space.

[0349] It should be noted that the above specific process can refer to the description of mapping image features and channel impulse response features to the same semantic space based on the multi-layer perceptron in the recognition model in the aforementioned embodiment, which will not be elaborated here.

[0350] Based on the modal attention mechanism of the initial model, a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature are determined according to the third feature and the fourth feature.

[0351] Exemplarily, modality-specific weights are dynamically calculated based on the modality attention mechanism in the initial model.

[0352] It should be noted that the above specific process can refer to the description of dynamically calculating the modality-specific weight based on the modal attention mechanism in the recognition model in the aforementioned embodiment, which will not be expanded here.

[0353] A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

[0354] Exemplarily, based on the initial model, a weighted summation of vectors is performed according to the third feature, the fourth feature, the first weight value corresponding to the third feature, and the second weight value corresponding to the fourth feature in the same semantic space to obtain a fused feature.

[0355] It should be noted that the above specific process can refer to the description of the weighted summation of vectors by the recognition model in the aforementioned embodiment to obtain the fusion feature, which will not be expanded here.

[0356] In the above example, during initial model training, the initial model maps image features and channel impulse response features to the same semantic space, generating fused features. This allows the initial model to better identify the fused features, thereby improving communication status recognition accuracy. Furthermore, by leveraging the resulting multimodal data features, rather than relying solely on channel impulse response data, communication status recognition accuracy can also be improved.

[0357] The multi-layer perceptron based on the initial model determines the predicted communication status according to the fusion features.

[0358] Exemplarily, the fused features are input into the multi-layer perceptron of the initial model, and recognition processing is performed based on the fused features based on the multi-layer perceptron, so as to output the communication state corresponding to the fused features.

[0359] It should be noted that the above specific process can refer to the description of the multi-layer perceptron of the recognition model in the aforementioned embodiment for determining the communication status, which will not be expanded here.

[0360] In the above embodiment, during the initial model training process, by training the initial model's multi-layer perceptron, the trained recognition model can improve its accuracy in recognizing communication status. Furthermore, training the initial model based on multimodal data, rather than single-modal data, further improves the trained recognition model's accuracy in recognizing communication status.

[0361] The recognition model training method provided in an embodiment of the present application obtains channel impulse response data from a receiving device and image data from the receiving device and / or the transmitting device to construct a training data set. An initial model is trained based on the training data set to thereby obtain a recognition model for identifying communication status. By introducing image data to train the initial model, the trained recognition model can be combined with the channel impulse response data and image data to jointly identify the communication status, thereby improving the recognition model's accuracy in identifying the communication status.

[0362] Figure 10 This is a schematic diagram of the structure of the communication status identification device provided by this application, such as Figure 10 As shown, the communication status identification device 100 provided in this embodiment includes:

[0363] A first acquisition module 1001 is configured to acquire a data group to be identified, the data group to be identified including channel impulse response data and image data; the channel impulse response data represents channel characteristics when a receiving device communicates with a transmitting device, and the image data represents environmental information when the receiving device communicates with the transmitting device;

[0364] The processing module 1002 is used to identify and process the channel impulse response data and image data in the data group to be identified according to the identification model to obtain a communication state; wherein the communication state represents the communication state when the receiving device communicates with the transmitting device.

[0365] In a possible implementation, the channel impulse response data and the image data in the data group to be identified are identified and processed according to the identification model to obtain the communication status. The processing module 1002 is configured to:

[0366] Performing feature extraction processing on the channel impulse response data based on the recognition model to obtain channel impulse response features; and performing feature extraction processing on the image data based on the recognition model to obtain image features, wherein the image features represent whether the receiving device and the transmitting device are in the same environment when communicating;

[0367] Based on the recognition model, the communication status is determined according to the image features and channel impulse response characteristics.

[0368] In one possible implementation, the image data includes first image data and second image data; wherein the first image data is an environmental image captured by the sending device when the receiving device communicates with the sending device; and the second image data is an environmental image captured by the receiving device when the receiving device communicates with the sending device.

[0369] In one possible implementation, feature extraction processing is performed on image data based on a recognition model to obtain image features. The processing module 1002 is configured to:

[0370] Performing feature extraction processing on the first image data based on the recognition model to obtain a first feature; wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device;

[0371] Performing feature extraction processing on the second image data based on the recognition model to obtain a second feature; wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the transmitting device;

[0372] The first feature and the second feature are subjected to feature fusion processing to obtain image features.

[0373] In a possible implementation, feature extraction processing is performed on the channel impulse response data based on the recognition model to obtain channel impulse response features. The processing module 1002 is configured to:

[0374] Based on the recognition model, the channel impulse response data is subjected to convolution and pooling processing to obtain local spatial features; wherein the local spatial features represent the data characteristics of the channel impulse response data;

[0375] Based on the gating mechanism of the recognition model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response characteristics.

[0376] In one possible implementation, based on the recognition model, the communication state is determined according to the image features and the channel impulse response features. The processing module 1002 is configured to:

[0377] Based on the recognition model, the image features and channel impulse response features are fused to obtain fused features;

[0378] The multi-layer perceptron based on the recognition model determines the communication status according to the fusion features.

[0379] In a possible implementation, based on the recognition model, feature fusion processing is performed on the image features and the channel impulse response features to obtain fused features. The processing module 1002 is configured to:

[0380] A multi-layer perceptron based on the recognition model determines, based on the image features and the channel impulse response features, a third feature corresponding to the image features and a fourth feature corresponding to the channel impulse response features; wherein the third feature and the fourth feature belong to the same semantic space;

[0381] Based on the modal attention mechanism of the recognition model, according to the third feature and the fourth feature, a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature are determined respectively;

[0382] A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

[0383] In a possible implementation, to obtain a data group to be identified, the first obtaining module 1001 is configured to:

[0384] In response to receiving the polling data packet, image data is acquired; and signal analysis processing is performed on the polling data packet to obtain original impulse response sequence data;

[0385] The original impulse response sequence data is truncated to obtain channel impulse response data;

[0386] A data group to be identified is constructed based on the image data and the channel impulse response data.

[0387] In a possible implementation, the image data and the channel impulse response data have the same round identifier; wherein the round identifier is used to indicate the rounds of the polling data packets corresponding to the image data and the channel impulse response data.

[0388] The communication status identification device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0389] Figure 11 A schematic diagram of the structure of the training device for the recognition model provided in this application, such as Figure 11 As shown, the recognition model training device 110 provided in this embodiment includes:

[0390] The second acquisition module 1101 is configured to acquire a training data set; wherein the training data set includes at least one data group, the data group including channel impulse response data and image data; the channel impulse response data represents channel characteristics when the receiving device communicates with the transmitting device, and the image data represents environmental information when the receiving device communicates with the transmitting device;

[0391] A training module 1102 is configured to train the initial model based on the channel impulse response data and image data in the training data set to obtain a recognition model;

[0392] The identification model is used to identify the communication state when the receiving device communicates with the sending device.

[0393] In a possible implementation, the initial model is trained based on the channel impulse response data and image data in the training data set to obtain a recognition model. The training module 1102 is configured to:

[0394] Performing feature extraction processing on the channel impulse response data based on the initial model to obtain channel impulse response features; and performing feature extraction processing on the image data based on the initial model to obtain image features, wherein the image features represent whether the receiving device and the transmitting device are in the same environment when communicating;

[0395] According to the image features and the channel impulse response features, the predicted communication state is obtained;

[0396] According to the predicted communication state and the actual communication state corresponding to the data group, the initial model is trained to obtain the recognition model.

[0397] In one possible implementation, the image data includes first image data and second image data; wherein the first image data is an environmental image captured by the sending device when the receiving device communicates with the sending device; and the second image data is an environmental image captured by the receiving device when the receiving device communicates with the sending device.

[0398] In one possible implementation, feature extraction processing is performed on the image data based on the initial model to obtain image features. The training module 1102 is used to:

[0399] Performing feature extraction processing on the first image data based on the initial model to obtain a first feature; wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device;

[0400] Performing feature extraction processing on the second image data based on the initial model to obtain a second feature; wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the transmitting device;

[0401] The first feature and the second feature are subjected to feature fusion processing to obtain image features.

[0402] In a possible implementation, feature extraction processing is performed on the channel impulse response data based on the initial model to obtain channel impulse response features. The training module 1102 is configured to:

[0403] Based on the initial model, convolution pooling is performed on the channel impulse response data to obtain local spatial features; wherein the local spatial features represent the data characteristics of the channel impulse response data;

[0404] Based on the gating mechanism of the initial model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response features.

[0405] In one possible implementation, the predicted communication state is obtained based on the image features and the channel impulse response features. The training module 1102 is configured to:

[0406] Based on the initial model, the image features and channel impulse response features are fused to obtain fused features;

[0407] The multi-layer perceptron based on the initial model determines the predicted communication status according to the fusion features.

[0408] In a possible implementation, based on the initial model, feature fusion processing is performed on the image features and the channel impulse response features to obtain fused features. The training module 1102 is used to:

[0409] A multilayer perceptron based on the initial model determines, according to the image features and the channel impulse response features, a third feature corresponding to the image features and a fourth feature corresponding to the channel impulse response features; wherein the third feature and the fourth feature belong to the same semantic space;

[0410] Based on the modal attention mechanism of the initial model, according to the third feature and the fourth feature, a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature are determined respectively;

[0411] A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

[0412] In one possible implementation, the initial model is trained based on the predicted communication state and the actual communication state corresponding to the data group to obtain a recognition model. The training module 1102 is used to:

[0413] Perform iterative training on the initial model based on the predicted communication state, the actual communication state corresponding to the data group, and a preset composite loss function;

[0414] Among them, when the number of iterative training reaches a preset number, the recognition model is obtained; the composite loss function includes: fusion classification result deviation term, image classification result deviation term, channel impulse response classification result deviation term, divergence minimization term and cosine similarity maximization term.

[0415] The training device for the recognition model provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.

[0416] Figure 12 This is a schematic diagram of the structure of the receiving device provided in this application. Figure 12 As shown, the receiving device 120 provided in this embodiment includes: at least one processor 1201 and a memory 1202. Optionally, the receiving device 120 also includes a communication component 1203. The processor 1201, the memory 1202, and the communication component 1203 are connected via a bus 1204.

[0417] During the specific implementation process, at least one processor 1201 executes the computer-executable instructions stored in the memory 1202, so that the at least one processor 1201 performs the above method.

[0418] The specific implementation process of the processor 1201 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0419] Figure 13 This is a schematic diagram of the structure of the host computer equipment provided in this application. Figure 13 As shown, the host computer device 130 provided in this embodiment includes: at least one processor 1301 and a memory 1302. Optionally, the host computer device 130 further includes a communication component 1303. The processor 1301, the memory 1302 and the communication component 1303 are connected via a bus 1304.

[0420] During the specific implementation process, at least one processor 1301 executes the computer-executable instructions stored in the memory 1302, so that the at least one processor 1301 performs the above method.

[0421] The specific implementation process of the processor 1301 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0422] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0423] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0424] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0425] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0426] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0427] The readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0428] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0429] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.

[0430] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0431] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0432] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0433] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0434] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.

Claims

1. A method for identifying a communication status, characterized in that: The method comprises: Acquire a data group to be identified, the data group to be identified including channel impulse response data and image data; the channel impulse response data represents channel characteristics when a receiving end device communicates with a transmitting end device, and the image data represents environmental information when the receiving end device communicates with the transmitting end device; The channel impulse response data and the image data in the data group to be identified are identified and processed according to the identification model to obtain a communication state; wherein the communication state represents the communication state when the receiving device communicates with the sending device.

2. The method according to claim 1, characterized in that Performing recognition processing on the channel impulse response data and the image data in the data group to be recognized according to the recognition model to obtain the communication state includes: performing feature extraction processing on the channel impulse response data based on the recognition model to obtain channel impulse response features; and performing feature extraction processing on the image data based on the recognition model to obtain image features, wherein the image features represent whether the receiving device and the transmitting device are in the same environment when communicating; Based on the recognition model, the communication state is determined according to the image features and the channel impulse response features.

3. The method according to claim 2, characterized in that The image data includes first image data and second image data; wherein, the first image data is the environmental image collected by the sending device when the receiving device communicates with the sending device; the second image data is the environmental image collected by the receiving device when the receiving device communicates with the sending device.

4. The method according to claim 3, characterized in that Performing feature extraction processing on the image data based on the recognition model to obtain image features includes: Performing feature extraction processing on the first image data based on the recognition model to obtain a first feature; wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device; Performing feature extraction processing on the second image data based on the recognition model to obtain a second feature; wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the sending device; Perform feature fusion processing on the first feature and the second feature to obtain the image feature.

5. The method according to claim 2, characterized in that Performing feature extraction processing on the channel impulse response data based on the recognition model to obtain channel impulse response features includes: Based on the recognition model, convolution pooling is performed on the channel impulse response data to obtain local spatial features; wherein the local spatial features represent data features of the channel impulse response data; Based on the gating mechanism of the recognition model, feature extraction processing is performed on the local spatial features to obtain the channel impulse response features.

6. The method according to claim 2, characterized in that Determining a communication state based on the recognition model and according to the image features and the channel impulse response features includes: Based on the recognition model, performing feature fusion processing on the image features and the channel impulse response features to obtain fused features; The communication state is determined based on the multi-layer perceptron of the recognition model and the fusion feature.

7. The method according to claim 6, characterized in that Based on the recognition model, the image features and the channel impulse response features are subjected to feature fusion processing to obtain fused features, including: A multilayer perceptron based on the recognition model determines, based on the image feature and the channel impulse response feature, a third feature corresponding to the image feature and a fourth feature of the channel impulse response feature; wherein the third feature and the fourth feature belong to the same semantic space; Based on the modal attention mechanism of the recognition model, determining a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature according to the third feature and the fourth feature respectively; A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

8. The method according to any one of claims 1 to 6, characterized in that Get the data group to be identified, including: In response to receiving a polling data packet, image data is acquired; and signal analysis processing is performed on the polling data packet to obtain original impulse response sequence data; performing truncation processing on the original impulse response sequence data to obtain the channel impulse response data; A data group to be identified is constructed based on the image data and the channel impulse response data.

9. The method according to any one of claims 1 to 6, characterized in that The image data and the channel impulse response data have the same round identifier; wherein the round identifier is used to represent the rounds of the polling data packets corresponding to the image data and the channel impulse response data.

10. A method for training a recognition model, characterized in that: The method comprises: Acquire a training data set; wherein the training data set includes at least one data group, the data group including channel impulse response data and image data; the channel impulse response data represents channel characteristics when the receiving device communicates with the transmitting device, and the image data represents environmental information when the receiving device communicates with the transmitting device; Training the initial model based on the channel impulse response data and image data in the training data set to obtain a recognition model; The identification model is used to identify the communication state when the receiving device communicates with the sending device.

11. The method according to claim 10, characterized in that Training the initial model based on the channel impulse response data and image data in the training data set to obtain a recognition model includes: performing feature extraction processing on the channel impulse response data based on the initial model to obtain channel impulse response features; and performing feature extraction processing on the image data based on the initial model to obtain image features, wherein the image features represent whether the receiving device and the transmitting device are in the same environment when communicating; Obtaining a predicted communication state according to the image features and the channel impulse response features; The initial model is trained according to the predicted communication state and the actual communication state corresponding to the data group to obtain the recognition model.

12. The method according to claim 11, characterized in that The image data includes first image data and second image data; wherein, the first image data is the environmental image collected by the sending device when the receiving device communicates with the sending device; the second image data is the environmental image collected by the receiving device when the receiving device communicates with the sending device.

13. The method according to claim 12, characterized in that Performing feature extraction processing on the image data based on the initial model to obtain image features includes: Performing feature extraction processing on the first image data based on the initial model to obtain a first feature; wherein the first feature represents an environmental feature of the environment in which the sending device is located when the receiving device communicates with the sending device; Performing feature extraction processing on the second image data based on the initial model to obtain a second feature; wherein the second feature represents an environmental feature of the environment in which the receiving device is located when the receiving device communicates with the transmitting device; Perform feature fusion processing on the first feature and the second feature to obtain the image feature.

14. The method according to claim 11, characterized in that Performing feature extraction processing on the channel impulse response data based on the initial model to obtain channel impulse response features includes: Based on the initial model, convolution pooling is performed on the channel impulse response data to obtain local spatial features; wherein the local spatial features represent data features of the channel impulse response data; The local spatial features are subjected to feature extraction processing based on the gating mechanism of the initial model to obtain the channel impulse response features.

15. The method according to claim 11, characterized in that Obtaining a predicted communication state according to the image feature and the channel impulse response feature, including: Based on the initial model, performing feature fusion processing on the image features and the channel impulse response features to obtain fused features; The predicted communication state is determined based on the multilayer perceptron of the initial model and according to the fusion features.

16. The method according to claim 15, characterized in that Based on the initial model, the image features and the channel impulse response features are subjected to feature fusion processing to obtain fused features, including: A multilayer perceptron based on the initial model determines, according to the image feature and the channel impulse response feature, a third feature corresponding to the image feature and a fourth feature of the channel impulse response feature; wherein the third feature and the fourth feature belong to the same semantic space; Based on the modal attention mechanism of the initial model, determining a first weight value corresponding to the third feature and a second weight value corresponding to the fourth feature according to the third feature and the fourth feature respectively; A weighted sum is performed according to the third feature, the first weight value, the fourth feature, and the second weight value to obtain a fusion feature.

17. The method according to any one of claims 11 to 16, characterized in that Training the initial model according to the predicted communication state and the actual communication state corresponding to the data group to obtain the recognition model includes: Performing iterative training on the initial model according to the predicted communication state, the actual communication state corresponding to the data group, and a preset composite loss function; Among them, when the number of iterative training executions reaches a preset number, the recognition model is obtained; the composite loss function includes: a fusion classification result deviation term, an image classification result deviation term, a channel impulse response classification result deviation term, a divergence minimization term, and a cosine similarity maximization term.

18. A communication status identification device, characterized in that: include: a first acquisition module, configured to acquire a data group to be identified, wherein the data group to be identified includes channel impulse response data and image data; the channel impulse response data represents channel characteristics when a receiving end device communicates with a transmitting end device, and the image data represents environmental information when the receiving end device communicates with the transmitting end device; The processing module is used to identify and process the channel impulse response data and image data in the data group to be identified according to the identification model to obtain a communication state; wherein the communication state represents the communication state when the receiving device communicates with the sending device.

19. A training device for a recognition model, characterized in that: include: a second acquisition module configured to acquire a training data set; wherein the training data set includes at least one data set, the data set including channel impulse response data and image data; the channel impulse response data represents channel characteristics when a receiving device communicates with a transmitting device, and the image data represents environmental information when the receiving device communicates with the transmitting device; A training module, configured to train an initial model based on the channel impulse response data and image data in the training data set to obtain a recognition model; The identification model is used to identify the communication state when the receiving device communicates with the sending device.

20. A receiving device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 9.

21. A host computer device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 10 to 17.

22. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 17 when executed by a processor.

23. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 17 when being executed by a processor.