Perception processing method and apparatus, and device and medium
By using AI models to process multiple input data in mobile communication systems, the problem of fusion perception is solved, and the perception resolution and system performance are improved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2025-05-09
- Publication Date
- 2026-04-23
AI Technical Summary
How can we achieve AI-based fusion sensing in future mobile communication systems to improve sensing resolution and overall performance?
By using a first AI model at the first node to process at least two different types of input data, second data related to the perception business is generated, thereby achieving AI-based fusion perception.
It improved the perception resolution, enhanced the overall performance of the communication and perception systems, and enabled more refined perception services.
Smart Images

Figure CN2025093852_23042026_PF_FP_ABST
Abstract
Description
Methods, devices, equipment and media for sensing and processing
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410597891.7, filed in China on May 14, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application belongs to the field of artificial intelligence technology, specifically relating to a method, apparatus, device, and medium for perception processing. Background Technology
[0004] Future mobile communication systems, such as Beyond 5th Generation (B5G) or 6th Generation (6G) systems, will possess sensing capabilities in addition to communication capabilities. One or more devices with sensing capabilities can perceive information such as the location, distance, and speed of target objects through the transmission and reception of wireless signals, or perform detection, tracking, identification, and imaging of target objects, events, or environments. Currently, how to achieve integrated sensing based on Artificial Intelligence (AI) is a problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a method, apparatus, device, and medium for perception processing, which can solve the problem of how to perform AI-based fusion perception.
[0006] Firstly, a method for sensory processing is provided, including:
[0007] The first node processes the first data using the first AI model to obtain the second data;
[0008] Wherein, the first data or the second data is data related to perception business, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing at least two different types of input data through an AI model; and intermediate data obtained by processing one type of input data and other types of input data through an AI model.
[0009] Secondly, a sensing processing apparatus is provided for a first node, the apparatus comprising: a first transceiver unit and a first processing unit;
[0010] The first processing unit is used to process the first data using a first AI model to obtain the second data;
[0011] Wherein, the first data or the second data is data related to perception business, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing at least two different types of input data through an AI model; and intermediate data obtained by processing one type of input data and other types of input data through an AI model.
[0012] Thirdly, a terminal is provided, comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method described in the first aspect.
[0013] Fourthly, a terminal is provided, including a processor and a communication interface, wherein the processor is used to process first data through a first AI model to obtain second data; wherein the first data or the second data is data related to perception services, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing at least two different types of input data through an AI model; and intermediate data obtained by processing one type of input data and other types of input data through an AI model.
[0014] Fifthly, a network-side device is provided, comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the sensing processing method as described in the first aspect.
[0015] In a sixth aspect, a network-side device is provided, including a processor and a communication interface, wherein the processor is used to process first data through a first AI model to obtain second data; wherein the first data or the second data is data related to perception services, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing at least two different types of input data through an AI model; and intermediate data obtained by processing one type of input data and other types of input data through an AI model.
[0016] In a seventh aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the perception processing method as described in the first aspect.
[0017] Eighthly, a wireless communication system is provided, comprising: a terminal and a network-side device, wherein the terminal is configured to perform the steps of the sensing processing method as described in the first aspect, or the network-side device is configured to perform the steps of the sensing processing method as described in the first aspect.
[0018] In a ninth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the perception processing method as described in the first aspect.
[0019] In a tenth aspect, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the perception processing method as described in the first aspect.
[0020] In this embodiment of the application, the first node processes the first data through the first AI model to obtain the second data; wherein, the first data or the second data is data related to perception business, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing the at least two different types of input data through the AI model; intermediate data obtained by processing one type of input data and other types of input data through the AI model, thereby realizing AI-based fusion perception. Attached Figure Description
[0021] Figure 1 is a schematic diagram of different sensing modes of integrated communication and sensing;
[0022] Figure 2 is a schematic diagram of the neural network structure;
[0023] Figure 3 is a schematic diagram of a neuron;
[0024] Figure 4 is a schematic diagram of the architecture of a wireless communication system provided in an embodiment of this application;
[0025] Figure 5 is a flowchart of a perception processing method provided in an embodiment of this application;
[0026] Figures 6a to 6d are schematic diagrams of AI fusion perception provided in the embodiments of this application;
[0027] Figure 7 is a schematic diagram of the extraction of a subset of time delay spectrum information provided in an embodiment of this application;
[0028] Figure 8 is a schematic diagram of time delay-Doppler spectrum information subset extraction provided in an embodiment of this application;
[0029] Figure 9 is a schematic diagram of the multipath channel response in the first dimension;
[0030] Figure 10 is a schematic diagram of a sensing processing apparatus provided in an embodiment of this application;
[0031] Figure 11 is a schematic diagram of a terminal provided in an embodiment of this application;
[0032] Figure 12 is a schematic diagram of a network-side device provided in an embodiment of this application. Detailed Implementation
[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0034] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0035] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0036] It is worth noting that the technologies described in this application are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in this application are often used interchangeably, and the described technologies can be used with the systems and radio technologies mentioned above, as well as with other systems and radio technologies. The following description describes New Radio (NR) systems for illustrative purposes, and the term NR is used in most of the following description; however, these technologies can also be applied to systems other than NR systems, such as 6th generation (6G) radio systems. th Generation 6G communication system.
[0037] To facilitate understanding of the embodiments of this application, the following technical points will be introduced first:
[0038] I. On the integration of communication and sensing.
[0039] Future mobile communication systems, such as Beyond 5th Generation (B5G) or Generation 6 (6G) systems, will possess sensing capabilities in addition to communication capabilities. Sensing capabilities refer to the ability of one or more devices to perceive information such as the location, distance, and speed of target objects through the transmission and reception of wireless signals, or to detect, track, identify, and image target objects, events, or environments. With the deployment of small base stations with high-frequency, high-bandwidth capabilities such as millimeter waves and terahertz waves in 6G networks, the resolution of sensing will be significantly improved compared to centimeter waves, enabling 6G networks to provide more refined sensing services. Typical sensing functions and application scenarios are shown in Table 1.
[0040] Table 1: Typical sensing functions and application scenarios.
[0041] Communication and sensing integration (referred to as wireless sensing integration) refers to the integrated design of communication and sensing functions within the same system through spectrum sharing and hardware sharing. While transmitting information, the system can sense information such as location, distance, and speed, and detect, track, and identify target devices or events. The communication system and the sensing system complement each other, thereby improving overall performance and bringing a better service experience.
[0042] The integration of communication and radar is a typical application of communication-sensing integration (communication-sensing fusion). In the past, radar systems and communication systems were strictly distinguished due to different research objects and focuses, and in most scenarios, the two systems were studied independently. In fact, radar and communication systems are both typical methods of information transmission, acquisition, processing, and exchange, and they share many similarities in terms of working principles, system architecture, and frequency bands. The design of integrated communication and radar systems is highly feasible, mainly in the following aspects: First, both communication and sensing systems are based on electromagnetic wave theory, using the transmission and reception of electromagnetic waves to complete information acquisition and transmission; second, both communication and sensing systems have structures such as antennas, transmitters, receivers, and signal processors, resulting in significant overlap in hardware resources; with technological advancements, their operating frequency bands also increasingly overlap; furthermore, they share similarities in key technologies such as signal modulation and reception detection, and waveform design. The integration of communication and radar systems can bring many advantages, such as cost savings, size reduction, power consumption reduction, improved spectral efficiency, and reduced mutual interference, thereby improving the overall system performance.
[0043] Based on the different target signal transmitting and receiving nodes, there are 6 basic sensing modes, as shown in Figure 1, which include:
[0044] (1) Base station self-transmitting and self-receiving sensing. In this sensing mode, base station A transmits a target signal and performs sensing measurements by receiving the echo of the target signal.
[0045] The aforementioned target signals include at least one of reference signals, synchronization signals, data signals, and dedicated signals. Receiving or transmitting target signals can support sensing services. For example, by receiving or transmitting target signals, sensing measurements or sensing results can be obtained. Sensing results refer to those that meet sensing requirements, such as: the shape of the sensing target, 2D or 3D environment reconstruction, spatial position, orientation, displacement, speed, and acceleration; radar-based sensing for velocity, distance, angle measurement, or imaging of target objects; the presence of people or objects; and sensing of targets such as human movements, gestures, respiratory rate, heart rate, and sleep quality.
[0046] (2) Inter-base station air interface sensing. Base station B receives the target signal sent by base station A and performs sensing measurements.
[0047] (3) Uplink air interface sensing. Base station A receives the target signal sent by terminal A and performs sensing measurements.
[0048] (4) Downlink air interface sensing. Terminal B receives the target signal sent by base station B and performs sensing measurements.
[0049] (5) Terminal self-transmitting and receiving sensing. Terminal A sends a first signal and performs sensing and measurement by receiving the echo of the target signal.
[0050] (6) Sidelink (SL) sensing between terminals. Terminal B receives the target signal sent by terminal A and performs sensing measurements.
[0051] It is worth noting that each sensing mode in Figure 1 uses a target signal transmitting node and a target signal receiving node as examples. In actual systems, one or more different sensing modes can be selected according to different sensing use cases and sensing requirements, and each sensing mode can have one or more transmitting and receiving nodes. The sensing targets in Figure 1 are people and vehicles as examples, and it is assumed that neither people nor vehicles carry or install signal receiving or transmitting devices. The sensing targets in real-world scenarios are much more diverse.
[0052] Optional configuration information for the sensing signal includes at least one of the following:
[0053] 1) Signal resource identifier (ID), used to distinguish different signal resource configurations;
[0054] 2) Signal Purpose: This indicates whether the signal is used for communication (e.g., channel measurement, channel estimation, synchronization, carrying data information, etc.), for sensing, or for both communication and sensing. Specifically, it can also specify which sensing service the signal is used for, or which type of sensing service it is used for.
[0055] 3) Waveforms, such as Orthogonal Frequency Division Multiplexing (OFDM), Single-Carrier Frequency-Division Multiple Access (SC-FDMA), Orthogonal Time-Frequency Space (OTFS), Frequency Modulated Continuous Wave (FMCW), pulse signals, etc.
[0056] 4) Subcarrier spacing, for example, the subcarrier spacing of an OFDM system is 30 kHz.
[0057] 5) Guard interval, which is the time interval from the moment the signal ends to the moment the latest echo signal of the signal is received; this parameter is proportional to the maximum sensing distance; for example, it can be expressed as c / (2R). max )Calculations show that R max For the maximum sensing distance (belonging to sensing demand information), such as for spontaneously generated and received sensing signals, R max This represents the maximum distance between the sensing signal transceiver point and the signal transmitter point; in some cases, the OFDM signal cyclic prefix (CP) can serve as a minimum guard interval; c is the speed of light.
[0058] 6) Starting frequency domain position, i.e., starting frequency point, can also be the starting resource element (RE) or resource block (RB) index;
[0059] 7) The starting time domain position, i.e. the starting time point, can also be the starting symbol index, time slot index, or frame index;
[0060] 8) The terminating frequency domain position, i.e., the terminating frequency point, can be represented by the terminating RE and RB indices;
[0061] 9) The termination time domain position, i.e. the termination time point, can be represented by the termination RE and RB indices;
[0062] 10) Frequency domain resource length, i.e. frequency domain bandwidth, which is inversely proportional to the distance resolution, wherein the frequency domain bandwidth B of each first signal is greater than or equal to c / (2ΔR), where c is the speed of light and ΔR is the distance resolution;
[0063] 11) Temporal resource length, also known as burst duration, is inversely proportional to Doppler resolution.
[0064] 12) Frequency domain resource spacing represents the spacing between adjacent signal frequency domain resource units. It can be represented by the number of REs or RBs, or by the density value Density. For example, Density = 1 means that there is one RE in each RB used to carry the signal. The frequency domain resource spacing is inversely proportional to the maximum unambiguous distance / delay. For OFDM systems, when subcarriers are continuously mapped, the frequency domain spacing is equal to the subcarrier spacing.
[0065] 13) Time-domain resource interval, which is the time interval between two adjacent signal resource units, and is associated with the maximum unambiguous Doppler frequency shift or the maximum unambiguous velocity.
[0066] 14) Time-domain resource characteristics: periodic transmission, semi-persistent transmission, and non-periodic transmission.
[0067] 15) Signal power, for example, a value taken in 2dBm increments from -20dBm to 23dBm.
[0068] 16) Sequence information, including sequence type information (ZC (Zadoff-Chu) sequence, pseudo-random sequence (PN sequence), etc.), sequence generation method, sequence length, etc.
[0069] 17) Signal direction, the angle information or beam information of the signal transmission.
[0070] 18) Quasi-co-located (QCL) relationships, for example, a sensing signal includes multiple resources, each resource is associated with a synchronization signal and PBCH block (SSB) QCL, and the QCL includes type A, B, C or D.
[0071] 19) Antenna port information, such as the maximum number of antenna ports and the antenna port index.
[0072] 20) Cyclic Prefix (CP) information, including CP type (e.g., Normal Cyclic Prefix (NCP), Extended Cyclic Prefix (ECP), or newly designed sensing measurement-specific CP), CP length, etc.
[0073] 2. Introduction to Artificial Intelligence.
[0074] Artificial intelligence (AI) has been widely applied in various fields. AI models can be implemented in various ways, such as neural networks, decision trees, support vector machines, and Bayesian classifiers. This application uses neural networks as an example, but does not limit the specific type of AI model. The structure of a neural network is shown in Figure 2.
[0075] The neural network is composed of neurons, and a schematic diagram of a neuron is shown in Figure 3. Where a1, a2, ... a K For input, w is the weight (multiplicative coefficient), b is the bias (additive coefficient), σ(.) is the activation function, and z = a1w1 + ... + a k w k +…+a K w K+b. Common activation functions include the sigmoid function, tanh function, rectified linear unit (ReLU), etc.
[0076] The parameters of a neural network can be optimized using optimization algorithms. An optimization algorithm is a class of algorithms that minimizes or maximizes an objective function (sometimes called a loss function). The objective function is often a mathematical combination of model parameters and data. For example, given data X and its corresponding label Y, we construct a neural network model f(.). With the model, we can obtain the predicted output f(x) based on the input x, and calculate the difference between the predicted value and the true value (f(x) - Y), which is the loss function. If we find suitable values W and b that minimize the value of the loss function, the smaller the loss value, the closer the model is to the reality.
[0077] Most common optimization algorithms are based on the error back propagation (BP) algorithm. The basic idea of the BP algorithm is that the learning process consists of two parts: forward propagation of the signal and backward propagation of the error. During forward propagation, the input sample is introduced from the input layer, processed layer by layer through the hidden layers, and then propagated to the output layer. If the actual output of the output layer does not match the expected output, the process transitions to the error back propagation stage. Error back propagation involves propagating the output error back to the input layer layer by layer through the hidden layers, distributing the error to all units in each layer, thus obtaining the error signal of each unit. This error signal serves as the basis for adjusting the weights of each unit. This process of adjusting the weights through forward and backward propagation is repeated continuously. This continuous adjustment of weights is the learning and training process of the network. This process continues until the error of the network output is reduced to an acceptable level, or until the predetermined number of learning iterations is reached.
[0078] Common optimization algorithms include gradient descent, stochastic gradient descent (SGD), mini-batch gradient descent, momentum method, momentum-driven stochastic gradient descent, adaptive gradient descent (Adagrad), Adadelta (an optimization algorithm with adaptive learning rate adjustment), root mean square propagation (RMSprop), and adaptive momentum estimation (Adam).
[0079] During error backpropagation, these optimization algorithms calculate the gradient by taking the derivative or partial derivative of the current neuron with respect to the error or loss obtained from the loss function, and then adding the effects of the learning rate, previous gradients, derivatives or partial derivatives, etc., and then pass the gradient to the previous layer.
[0080] In this application, the AI model may also be referred to as an AI unit, machine learning (ML) model, ML unit, AI structure, AI function, AI characteristic, machine learning model, neural network, neural network function, neural network functionality, etc. Alternatively, the AI model may refer to a processing unit capable of implementing specific algorithms, formulas, processing flows, capabilities, etc., related to AI. Or, the AI model may be a processing method, algorithm, function, module, or unit for a specific dataset. Alternatively, the AI model may be a processing method, algorithm, function, module, or unit running on AI or ML-related hardware such as a graphics processing unit (GPU), neural processing unit (NPU), tensor processing unit (TPU), or application-specific integrated circuit (ASIC). This application does not impose specific limitations in this regard. Optionally, the specific dataset includes the input or output of the AI model.
[0081] Optionally, the identifier of the AI model may also be referred to as an AI unit identifier, an AI structure identifier, an AI algorithm identifier, or an identifier of a specific dataset associated with the AI model, or an identifier of a specific scenario, environment, channel characteristics, or device related to AI or ML, or an identifier of a function, characteristic, capability, or module related to AI or ML. This application does not make any specific limitations in this regard.
[0082] The aforementioned AI functionality can be understood as an AI algorithm function. For a terminal (e.g., user equipment (UE)), the AI functionality may include one or more AI models.
[0083] Figure 4 shows a block diagram of a wireless communication system applicable to an embodiment of this application. The wireless communication system includes a terminal 41 and a network-side device 42. The terminal 41 can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM, or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the specific type of terminal 41 is not limited in this application embodiment. Network-side equipment 42 may include access network equipment or core network equipment, wherein access network equipment may also be referred to as Radio Access Network (RAN) equipment, radio access network function, or radio access network unit. Access network equipment may include base stations, Wireless Local Area Network (WLAN) access points (APs), or Wireless Fidelity (WiFi) nodes, etc.The term "base station" can be referred to as Node B (NB), Evolved Node B (eNB), Next Generation Node B (gNB), New Radio Node B (NR Node B), Access Point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), Radio Base Station, Radio Transceiver, Basic Service Set (BSS), Extended Service Set (ESS), Home Node B (HNB), Home Evolved Node B, Transmit / Receive Point (TRP), or any other suitable term in the relevant field, as long as the same technical effect is achieved. The term "base station" is not limited to any specific technical terminology. It should be noted that this application embodiment only uses a base station in an NR system as an example for description and does not limit the specific type of base station.
[0084] Core network equipment, also known as core network nodes, core network functions, or core network elements, includes, but is not limited to, at least one of the following: Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), Home Subscriber Server (HSS), Centralized network configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (or L-NEF), and Binding Support. The core network functions include: BSF (Block Network Function), Application Function (AF), Location Management Function (LMF), Gateway Mobile Location Centre (GMLC), and Network Data Analytics Function (NWDAF). It should be noted that this application embodiment only uses core network equipment in the NR system as an example and does not limit the specific type of core network equipment. If the name of the core network equipment mentioned in this application embodiment changes in subsequent protocol versions (e.g., 6G), it will still be within the scope of protection of this application.
[0085] Optionally, the core network equipment can be implemented by one or more functional modules in a single device, or by multiple devices working together; this application does not specifically limit this. It is understood that the aforementioned functional modules can be network elements in hardware devices, software functional modules running on dedicated hardware, or virtualized functional modules instantiated on a platform (e.g., a cloud platform).
[0086] The following description, in conjunction with the accompanying drawings, details a method, apparatus, device, and medium for sensing processing provided in this application, through some embodiments and application scenarios.
[0087] Referring to Figure 5, an embodiment of this application provides a method for perception processing, the specific steps of which include:
[0088] Step 51: The first node processes the first data using the first AI model to obtain the second data;
[0089] The first node processes the first data using a first AI model, including but not limited to inference. That is, the first node performs AI inference on the first data using the first AI model to obtain the second data, i.e., the inference result.
[0090] Wherein, the first data or the second data is data related to perception business, and the first data includes at least one of the following:
[0091] 1) At least two different types of input data;
[0092] In other words, the first node can process at least two different types of input data through the first AI model to obtain the second data, thereby achieving AI-based fusion perception processing.
[0093] Optionally, the at least two different types of input data include first input data and second input data, see Figure 6a.
[0094] Optionally, the first input data and the second input data have at least the following differences:
[0095] 1) The devices used to sense the data are different;
[0096] 2) The ways of perceiving data are different;
[0097] 3) Different RATs;
[0098] 4) Different perception modes;
[0099] For example, monostatic sensing or bistatic sensing.
[0100] 5) Different frequency bands.
[0101] For example, frequency bands include, but are not limited to: sub-6GHz band and millimeter wave band.
[0102] Optionally, the sensor includes at least one of the following: a visible light camera, an infrared camera, a Global Navigation Satellite System (GNSS), a lidar, a millimeter-wave radar, a thermometer, a hygrometer, a barometer, a gyroscope, an accelerometer, a magnetometer, a gravity sensor, a sonar, and a rain gauge.
[0103] Optionally, the data sensed by the sensor includes sensing data generated by the different sensors mentioned above; wherein, the sensing data includes at least one of the following: received signal, channel information, spectral information calculated based on channel information or received signal, measurement quantity, attribute of the sensed target, state of the sensed target, and sensing result.
[0104] 2) Intermediate data obtained by processing at least two different types of input data through an AI model;
[0105] In other words, the first node can process at least two different types of intermediate data to obtain the second data through the first AI model, so as to achieve AI-based fusion perception processing. The at least two different types of intermediate data are obtained by processing at least two different types of input data through the AI model.
[0106] Optionally, intermediate data obtained by processing at least two different types of input data through an AI model include: first intermediate data and second intermediate data. The first intermediate data is obtained by processing the first input data through a second AI model, and the second intermediate data is obtained by processing the second input data through a third AI model, as shown in Figure 6b.
[0107] It is understandable that the second or third AI model can be deployed on the first node along with the first AI model, or the second or third AI model can be deployed on a node other than the first node.
[0108] 3) Intermediate data obtained by processing one type of input data from at least two different types of input data and other input data from at least two different types of input data through an AI model.
[0109] Optionally, intermediate data obtained by processing one type of input data from at least two different types of input data and other input data from at least two different types of input data through an AI model includes at least one of the following:
[0110] 1) The first input data and the third intermediate data, wherein the third intermediate data is obtained by processing the second input data through the fourth AI model, see Figure 6c;
[0111] 2) The second input data and the fourth intermediate data, wherein the fourth intermediate data is obtained by processing the first input data through the fifth AI model, see Figure 6d.
[0112] It is understandable that the fourth or fifth AI model can be deployed on the first node along with the first AI model, or the fourth or fifth AI model can be deployed on a node other than the first node.
[0113] In the embodiments of this application, the first AI model, the second AI model, the third AI model, the fourth AI model, and the fifth AI model can be the same AI model or different AI models.
[0114] Optionally, the first, second, third, fourth, or fifth node in this document may include, but is not limited to, at least one of a terminal and a network-side device. The network-side device may include, but is not limited to, at least one of a base station, a core network device, a third-party server, and an operation administration and maintenance (OAM) management device. The core network device may include, but is not limited to, at least one of a sensing function network element and a positioning management function.
[0115] Optionally, the base station may include a base station responsible for transmitting or receiving sensing signals and processing sensing data, or a base station that is not responsible for transmitting or receiving sensing signals but only for processing sensing data.
[0116] Optionally, the terminal may include a terminal responsible for transmitting or receiving sensing signals and processing sensing data, or a terminal that is not responsible for transmitting or receiving sensing signals but only for processing sensing data.
[0117] Specifically, signaling transmission between the base station and the terminal, and between terminal A and terminal B, can be via Radio Resource Control (RRC) signaling, Medium Access Control (MAC) Control Element (CE), Layer 1 signaling, or other newly defined sensing signaling; signaling transmission between the sensing function network element and the terminal can be via Non-Access Stratum (NAS) signaling (forwarded via AMF) and / or via RRC signaling, MAC CE, Layer 1 signaling, or other newly defined sensing signaling; interaction between the sensing function network element and the base station can be via AMF forwarding to the radio access network through the N2 interface; or the sensing function network element sends to the UPF, and the UPF sends to the radio access network (base station) through the N3 interface; or it can send to the radio access network through a newly defined interface; signaling transmission between base stations can be via the Xn interface.
[0118] Optionally, the sensing function network element, also known as the sensing network element or sensing network function, can be located on the access network side or the core network side. It refers to a network node in the core network and / or access network responsible for at least one of the following functions: sensing request processing, sensing resource scheduling, sensing information interaction, and sensing data processing. It can be an upgrade based on the AMF or LMF in the 5G network, or it can be other network nodes or newly defined network nodes. Specifically, the functional characteristics of the sensing function network element can include at least one of the following:
[0119] 1) Interact with wireless signal transmitting equipment or wireless signal measuring equipment (including the target terminal or the serving base station of the target terminal or the base station associated with the target area) to exchange target information. The target information includes sensing processing requests, sensing capabilities, sensing auxiliary data, sensing measurement types, sensing resource configuration information, etc., in order to obtain the value of the target sensing result or sensing measurement (uplink measurement or downlink measurement) sent by the wireless signal measuring equipment. The wireless signal can also be referred to as the sensing signal.
[0120] 2) The sensing method to be used is determined based on factors such as the type of sensing service, the information of sensing service consumers, the required Quality of Service (QoS) requirements, the sensing capabilities of the wireless signal transmitting equipment, and the sensing capabilities of the wireless signal measuring equipment. The sensing method may include: base station A transmitting and base station B receiving, or base station transmitting and terminal receiving, or base station A transmitting and receiving, or terminal transmitting and base station receiving, or terminal transmitting and receiving, or terminal A transmitting and terminal B receiving, etc.
[0121] 3) The sensing equipment serving the sensing service is determined based on factors such as the type of sensing service, the information of the sensing service consumers, the required sensing QoS requirements, the sensing capabilities of the wireless signal transmitting equipment, and the sensing capabilities of the wireless signal measuring equipment. The sensing equipment includes wireless signal transmitting equipment or wireless signal measuring equipment.
[0122] 4) Manage the overall coordination and scheduling of resources required for sensing services, such as configuring sensing resources for base stations or terminals accordingly;
[0123] 5) Process the values of the sensed measurements or perform calculations to obtain the sensing results. Further, verify the sensing results and estimate the sensing accuracy.
[0124] In one embodiment of this application, the first data satisfies at least one of the following: the first data is generated by the first node, or the first data is obtained by the first node from the second node.
[0125] In one embodiment of this application, if the first data is obtained by the first node from the second node, the method further includes:
[0126] The first node sends first information to the second node, the first information being used to indicate the content or format of the first data.
[0127] In one embodiment of this application, the content of the first data, the second data, the first input data, the second input data, the first intermediate data, the second intermediate data, the third intermediate data, the fourth intermediate data, the sensed data, or the sensed data includes at least one of the following:
[0128] 1) Receiving signals;
[0129] The aforementioned received signals include at least one of reference signals, synchronization signals, data signals, and dedicated signals. Acquiring these signals can support sensing services. For example, acquiring these signals can yield sensing measurements or sensing results. Sensing results refer to those that meet sensing requirements, such as: the shape of a sensing target, 2D or 3D environment reconstruction, spatial location, orientation, displacement, speed, and acceleration; radar-based sensing for velocity, distance, angle measurement, or imaging of target objects; the presence of people or objects; and sensing of targets such as human movements, gestures, respiratory rate, heart rate, and sleep quality.
[0130] 2) Channel information;
[0131] Optionally, the channel information includes at least one of the following: time-domain channel response, frequency-domain channel response, complex result of the channel response, amplitude or phase, I-channel or Q-channel data.
[0132] 3) Spectral information calculated based on channel information or received signals;
[0133] Optionally, the spectral information includes at least one of the following: time delay (distance) spectrum, Doppler (velocity) spectrum, angle spectrum, time delay (distance)-Doppler (velocity) spectrum, time delay (distance)-angle spectrum, time delay (distance)-Doppler (velocity)-angle spectrum, time-Doppler spectrum (micro-Doppler spectrum);
[0134] Optionally, the spectral information can refer to complex results, such as a time-delay-Doppler spectrum, which refers to the time delay, Doppler index, and corresponding complex values (including phase information) in a two-dimensional spectrum; the spectral information can also refer to a power spectrum, such as a time-delay-Doppler spectrum, which refers to the time delay, Doppler index, and corresponding power values (excluding phase information) in a two-dimensional spectrum.
[0135] Optionally, the spectral information mentioned above can also be spectral information calculated based on the received signal.
[0136] 4) Quantity to be measured;
[0137] Optionally, the measured quantities include at least one of the following: time delay, Doppler, angle, and intensity (power); wherein the time delay can be the time delay of different diameters.
[0138] Optionally, the basic measurement can be a quantified result of an actual value or soft information; wherein, soft information can be data described by mean and variance, or confidence interval and confidence level. For example, the mean and variance of a Gaussian distribution, or two values and their respective probabilities.
[0139] 5) Perceive the attributes or state of the target;
[0140] Optionally, the attributes or state of the perceived target include at least one of the following: distance, velocity, orientation, spatial location, and acceleration.
[0141] 6) Perception results.
[0142] Optionally, the perception results include at least one of the following: the presence, trajectory, movement, expression, vital signs, number of perception targets, imaging results, weather, air quality, shape, material, and composition of the perceived target.
[0143] Optionally, the first data, the second data, the first input data, the second input data, the first intermediate data, the second intermediate data, the third intermediate data, or the fourth intermediate data may include different contents from the six contents mentioned above (1) to (6). For example, the first data, the first input data, or the second input data may include channel information as input data, and the second data may include basic measurement quantities as output data. As another example, the first data, the first input data, or the second input data may include channel information as input data, and the second data may include sensing results as output data. As yet another example, the first data, the first input data, or the second input data may include channel information, the first intermediate data, the second intermediate data, the third intermediate data, or the fourth intermediate data may include spectral information calculated based on channel information or received signals, and the second data may include sensing results as output data.
[0144] Optionally, the first data, the second data, the first input data, the second input data, the first intermediate data, the second intermediate data, the third intermediate data, or the fourth intermediate data may include the same content among the six types of content mentioned above (1) to (6); for example, the input data and output data include received signals or channel information, in which case the AI model in the sensing node can overcome channel estimation errors and noise to obtain more accurate data.
[0145] In one embodiment of this application, the content of the first data, the second data, the first input data, the second input data, the first intermediate data, the second intermediate data, the third intermediate data, or the fourth intermediate data further includes at least one of the following:
[0146] 1) Perceive signal identification information;
[0147] Optionally, the sensed signal identification information may include the index of the reference signal, etc.
[0148] 2) Sensing and measurement configuration identification information;
[0149] 3) Perceive business information;
[0150] Optionally, the perceived service information includes the perceived service identifier.
[0151] 4) Data subscription identifier;
[0152] 5) Purpose of the measurement;
[0153] Optionally, the measurement application includes at least one of the following: communication, sensing, wireless sensing.
[0154] 6) Time information;
[0155] 7) Sensing node information;
[0156] Optionally, the sensing node information includes at least one of the following: sensing node identifier, sensing node location, and sensing direction of the sensing node.
[0157] 8) Sensing link information;
[0158] Optionally, the sensing link information includes at least one of the following: sensing link sequence number and transceiver node identifier.
[0159] 9) Description of measured quantities;
[0160] Optionally, the measurement description information includes at least one of the following: a) form, such as amplitude value, phase value, or a complex value combining amplitude and phase; b) resource type, such as time-domain measurement result or frequency-domain resource measurement result.
[0161] 10) Measurement index information.
[0162] Optionally, the measurement metrics include at least one of the following: signal-to-noise ratio (SNR), perceived SNR, and reference signal receiving power (RSRP). In one embodiment of this application, the sensing service includes at least one of the following: detecting the presence of a target, detecting the number of targets, positioning, trajectory tracking, velocity detection, distance detection, angle detection, acceleration detection, material analysis, composition analysis, shape detection, category classification, and radar cross section (RCS). Section (RCS) detection, polarization scattering characteristic detection, fall detection, intrusion detection, indoor positioning, gesture recognition, lip reading, gait recognition, facial expression recognition, face recognition, respiration monitoring, heart rate monitoring, pulse monitoring, humidity, brightness, temperature, or atmospheric pressure monitoring, air quality monitoring, weather condition monitoring, environmental reconstruction, terrain, building, or vegetation distribution detection, pedestrian or vehicle flow detection, crowd density, vehicle density detection, etc.; or, sensing services can also refer to a category of sensing services, that is, classifying multiple different sensing services according to certain characteristics, such as dividing them according to function into detection-type sensing services (e.g., including intrusion detection, fall detection), parameter estimation-type sensing services (distance, angle, speed calculation), recognition-type sensing services (action recognition, identity recognition), etc.; they can also be divided according to the sensing range (near-range sensing, medium-range sensing, long-range sensing), according to the level of sensing fineness (coarse-grained sensing, fine-grained sensing, etc.), according to power consumption or energy consumption, according to resource consumption, etc.
[0163] In one embodiment of this application, the format of the first data, the second data, the first input data, the second input data, the first intermediate data, the second intermediate data, the third intermediate data, or the fourth intermediate data includes at least one of the following:
[0164] 1) Dimensions of channel information;
[0165] For example, frequency domain channel response information on a single symbol (one-dimensional channel information), time-frequency domain channel response information on multiple symbols (two-dimensional channel information), time-frequency spatial domain channel response information of multiple symbols and multiple antennas (three-dimensional channel information), and different dimensions of scale, i.e., the number of sampling points or sampling interval (e.g., time-frequency domain density).
[0166] 2) The range of spectral information;
[0167] It can be a limitation on the range of different spectral information (or a truncation window for different dimensions of spectral information). That is, the spectral information can be the complete spectral information calculated based on the channel information, or it can be a subset of the complete spectral information, such as a subset of spectral information corresponding to a specific time delay or Doppler range in the time-delay-Doppler spectrum, or information on paths or sampling points in the time-delay-Doppler spectrum whose power or amplitude exceeds a preset threshold; it can also indicate the upper limit of the number of paths or sampling points of a specific input spectrum that the AI model supports; or it can indicate the minimum granularity of the specific input spectrum that the AI model supports, that is, the interval between two adjacent sampling points (corresponding to the sensing resolution).
[0168] For example, the spectral information includes partial spectral information, such as the information of the N2-N3 sampling points or the N2-N3 paths in the time delay spectrum (as shown in the boxed part in Figure 7).
[0169] For example, the subset of spectral information is the portion of the time-delay-Doppler spectrum where the absolute Doppler value is less than X1 and the time delay value is less than X2 (as shown in the boxed portion in Figure 8).
[0170] 3) Supports the number of measurements that can be input into the AI model at one time;
[0171] 4) Supported quantization methods for measured values;
[0172] Optionally, quantization methods include quantization granularity.
[0173] 5) Soft information type.
[0174] Soft information can be data described by mean and variance, or confidence interval and confidence level.
[0175] In one embodiment of this application, the first AI model, the second AI model, the third AI model, the fourth AI model, or the fifth AI model is determined by at least one of the following methods: determined by the first node based on second information configured by other nodes, determined autonomously by the first node, agreed upon by the protocol, determined by the first node based on higher-level signaling, or selected by the first node from the first AI model pool based on third information.
[0176] In one embodiment of this application, the second information includes at least one of the following:
[0177] 1) Model structure information;
[0178] Optionally, the model structure information may include at least one of the following: the type of AI model (such as Gaussian process, support vector machine, various neural networks (fully connected neural network, convolutional neural network, recurrent neural network or residual network, combination of multiple small networks, such as fully connected + convolution, convolution + residual) etc.) and the structure of the model (such as the number of layers of the neural network, the number of neurons in each layer, activation function etc.).
[0179] 2) Supported AI model hyperparameter configuration;
[0180] Optionally, the supported hyperparameter configurations for AI models include at least one of the following: relevant parameters in the kernel function, relevant parameters in the activation function, relevant parameters in the normalization layer, etc.
[0181] 3) Supported AI model data processing methods, i.e., the preprocessing methods for data before it is input into the AI model. Optionally, the model data processing methods may include, but are not limited to, at least one of the following: normalization, upsampling, downsampling, etc.
[0182] 4) Supported AI model execution cycles, i.e., how often the AI model is executed;
[0183] 5) Supported AI model update cycle, i.e. how often the AI model is updated;
[0184] 6) Supported AI model update information;
[0185] Optionally, the AI model update information includes at least one of the following: kernel function update information, hyperparameter update information, prediction mode update information, and computation mode update information.
[0186] 7) Complexity information of supported AI models, such as the number of floating point operations (FLOPs) for model inference, such as 100 iterations, hardware conditions, and computing conditions;
[0187] 8) Available AI resources, including computing or storage resources available for AI use;
[0188] 9) Supported AI frameworks or algorithms;
[0189] 10) Information on the number of AI models;
[0190] 11) AI model category information;
[0191] 12) AI model identification information;
[0192] 13) Priority information of AI models;
[0193] 14) AI model attribute information;
[0194] 15) AI model accuracy information;
[0195] 16) AI model error information;
[0196] 17) AI model feature information;
[0197] 18) Adapt to environmental information;
[0198] 19) Process delay information;
[0199] 20) Information on the fusion method of the AI model output results;
[0200] 21) AI model lifecycle information;
[0201] 22) Information about the input data of the AI model;
[0202] 23) Information from the AI model's output data.
[0203] In one embodiment of this application, the third information includes at least one of the following: model error information, mobility information of network-side devices, mobility information of terminals, environmental information of network-side devices, environmental information of terminals, perception accuracy requirement information, perception service information, and model priority information.
[0204] In one embodiment of this application, the input data is obtained through at least one of the following methods:
[0205] 1) Sensing data processing based on a single device;
[0206] 2) Joint processing of sensing data from multiple different devices;
[0207] 3) Processing of sensory data acquired based on a single sensing mode;
[0208] Optionally, the sensing mode includes at least one of the following: base station self-transmission and self-reception, base station A transmits and base station B receives, base station transmits and terminal receives, terminal transmits and base station receives, terminal self-transmission and self-reception, terminal A transmits and terminal B receives; or the sensing mode may also refer to single-base sensing or dual-base sensing.
[0209] 4) Joint processing of sensing data acquired from multiple different sensing modes;
[0210] 5) Sensor-based sensing data processing;
[0211] 6) Joint processing of sensor data based on sensors and wireless sensing;
[0212] Optionally, the sensor includes at least one of the following: a visible light camera, an infrared camera, a Global Navigation Satellite System (GNSS), a lidar, a millimeter-wave radar, a thermometer, a hygrometer, a barometer, a gyroscope, an accelerometer, a magnetometer, a gravity sensor, a sonar, a rain gauge, a rotation vector sensor, etc.
[0213] Optionally, the sensor's data format can be binary, ASCII, or a data format associated with a specific sensor type. For example, for an image sensor (camera), it could be YUV, RGB, etc.
[0214] Optionally, the accelerometer is used to measure the acceleration applied to the device, including acceleration along the x-axis, y-axis, and z-axis. Further, the results can be categorized as including gravitational acceleration, not including gravitational acceleration, including bias compensation, and not including bias compensation.
[0215] Optionally, a gyroscope is used to measure the rate of rotation (radians per second) around the device's x, y, and z axes, and can be represented as a three-dimensional vector, similar to an accelerometer.
[0216] Optionally, a magnetometer is used to monitor changes in the Earth's magnetic field, measuring the intensity of the geomagnetic field (in microtesla) along each of the three coordinate axes. Typically, this sensor is not used directly, but rather in conjunction with other sensors to obtain rotational angle information.
[0217] Optionally, a rotation vector sensor can be used to obtain the terminal angle by combining different sensors. The format is (the angle of rotation of the terminal around the x, y, and z axes, respectively, or the rotation angle relative to the East-North-Up (northeast sky) / North-East-Down (northeast earth) coordinate axes).
[0218] 7) Obtained by joint processing of sensing data from different frequency bands;
[0219] Optionally, sensing data from different frequency bands may include, but are not limited to: sensing data from the sub-6GHz band, sensing data from the millimeter-wave band, and sensing data from the THz band.
[0220] 8) Obtained by joint processing of sensing data based on different Radio Access Technologies (RATs).
[0221] Optional, RAT includes, but is not limited to, 4G, 5G, 6G, Wireless Fidelity (Wifi), Ultra Wide Band (UWB), Bluetooth, etc.
[0222] In one embodiment of this application, different input data, or different intermediate data, or different weights or proportions of input data and intermediate data, wherein the weights or proportions are associated with perception-related indicators corresponding to the input data.
[0223] The perception processing method of this application can be used in at least the following scenarios:
[0224] 1) Fusion of wireless sensing and sensor sensing; where sensors include visible light cameras, infrared cameras, GNSS, lidar, millimeter-wave radar, thermometers, hygrometers, barometers, gyroscopes, accelerometers, magnetometers, gravity sensors, sonar, rain gauges, etc.; among them, lidar, by emitting ultra-narrow laser beams and receiving reflected echoes, can obtain ultra-high angular resolution; at the same time, the optical frequency band has ultra-high bandwidth, enabling lidar to have ultra-high distance resolution. However, existing commercial lidar generally does not have speed measurement capabilities. The advantage of visual sensors such as visible light cameras and infrared cameras compared to wireless sensing is that they can image and, based on visual images and algorithms, can identify the visual features of targets (e.g., people, vehicles, etc.). The disadvantage of visual sensors compared to wireless sensing is that they cannot measure speed and have poorer ranging performance. The fusion of wireless sensing and sensor sensing includes the following:
[0225] a) The fusion of the sensing results of wireless sensing systems and sensors on the same target or the same spatial area improves the sensing accuracy;
[0226] b) The fusion of the sensing results of wireless sensing systems and sensors on different targets or different spatial areas expands the spatial coverage of sensing to meet sensing needs;
[0227] c) The wireless sensing system and sensors perform sensing results at different times, thereby improving the sensing update rate.
[0228] 2) Fusion of multi-point collaborative sensing; that is, fusion processing of sensing data from different devices or fusion processing of data from different sensing modes of the same device; the sensing data from different devices can be the processing of sensing data acquired from one sensing mode, or the processing of sensing data acquired from multiple different sensing modes; the sensing mode includes at least one of the following: base station self-transmission and self-reception, base station A transmits and base station B receives, base station transmits and terminal receives, terminal transmits and base station receives, terminal self-transmission and self-reception, terminal A transmits and terminal B receives; or the sensing mode can also refer to monostatic sensing or bistatic sensing.
[0229] 3) Fusion of sensing data from multiple RATs, i.e., support for processing sensing data acquired based on more than one wireless access technology (including but not limited to 4G, 5G, 6G, Wifi, UWB, Bluetooth, etc.).
[0230] In this embodiment, the first node processes the first data using a first AI model to obtain the second data. The first or second data is data related to perception services. The first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing the at least two different types of input data using an AI model; and intermediate data obtained by processing one type of input data and other types of input data using an AI model, thus achieving AI-based fusion perception.
[0231] The optional implementation methods of this application are described below with reference to Embodiments 1, 2 and 3.
[0232] Example 1
[0233] Referring to Figure 6a, the specific steps are as follows:
[0234] Step 1: The first node acquires the first input data and the second input data of the first AI model; wherein, the first AI model is deployed on the first node;
[0235] Optionally, the first input data and the second input data can be one of the following:
[0236] 1) The first input data and the second input data are wireless sensing data and sensor sensing data, respectively;
[0237] 2) The first input data and the second input data are wireless sensing data or sensor-sensed data from different devices, respectively.
[0238] 3) The first input data and the second input data are data from different sensing modes of the same device; for example, the first input data is data from monostatic sensing, and the second input data is data from bistatic sensing.
[0239] 4) The first input data and the second input data are sensing data from different frequency bands, such as sensing data from the sub-6GHz band and sensing data from the millimeter-wave band.
[0240] 5) The first input data and the second input data are sensing data from different RATs; including but not limited to sensing data from 4G, 5G, 6G, Wifi, UWB, Bluetooth, etc.
[0241] Optionally, the first input data or the second input data includes at least one of the data generated by the first node and the data obtained by the first node from other nodes;
[0242] Optionally, if the first input data or the second input data is data obtained by the first node (e.g., a terminal) from other nodes (e.g., network-side devices), then before other nodes send the first data to the first node, the other nodes need to receive first information, which indicates at least one of the following: relevant information about the content of the first input data or the second input data, and relevant information about the format of the first input data or the second input data.
[0243] Optionally, prior to step 1, the first node determines one or more first AI models, specifically in any of the following ways:
[0244] 1) A first node (e.g., a terminal) receives second information sent by other nodes (e.g., network-side devices), the second information including configuration information of one or more first AI models; the first node determines one or more first AI models based on the second information;
[0245] 2) A first node (e.g., a terminal) sends a request message to other nodes (e.g., network-side devices), the request message being used to request the configuration of one or more first AI models; the first node receives second information sent by other nodes; wherein, the second information includes configuration information of one or more first AI models; the first node determines one or more first AI models based on the second information;
[0246] 3) The first node is based on autonomously determining one or more first AI models;
[0247] 4) The first node determines one or more first AI models based on the protocol;
[0248] 5) The first node determines one or more first AI models based on higher-level signaling;
[0249] 6) Based on the third information, the first node selects one or more first AI models from the first AI model pool; wherein the first AI model pool includes K AI models, and K is a positive integer.
[0250] Step 2: The first node processes the first input data and the second input data through the first AI model (e.g., reasoning based on the first AI model) to obtain the second data (i.e., the output data);
[0251] Optionally, the first node sends the second data to other nodes, for example, the second data is processed by the AI model or non-AI model of other nodes to obtain the final output data.
[0252] Example 2
[0253] Referring to Figure 6b, the specific steps are as follows:
[0254] Step 1: The third node obtains the first input data of the second AI model, and the fourth node obtains the second input data of the third AI model; wherein, the third node or the fourth node and the first node can be the same node or different nodes, that is, the second AI model or the third AI model can be deployed on the first node along with the first AI model or deployed on a node other than the first node;
[0255] Optionally, the first input data and the second input data can be one of the following:
[0256] 1) The first input data and the second input data are wireless sensing data and sensor sensing data, respectively;
[0257] 2) The first input data and the second input data are wireless sensing data or sensor-sensed data from different devices, respectively;
[0258] 3) The first input data and the second input data are data from different sensing modes of the same device; for example, the first input data is data from monostatic sensing and the second input data is data from bistatic sensing.
[0259] 4) The first input data and the second input data are sensing data of different frequency bands, such as sensing data of the sub-6GHz frequency band and sensing data of the millimeter wave frequency band.
[0260] 5) The first input data and the second input data are sensing data from different RATs; including but not limited to sensing data from 4G, 5G, 6G, Wifi, UWB, Bluetooth, etc.
[0261] Optionally, if the third or fourth node is the same node as the first node, the first input data or the second input data includes at least one of the data generated by the first node and the data obtained by the first node from other nodes;
[0262] Optionally, if the first input data or the second input data is data obtained by the first node (e.g., a terminal) from other nodes (e.g., network-side devices), then before other nodes send the first data to the first node, the other nodes need to receive first information, which indicates at least one of the following: relevant information about the content of the first input data or the second input data, and relevant information about the format of the first input data or the second input data.
[0263] Optionally, before step 1, the first node, the third node, or the fourth node determines the first AI model, the second AI model, or the third AI model, specifically in any of the following ways:
[0264] 1) A first node, a third node, or a fourth node (e.g., a terminal) receives second information sent by other nodes (e.g., network-side devices), the second information including configuration information of one or more first AI models; the first node, the third node, or the fourth node determines a first AI model, a second AI model, or a third AI model based on the second information;
[0265] 2) A first node, third node, or fourth node (e.g., a terminal) sends a request message to other nodes (e.g., network-side devices), the request message being used to request the configuration of a first AI model, a second AI model, or a third AI model; the first node, third node, or fourth node receives second information sent by other nodes; wherein, the second information includes configuration information of the first AI model, the second AI model, or the third AI model; the first node, third node, or fourth node determines the first AI model, the second AI model, or the third AI model based on the second information;
[0266] 3) The first, third, or fourth node is based on the autonomous determination of the first, second, or third AI model;
[0267] 4) The first, third, or fourth node determines the first, second, or third AI model based on the protocol;
[0268] 5) The first, third, or fourth node determines the first, second, or third AI model based on higher-level signaling;
[0269] 6) The first node, the third node, or the fourth node selects a first AI model, a second AI model, or a third AI model from the first AI model pool based on the third information; wherein, the first AI model pool includes K AI models, and K is a positive integer.
[0270] Step 2: The second AI model of the third node processes the first input data (e.g., inference based on the second AI model) to obtain the first intermediate data. The third AI model of the fourth node processes the second input data (e.g., inference based on the third AI model) to obtain the second intermediate data. The first node processes the first intermediate data and the second intermediate data through the first AI model (e.g., inference based on the first AI model) to obtain the second data (i.e., the output data).
[0271] Optionally, the first node sends the second data to other nodes, for example, the second data is processed by the AI model or non-AI model of other nodes to obtain the final output data.
[0272] Example 3
[0273] Referring to Figure 6d, the specific steps are as follows:
[0274] Step 1: The fifth node obtains the first input data of the fifth AI model, and the first node obtains the second input data of the first AI model;
[0275] The fifth node and the first node can be the same node or different nodes. That is, the fifth AI model can be deployed on the first node as well as the first AI model, or deployed on a node other than the first node.
[0276] Optionally, the first input data and the second input data can be one of the following:
[0277] 1) The first input data and the second input data are wireless sensing data and sensor sensing data, respectively;
[0278] 2) The first input data and the second input data are wireless sensing data or sensor-sensed data from different devices, respectively;
[0279] 3) The first input data and the second input data are data from different sensing modes of the same device; for example, the first input data is data from monostatic sensing and the second input data is data from bistatic sensing.
[0280] 4) The first input data and the second input data are sensing data of different frequency bands, such as sensing data of the sub-6GHz frequency band and sensing data of the millimeter wave frequency band.
[0281] 5) The first input data and the second input data are sensing data from different RATs; including but not limited to sensing data from 4G, 5G, 6G, Wifi, UWB, Bluetooth, etc.
[0282] Optionally, if the third or fourth node is the same node as the first node, the first input data or the second input data includes at least one of the data generated by the first node and the data obtained by the first node from other nodes;
[0283] Optionally, if the first input data or the second input data is data obtained by the first node (e.g., a terminal) from other nodes (e.g., network-side devices), then before other nodes send the first data to the first node, the other nodes need to receive first information, which indicates at least one of the following: relevant information about the content of the first input data or the second input data, and relevant information about the format of the first input data or the second input data.
[0284] Optionally, before step 1, the first node or the fifth node determines the first AI model or the fifth AI model, specifically in any of the following ways:
[0285] 1) A first node or a fifth node (e.g., a terminal) receives second information sent by other nodes (e.g., network-side devices), the second information including configuration information of one or more first AI models; the first node or the fifth node determines a first AI model or a fifth AI model based on the second information;
[0286] 2) A first node or a fifth node (e.g., a terminal) sends a request message to other nodes (e.g., network-side devices), the request message being used to request the configuration of a first AI model or a fifth AI model; the first node or the fifth node receives second information sent by other nodes; wherein, the second information includes configuration information of the first AI model or the fifth AI model; the first node or the fifth node determines the first AI model or the fifth AI model based on the second information;
[0287] 3) The first or fifth node is based on the autonomously determined first or fifth AI model;
[0288] 4) The first or fifth node determines the first or fifth AI model based on the protocol;
[0289] 5) The first or fifth node determines the first or fifth AI model based on higher-level signaling;
[0290] 6) The first node or the fifth node selects the first AI model or the fifth AI model from the first AI model pool based on the third information; wherein the first AI model pool includes K AI models, and K is a positive integer.
[0291] Step 2: The fifth AI model of the fifth node processes the first input data (e.g., inference based on the fifth AI model) to obtain the fourth intermediate data. The first node processes the fourth intermediate data and the second input data through the first AI model (e.g., inference based on the first AI model) to obtain the second data (i.e., the output data).
[0292] Optionally, the first node sends the second data to other nodes, for example, the second data is processed by the AI model or non-AI model of other nodes to obtain the final output data.
[0293] In embodiments 1 to 3 above, when an AI model corresponds to multiple input data, it is also necessary to determine the weights or proportions of the multiple input data (e.g., the first input data and the second input data). These weights or proportions can be associated with perception-related indicators corresponding to the input data. For example, the larger the perception-related indicator, the higher the weight or proportion of the input data. Similarly, when there are multiple intermediate data (e.g., in embodiment 2), it is also necessary to determine the weights or proportions of the multiple intermediate data (e.g., the first intermediate data and the second intermediate data). These weights or proportions can be associated with perception-related indicators corresponding to the input data. Optionally, the respective weights or proportions of the multiple input data can be input into the AI model to assist the AI model in outputting more accurate output information, such as perception results.
[0294] In the above embodiments 1 to 3, when there are multiple input data or multiple intermediate data, it is necessary to preprocess the multiple input data or multiple intermediate data, such as unifying the coordinate system; for example, when the multiple input data or multiple intermediate data are spectral information, the coordinate ranges or granularities of the multiple spectral information are aligned.
[0295] In the above embodiments 1 to 3, the preprocessing process for multiple input data or multiple intermediate data can be performed on other nodes or on the node where the AI model is located before other nodes send multiple input data or multiple intermediate data to the node where the AI model is deployed; the preprocessing method can be notified to the node performing the preprocessing.
[0296] Optionally, the perception-related indicators include at least one of the following: a first indicator, a second indicator, and a third indicator, wherein the first indicator is related to the received power, the second indicator is related to the interference and noise power, and the third indicator is related to both the received power and the interference and noise power.
[0297] The perception-related indicators include at least one of the following (1) to (3):
[0298] (1) First indicator;
[0299] The first metric is the linear average value (in W) of the received power of the target-correlation path in the channel response measured from the target signal on the resource unit carrying the target signal;
[0300] (2) Second indicator;
[0301] The second indicator includes at least one of the following (2a) to (2c):
[0302] (2a) Fourth indicator;
[0303] The fourth metric is the sum of the linear average power of the paths other than the sensing target associated path in the channel response of the target signal on the target resource and the linear average power of the interference and noise from other signals other than the target signal on the first resource (in W). The first resource is the target resource or other resources other than the target resource. The target resource includes resource units carrying the first signal. The resource units can be time-domain resource units or frequency-domain resource units.
[0304] Optionally, the fourth metric = total received power - the first metric; where total received power can be expressed as: the linear average of the total received power on the target resource (including the received power of signals from the serving cell and non-serving cells, adjacent channel interference and thermal noise, etc.) (in W); or, total received power = RSSI * K1, where K1 is a coefficient, and the resource for measuring RSSI is the target resource or other resources (e.g., resources configured by higher-layer signaling).
[0305] (2b) The fifth indicator;
[0306] The fifth metric is the linear average of the interference and noise power from signals other than the target signal on the second resource, where the second resource is the target resource or other resources besides the target resource.
[0307] Optionally, the fifth indicator = total received power - first signal received power; where the first signal received power is the reference signal received power (RSRP) of the first signal.
[0308] (2c) The sixth indicator;
[0309] The sixth indicator is the linear average power (in W) of the power of the other paths (excluding the sensing target associated path) in the channel response of the target signal on the target resource.
[0310] Optional, the sixth indicator = RSRP of the first signal - the first indicator;
[0311] (3) The third indicator;
[0312] The third indicator includes at least one of the following (3a) to (3d):
[0313] (3a) The seventh indicator;
[0314] The seventh indicator represents the first indicator divided by the fourth indicator, that is, the seventh indicator = the first indicator / the fourth indicator;
[0315] (3b) The eighth indicator;
[0316] The eighth indicator represents the first indicator divided by the fifth indicator, that is, the eighth indicator = the first indicator / the fifth indicator;
[0317] (3c) Ninth indicator;
[0318] The ninth indicator represents the first indicator divided by the sixth indicator, that is, the ninth indicator = the first indicator / the sixth indicator;
[0319] (3d) Tenth indicator;
[0320] The tenth index represents the first index divided by the first received power and then multiplied by a preset first coefficient. The first received power represents the total received power on the target resource, or the first received power represents the product of the Received Signal Strength Indication (RSSI) and a preset second coefficient. The RSSI measurement resource is the target resource or other resources. That is, the tenth index = K2 * the first index / total received power, where K2 is the coefficient.
[0321] Optionally, the method for obtaining the perception target association path includes:
[0322] The first node (e.g., a terminal) performs channel estimation based on the target signal and the received signal corresponding to the target signal to obtain the channel response;
[0323] The first node transforms the channel response to the first dimension;
[0324] The first node determines the associated path of the perceived target in the path corresponding to the first dimension;
[0325] The first dimension includes at least one of the following: time delay dimension; Doppler dimension; azimuth dimension; pitch dimension.
[0326] Optionally, the first node determines the perception target association path in the path corresponding to the first dimension, including:
[0327] The first node selects the path that satisfies the first condition from the paths corresponding to the first dimension as the path associated with the perceived target;
[0328] The first condition includes at least one of the following:
[0329] 1) The first parameter of the path is greater than or equal to the first threshold or is within the first interval range;
[0330] 2) The difference between the first parameter of the path and the first-reach path or reference path is greater than or equal to the second threshold or falls within the second interval range;
[0331] 3) The second parameter of the path satisfies the preset modulation rule;
[0332] The first parameter includes at least one of the following: amplitude, power, intensity, energy, Doppler, time delay, and angle;
[0333] The second parameter includes at least one of the following: amplitude, power, intensity, energy, and phase.
[0334] Optionally, the first node selects a path that satisfies the second condition from the paths corresponding to the first dimension as the associated path of the perceived target, including:
[0335] The first node determines a first path set in the path corresponding to the first dimension, and the third parameter of each path in the first path set is greater than or equal to a third threshold. The third parameter includes at least one of the following: amplitude, power, intensity, and energy.
[0336] The first node determines the path that satisfies the first condition in the first path set as the associated path of the sensing target.
[0337] Optionally, the first indicator can be calculated as follows:
[0338] The first node performs channel estimation based on the transmitted first signal (hereinafter referred to as X(k)) and the corresponding received signal (hereinafter referred to as Y(k)) to obtain the channel response, i.e., H(k) = Y(k) / X(k), where k = 0, 1, 2, ..., K-1 represents the resource unit index. After obtaining the channel response H(k), the first node transforms it to the first dimension and determines the sensing target association path in the first dimension. Then, the power of the sensing target association path is calculated as the first index. If the sensing target association path includes multiple paths, the sum of the power of the multiple paths is calculated as the first index.
[0339] The first dimension includes at least one of the following: time delay dimension; Doppler dimension; azimuth dimension; pitch dimension, for example, time delay-Doppler dimension, time delay-Doppler-angle dimension, etc.
[0340] For example, H(f) is the channel response, where f = 0, 1, 2, ..., N-1 represents frequency domain sampling points (e.g., subcarrier index). Then, by performing an inverse Fourier transform on H(f), it can be transformed to the time delay dimension (the first dimension). As another example, H(f,t) is the channel response, where f = 0, 1, 2, ..., N-1 represents frequency domain sampling points (e.g., subcarrier index), and t = 0, 1, 2, ..., M-1 represents time domain sampling points (e.g., OFDM symbol index). Then, by performing an inverse Fourier transform along the frequency domain and a Fourier transform along the time domain, H(f,t) can be transformed to the time delay-Doppler dimension (the first dimension). As yet another example, H(f,t,s) is the channel response, where f = 0, 1, 2, ..., N-1 represents frequency domain sampling points (e.g., subcarrier index), and t = 0, 1, 2, ..., M-1 represents time domain sampling points (e.g., Orthogonal Frequency Division Multiplexing (OFDM)). Division Multiplexing (OFDM) symbol index), s = 0, 1, 2, ..., P-1 represents the spatial sampling point (antenna index or port index). Then, by performing inverse Fourier transform along the frequency domain dimension, Fourier transform along the time domain dimension, and Fourier transform along the antenna domain dimension on H(f,t,s), it can be transformed to the time delay-Doppler-angle dimension (the first dimension).
[0341] Optionally, the method for determining the path associated with the sensing target (referred to as the sensing path) in the channel response obtained from the first signal measurement is as follows:
[0342] Step 1: Determine the first path set. The paths in the first path set include those whose amplitude, power, intensity, or energy exceeds a preset threshold after the channel response is transformed to the first dimension. (For example, in Figure 9, paths 0, 1, 2, and 3 are the paths in the first path set).
[0343] Optionally, the preset threshold can be set to be higher than the noise threshold or higher than the noise interference threshold.
[0344] Understandably, the step of determining the first set of paths is optional, or the paths associated with the perceived target can be determined solely based on step 2.
[0345] Step 2: Select a path that satisfies the first condition from the first path set or from all paths, as the path associated with the perceived target.
[0346] Optionally, the first condition includes at least one of the following:
[0347] 1) The amplitude, power, intensity, or energy of the noise level exceeds a preset threshold or falls within a preset range; for example, the preset threshold is 5 times the noise threshold.
[0348] 2) The Doppler amplitude of the path exceeds the preset threshold or is within the preset range;
[0349] 3) The path delay exceeds a preset threshold or falls within a preset range;
[0350] 4) The angle of the radius exceeds the preset threshold or is within the preset range;
[0351] 5) The difference between the amplitude, power, intensity, or energy of the first-reaching path (e.g., line-of-sight (LOS) path) or the reference path (e.g., the path of a signal reflected by a known target (e.g., a reconfigurable intelligence surface (RIS) or backscatter or other known passive target)) exceeds a preset threshold or is within a preset range.
[0352] 6) The Doppler difference between the path and the first path (e.g., the LOS path) or the reference path (e.g., the signal path reflected by a known target (e.g., a RIS or Backscatter device or other known passive target)) exceeds a preset threshold or is within a preset range;
[0353] 7) The time delay difference between the path and the first path (e.g., the LOS path) or the reference path (e.g., the signal path reflected by a known target (e.g., a RIS or Backscatter device or other known passive target)) exceeds a preset threshold or is within a preset range.
[0354] 8) The angle difference between the path and the first path (e.g., the LOS path) or the reference path (e.g., the signal path reflected by a known target (e.g., a RIS or Backscatter device or other known passive target)) exceeds a preset threshold or is within a preset range;
[0355] 9) The amplitude, power, intensity, energy, or phase of the path satisfies a specific modulation rule, which is the modulation rule of the tag, backscatter device, or RIS. That is, the path associated with the sensing target can be a path that has been modulated and reflected by the tag, backscatter device, or RIS.
[0356] It should be noted that the first condition of each of the above can also be based on the results of statistics over a period of time; for example, the proportion of the above indicators (such as Doppler of the path, delay of the path, etc.) exceeding the preset threshold or falling within the preset range within the preset time window reaches the preset proportion, or the number of times the above indicators (such as Doppler of the path, delay of the path, etc.) exceed the preset threshold or fall within the preset range within the preset time window reaches the preset number.
[0357] The preset threshold or set range is sent to the receiving device by other devices, and determined by those other devices based on prior sensing information or sensing requirements. Alternatively, the preset threshold or set range is determined by the receiving device based on prior sensing information or sensing requirements.
[0358] Among them, the perception of prior information or perception needs includes at least one of the following:
[0359] 1) Sensing services or sensing service types;
[0360] Optionally, the sensing services may include, but are not limited to, at least one of the following: detecting the presence of a target, positioning, speed detection, distance detection, angle detection, acceleration detection, material analysis, composition analysis, shape detection, category classification, radar cross-section detection, polarization scattering characteristic detection, fall detection, intrusion detection, quantity statistics, indoor positioning, gesture recognition, lip reading, gait recognition, expression recognition, facial recognition, respiration monitoring, heart rate monitoring, pulse monitoring, humidity, brightness, temperature, or atmospheric pressure monitoring, air quality monitoring, weather condition monitoring, environmental reconstruction, topography, building or vegetation distribution detection. The sensing services can be categorized according to certain characteristics, such as detection of pedestrian or vehicle traffic, crowd density, and vehicle density. These services can be classified by function (e.g., intrusion detection, fall detection), parameter estimation (distance, angle, speed calculation), and recognition (action recognition, identity recognition). They can also be categorized by sensing range (near-range, medium-range, long-range), sensing precision (coarse-grained, fine-grained), power consumption, or resource usage. For example, if the sensing service is respiratory monitoring, the normal breathing rate can be determined based on the person's gender and age (e.g., males: 13-21 breaths / minute, females: 15-20 breaths / minute; adults: 12-20 breaths / minute, children: approximately 30-40 breaths / minute), which can serve as prior information for sensing.
[0361] 2) Perceive the target area;
[0362] Optionally, the sensing target area includes the location area of the sensing object, or the location area that needs to be imaged or reconstructed; for example, a preset range of the time delay of the sensing target association path is determined based on the approximate location or distance of the sensing object.
[0363] 3) Perceive the object type;
[0364] Optionally, the sensing objects can be classified according to their possible motion characteristics. Each type of sensing object includes information such as the motion velocity range, motion acceleration range, and typical RCS range of typical sensing objects.
[0365] 4) The number of targets perceived;
[0366] Optionally, the camera perception results, as a form of prior perception information, can be used to determine the number of perceived targets;
[0367] For example, in Figure 9, paths 0, 1, 2, and 3 are paths in the first path set, where paths 2 and 3 are sensing paths that satisfy a first condition (e.g., their time delay meets a preset threshold), and paths 0 and 1 are paths associated with other scatterers. In Figure 9, the horizontal axis represents the first dimension, and the vertical axis represents the normalized amplitude, power, intensity, or energy.
[0368] For frequency range 1, the reference point for the first indicator can be the antenna connector of the receiving device, such as a terminal. For frequency range 1, if the receiving device has multiple receiving channels, the first indicator measured and reported by the receiving device cannot be lower than the indicator of any single receiving channel. For frequency range 2, the first indicator measured for a certain receiving channel needs to be obtained by measuring the combined signal on multiple antenna elements corresponding to that receiving channel.
[0369] Optionally, the first indicator can be calculated as follows:
[0370] Optionally, when calculating the received power of the sensing target correlation path, it can also be the power of the sensing target correlation path in the first dimension and... The difference is used as the first indicator, where N1 represents the number of paths associated with the perceived target. It represents the average power of multiple paths outside the first path set in the first dimension.
[0371] In one embodiment of this application, the received power of the first signal is calculated as follows:
[0372] The received power of the first signal can be obtained by the receiving device after obtaining the channel response H(k), transforming it to the first dimension, determining the first path set in the first dimension, and then calculating the sum of the power of all paths in the first path set.
[0373] Optionally, the received power of the first signal can be calculated as follows:
[0374] The received power of the first signal can also be the sum of the powers of all paths in the first path set in the first dimension. The difference, where N2 represents the number of paths in the first path set.
[0375] Optional method for calculating total received power: Total Received Power
[0376] Optional, the calculation method for the third indicator:
[0377] The channel response H(k) is processed by the first filter to obtain H. filter1 (k), then according to H filter1 The received signal Y after the first filtering process is calculated from (k) and the first signal X(k). filter1 (k), i.e., Y filter1 (k)=H filter1 (k)X(k). Then subtract the received signal Y(k) after the first filtering process from the received signal Y(k). filter1 (k) thus obtaining the interference and noise signal Y σ1 (k), i.e., Y σ1 (k)=Y(k)-Y filter1 (k), and then calculate the third index.
[0378] The first filtering process is used to eliminate noise and interference in the first dimension, as well as paths associated with non-perceived targets. For example, the first filtering process sets the amplitude, power, intensity, or energy of paths other than those associated with perceived targets in Figure 9 to zero. The channel response H after the first filtering process... filter1 (k) does not contain noise and interference, nor does it contain paths associated with non-perceived targets; it only contains paths associated with perceived targets.
[0379] Optionally, the fourth indicator can be calculated as follows:
[0380] The channel response H(k) is processed by a second filter to obtain H. filter2 (k), then according to H filter2 The received signal Y after the second filtering process is calculated from the first signal X(k) and the first signal X(k). filter2 (k), i.e., Y filter2 (k)=H filter2 (k)X(k). Then subtract the received signal Y(k) after the second filtering process from the received signal Y(k). filter2 (k) thus obtaining the interference and noise signal Y σ2 (k), i.e., Y σ2 (k)=Y(k)-Y filter2 (k), and then calculate the fourth index.
[0381] The second filtering process can be noise interference suppression processing in the first dimension (e.g., setting the amplitude, power, intensity, or energy of other paths besides the first path set in Figure 9 to zero), or minimum mean square error (MMSE) filtering. The channel response H after the second filtering process... filter2 (k) does not contain noise and interference, but only contains paths from the first path set.
[0382] Optionally, the fourth indicator can be calculated as follows:
[0383] Based on the average power of multiple paths outside the first path set in the first dimension The fourth index P was calculated. σ2 ,Right now Where N represents the number of sampling points in the first dimension.
[0384] If the receiving device identifies multiple sensing targets, or if the receiving device obtains the number of sensing targets based on prior sensing information or sensing requirements, the following methods are available:
[0385] Method 1: Calculate the perception-related indicators for each sensing target separately. For example, in Figure 9, determine the paths associated with each sensing target, and then calculate the perception-related indicators for each sensing target. When calculating the third indicator for a certain sensing target (e.g., sensing target A), there are two methods: the third indicator of sensing target A = total received power - the first indicator of sensing target A; or, the third indicator of sensing target A = total received power - the first indicator of sensing target A - the first indicator of sensing target B; (assuming there are two sensing targets: A and B). Similarly, there are two ways to calculate the fifth indicator: the fifth indicator of sensing target A = the RSRP of the first signal - the first indicator of sensing target A; or, the fifth indicator of sensing target A = the RSRP of the first signal - the first indicator of sensing target A - the first indicator of sensing target B; (assuming there are two sensing targets: A and B).
[0386] Method 2: Calculate a perception-related index for multiple perception targets. For example, in Figure 9, determine the paths associated with any perception target, and then treat all these paths as paths associated with the perception target; this is equivalent to treating multiple perception targets as a virtual perception target, and then calculating the perception-related index corresponding to this virtual perception target.
[0387] Referring to Figure 10, an embodiment of this application provides a sensing processing device applied to a first node. The device 100 includes a first transceiver unit 101 and a first processing unit 102.
[0388] The first processing unit 102 is used to process the first data through the first AI model to obtain the second data;
[0389] Wherein, the first data or the second data is data related to perception business, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing at least two different types of input data through an AI model; and intermediate data obtained by processing one type of input data and other types of input data through an AI model.
[0390] In one embodiment of this application, the at least two different types of input data include first input data and second input data, and the first input data and second input data have at least the following differences:
[0391] 1) The devices used to sense the data are different;
[0392] 2) The ways of perceiving data are different;
[0393] 3) Different RATs;
[0394] 4) Different perception modes;
[0395] For example, monostatic sensing or bistatic sensing.
[0396] 5) Different frequency bands.
[0397] For example, frequency bands include, but are not limited to: sub-6GHz band and millimeter wave band.
[0398] In one embodiment of this application, intermediate data obtained by processing at least two different types of input data through an AI model includes: first intermediate data and second intermediate data. The first intermediate data is obtained by processing the first input data through a second AI model, and the second intermediate data is obtained by processing the second input data through a third AI model.
[0399] In one embodiment of this application, intermediate data obtained by processing one type of input data from at least two different types of input data and other input data from at least two different types of input data through an AI model includes at least one of the following:
[0400] The first input data and the third intermediate data, wherein the third intermediate data is obtained by processing the second input data through the fourth AI model;
[0401] The second input data and the fourth intermediate data, wherein the fourth intermediate data is obtained by processing the first input data through the fifth AI model.
[0402] In one embodiment of this application, the sensor includes at least one of the following: a visible light camera, an infrared camera, a global navigation satellite system, a lidar, a millimeter-wave radar, a thermometer, a hygrometer, a barometer, a gyroscope, an accelerometer, a magnetometer, a gravity sensor, a sonar, and a rain gauge.
[0403] In one embodiment of this application, the second, third, fourth, or fifth AI model is deployed on the first node, or deployed on a node other than the first node.
[0404] In one embodiment of this application, the first data satisfies at least one of the following: the first data is generated by the first node, or the first data is obtained by the first node from the second node.
[0405] In one embodiment of this application, the first transceiver unit 101 is used to send first information to the second node, the first information being used to indicate the content of the first data or the format of the first data.
[0406] In one embodiment of this application, the content of the first data or the second data or the first input data or the second input data or the first intermediate data or the second intermediate data or the third intermediate data or the fourth intermediate data includes at least one of the following: received signal, channel information, spectral information calculated based on channel information or received signal, measurement quantity, attribute of the perceived target, state of the perceived target, and perception result.
[0407] In one embodiment of this application, the content of the first data or the second data or the first input data or the second input data or the first intermediate data or the second intermediate data or the third intermediate data or the fourth intermediate data may further include at least one of the following: sensing signal identification information, sensing measurement configuration identification information, sensing service information, data subscription identification, measurement purpose, time information, sensing node information, sensing link information, measurement description information, and measurement index information.
[0408] In one embodiment of this application, the format of the first data or the second data or the first input data or the second input data or the first intermediate data or the second intermediate data or the third intermediate data or the fourth intermediate data includes at least one of the following: the dimension of channel information, the range of spectral information, the number of measurements that can be input into the AI model at one time, the quantization method of the values of the supported measurements, and the type of soft information.
[0409] In one embodiment of this application, the first AI model, the second AI model, the third AI model, the fourth AI model, or the fifth AI model is determined by at least one of the following methods: determined by the first node based on second information configured by other nodes, determined autonomously by the first node, agreed upon by the protocol, determined by the first node based on higher-level signaling, or selected by the first node from the first AI model pool based on third information.
[0410] In one embodiment of this application, the second information includes at least one of the following: model structure information, hyperparameter configuration of supported AI models, supported model data processing methods, supported model allowed cycles, supported model update cycles, supported model update information, supported model complexity information, available AI resources, supported AI frameworks or algorithms, number of models, model category information, model identification information, model priority information, model attribute information, model accuracy information, model error information, model feature information, adaptation environment information, processing latency information, fusion method information of model output results, model lifecycle information, information on model input data, and information on model output data.
[0411] In one embodiment of this application, the third information includes at least one of the following: model error information, mobility information of network-side devices, mobility information of terminals, environmental information of network-side devices, environmental information of terminals, perception accuracy requirement information, perception service information, and model priority information.
[0412] In one embodiment of this application, the input data is obtained through at least one of the following methods:
[0413] Based on the processing of sensing data from a single device;
[0414] Joint processing of sensing data from multiple different devices;
[0415] Processing of sensory data acquired based on a sensing mode;
[0416] Joint processing of sensing data acquired from multiple different sensing modes;
[0417] Sensor-based sensing data processing;
[0418] Joint processing of sensor data based on sensor and wireless sensing;
[0419] Joint processing of sensing data from different frequency bands;
[0420] Joint processing of perception data based on different RATs.
[0421] In one embodiment of this application, different input data, or different intermediate data, or different weights or proportions of input data and intermediate data, wherein the weights or proportions are associated with perception-related indicators corresponding to the input data.
[0422] The apparatus provided in this application embodiment can implement the various processes implemented in the method embodiment of FIG5 and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0423] This application also provides a terminal, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps in the method embodiment shown in FIG5. This terminal embodiment corresponds to the above-described terminal-side method embodiment, and all implementation processes and methods of the above-described method embodiments can be applied to this terminal embodiment and can achieve the same technical effect. The terminal may be the sensing processing device shown in FIG10. Specifically, FIG11 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of this application.
[0424] The terminal 1100 includes, but is not limited to, at least some of the following components: radio frequency unit 1101, network module 1102, audio output unit 1103, input unit 1104, sensor 1105, display unit 1106, user input unit 1107, interface unit 1108, memory 1109, and processor 1110.
[0425] Those skilled in the art will understand that terminal 1100 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to processor 1110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The terminal structure shown in Figure 11 does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0426] It should be understood that, in this embodiment, the input unit 1104 may include a graphics processor 11041 and a microphone 11042. The graphics processor 11041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1106 may include a display panel 11061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1107 includes at least one of a touch panel 11071 and other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include a touch detection device and a touch controller. Other input devices 11072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0427] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1101 can transmit it to the processor 1110 for processing; in addition, the radio frequency unit 1101 can send uplink data to the network-side device. Typically, the radio frequency unit 1101 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0428] The memory 1109 can be used to store software programs or instructions, as well as various data. The memory 1109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1109 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1109 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0429] Processor 1110 may include one or more processing units; optionally, processor 1110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1110.
[0430] The processor 1110 is used to process the first data through a first AI model to obtain the second data; the first data or the second data is data related to perception business, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing the at least two different types of input data through the AI model; intermediate data obtained by processing one type of input data and other types of input data through the AI model.
[0431] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description in Figure 5 of the method embodiment and achieve the same or corresponding technical effects. To avoid repetition, it will not be described again here.
[0432] This application also provides a network-side device. As shown in FIG12, the network-side device 1200 includes a processor 1201, a network interface 1202, and a memory 1203. The network-side device may be the sensing processing device shown in FIG10. The network interface 1202 is, for example, a Common Public Radio Interface (CPRI).
[0433] Specifically, the network-side device 1200 in this application embodiment further includes: instructions or programs stored in memory 1203 and executable on processor 1201. Processor 1201 calls the instructions or programs in memory 1203 to execute the methods executed by each unit shown in FIG10 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0434] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the method embodiment in Figure 5 above and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0435] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0436] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the method embodiment in Figure 5 above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0437] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0438] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the method embodiment in FIG5 above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0439] This application also provides a communication system, including: a terminal and a network-side device, wherein the terminal can be used to perform the steps of the method shown in Figure 5 above, or the network-side device can be used to perform the steps of the method shown in Figure 5 above.
[0440] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0441] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0442] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. A method for sensory processing, comprising: The first node processes the first data using the first artificial intelligence (AI) model to obtain the second data; Wherein, the first data or the second data is data related to perception business, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing at least two different types of input data through an AI model; and intermediate data obtained by processing one type of input data and other types of input data through an AI model.
2. The method of claim 1, wherein, The at least two different types of input data include first input data and second input data, and the first input data and second input data have at least the following differences: Different devices are used to sense data; The ways of perceiving data are different; Different RATs; Different perception modes; Different frequency bands.
3. The method of claim 1 or 2, wherein, Intermediate data obtained by processing at least two different types of input data through an AI model include: first intermediate data and second intermediate data. The first intermediate data is obtained by processing the first input data through a second AI model, and the second intermediate data is obtained by processing the second input data through a third AI model.
4. The method of claim 1 or 2, wherein, Intermediate data obtained by processing one type of input data from at least two different types of input data and other input data from at least two different types of input data through an AI model, including at least one of the following: The first input data and the third intermediate data, wherein the third intermediate data is obtained by processing the second input data through the fourth AI model; The second input data and the fourth intermediate data, wherein the fourth intermediate data is obtained by processing the first input data through the fifth AI model.
5. The method of claim 3 or 4, wherein, The second, third, fourth, or fifth AI model is deployed on the first node, or on a node other than the first node.
6. The method of claim 1, wherein, The first data satisfies at least one of the following: the first data is generated by the first node, or the first data is obtained by the first node from the second node.
7. The method of claim 7, wherein, If the first data is obtained by the first node from the second node, the method further includes: The first node sends first information to the second node, the first information being used to indicate the content or format of the first data.
8. The method of claim 1 or 2 or 3 or 4 or 7, wherein, The content of the first data or the second data or the first input data or the second input data or the first intermediate data or the second intermediate data or the third intermediate data or the fourth intermediate data includes at least one of the following: received signal, channel information, spectral information calculated based on channel information or received signal, measurement quantity, attribute of the perceived target, state of the perceived target, and perception result.
9. The method of claim 8, wherein, The content of the first data or the second data or the first input data or the second input data or the first intermediate data or the second intermediate data or the third intermediate data or the fourth intermediate data may further include at least one of the following: sensing signal identification information, sensing measurement configuration identification information, sensing service information, data subscription identification, measurement purpose, time information, sensing node information, sensing link information, measurement description information, and measurement index information.
10. The method of claim 1 or 2 or 3 or 4 or 7, wherein, The format of the first data, the second data, the first input data, the second input data, the first intermediate data, the second intermediate data, the third intermediate data, or the fourth intermediate data includes at least one of the following: the dimension of the channel information, the range of the spectral information, the number of measurements that can be input into the AI model at one time, the quantization method of the values of the supported measurements, and the type of soft information.
11. The method of claim 1 or 3 or 4, wherein, The first AI model, the second AI model, the third AI model, the fourth AI model, or the fifth AI model is determined by at least one of the following methods: determined by the first node based on second information configured by other nodes, determined autonomously by the first node, agreed upon by the protocol, determined by the first node based on higher-level signaling, or selected by the first node from the first AI model pool based on third information.
12. The method of claim 11, wherein, The second information includes at least one of the following: model structure information, hyperparameter configuration of supported AI models, supported model data processing methods, allowed model cycles, supported model update cycles, supported model update information, supported model complexity information, available AI resources, supported AI frameworks or algorithms, number of models, model category information, model identification information, model priority information, model attribute information, model accuracy information, model error information, model feature information, adaptation environment information, processing latency information, fusion method information of model output results, model lifecycle information, information on model input data, and information on model output data.
13. The method of claim 11, wherein, The third information includes at least one of the following: model error information, network-side device mobility information, terminal mobility information, network-side device environmental information, terminal environmental information, perception accuracy requirement information, perception service information, and model priority information.
14. The method of claim 1, wherein, The input data is obtained through at least one of the following methods: Based on the processing of sensing data from a single device; Joint processing of sensing data from multiple different devices; Processing of sensory data acquired based on a sensing mode; Joint processing of sensing data acquired from multiple different sensing modes; Sensor-based sensing data processing; Joint processing of sensor data based on sensor and wireless sensing; Joint processing of sensing data from different frequency bands; Joint processing of perception data based on different RATs.
15. The method of claim 1, wherein, Different input data, or different intermediate data, or different weights or proportions of input data and intermediate data, wherein the weights or proportions are associated with perception-related indicators corresponding to the input data.
16. An apparatus for perception processing, applied to a first node, the apparatus comprising: First transceiver unit and first processing unit; The first processing unit is used to process the first data through a first AI model to obtain the second data; wherein the first data or the second data is data related to perception business, and the first data includes at least one of the following: at least two different types of input data; intermediate data obtained by processing the at least two different types of input data through the AI model; intermediate data obtained by processing one type of input data and other types of input data through the AI model.
17. The apparatus of claim 16, wherein, The at least two different types of input data include first input data and second input data. The first input data and the second input data have at least the following differences: Different devices are used to sense data; The ways of perceiving data are different; Different RATs; Different perception modes; Different frequency bands.
18. A terminal comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the perceptual processing method as claimed in any one of claims 1 to 15.
19. A network-side device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the sensing processing method as claimed in any one of claims 1 to 15.
20. A readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the perceptual processing method as claimed in any one of claims 1 to 15.
21. A computer program product comprising computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 15.