A three-dimensional monitoring and perception method for rich media data

Through the stereo monitoring and perception model, the problem of difficult to process multi-dimensional media-rich signal data in the prior art is solved, and high-dimensional feature fusion and deep environment perception are realized to obtain more accurate perception results.

CN119202627BActive Publication Date: 2025-06-1752ND RES INST AT CHINA ELECTRONICS TECH GRP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411708302.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-06-17
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

When existing perception methods process large-scale, multi-dimensional, and high-complex rich media signal data, it is difficult to adapt to different types of data, and it is impossible to effectively mine hidden information and complex intrinsic relationships in the data.

Method used

The stereoscopic monitoring and perception model is adopted, including feature extraction module, feature fusion module, feature alignment module and large language model. Through feature extraction, fusion and mapping, ultraviolet light, visible light, infrared light and radio signal data are processed to achieve high-dimensional feature fusion and deep environment perception.

Benefits of technology

It realizes effective monitoring of different types of data, minimizes information loss, obtains more accurate perceptual results, and can effectively mine hidden information and complex intrinsic relationships in the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119202627B_ABST
    Figure CN119202627B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional monitoring and perception method for rich media data, which includes obtaining data to be monitored and inputting it into a trained three-dimensional monitoring and perception model to obtain a perception result including a situation description and a target recognition result. The three-dimensional monitoring and perception model includes a feature extraction module, a feature fusion module, a feature alignment module, a large language model, and a feature mapper that are sequentially connected from the data input to the output direction. The three-dimensional monitoring and perception method for rich media data monitors different types of data, effectively realizes the high-dimensional feature fusion of data in different modalities and different source domains, and realizes the three-dimensional perception of the deep environment; during the feature extraction process by the feature extraction module, features are extracted from both the data and the image, and shallow and deep fusion are performed by the feature fusion module to minimize information loss to the greatest extent and obtain a more accurate perception result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent perception monitoring, and particularly relates to a three-dimensional monitoring and perception method for rich media data. Background Art

[0002] The three-dimensional monitoring and perception of rich media data is for three-dimensional monitoring in different fields such as sea and air, and realizes in-depth environmental three-dimensional perception of rich media signal data in different bands such as ultraviolet light, visible light, infrared light, and radio signals, aiming to comprehensively and accurately analyze information such as space, targets, and situations in complex environments. However, with the explosive growth of data, traditional data analysis methods face many challenges in processing large-scale, multi-dimensional, and high-complexity rich media signal data.

[0003] Most of the existing perception methods are based on simple statistical models, and these methods have obvious limitations in processing ability, model generalization ability, and intelligent processing means. They often rely on prior knowledge in specific fields to process a certain type of data alone, are difficult to adapt to different rich media signal data, and cannot effectively mine the hidden information and complex internal correlation relationships in the data. Summary of the Invention

[0004] The purpose of the present invention is to propose a three-dimensional monitoring and perception method for rich media data to solve the problems raised in the background art.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A three-dimensional monitoring and perception method for rich media data proposed by the present invention includes obtaining data to be monitored, inputting it into a trained three-dimensional monitoring and perception model, and obtaining a perception result including a situation description and a target recognition result;

[0007] Among them, the three-dimensional monitoring and perception model includes a feature extraction module, a feature fusion module, a feature alignment module, a large language model, and a feature mapper connected in sequence from the data input to the output direction;

[0008] Taking the data to be monitored as the input of the three-dimensional monitoring and perception model, the first signal feature and the second signal feature are obtained through the feature extraction module;

[0009] Both the first signal feature and the second signal feature are input into the feature fusion module to obtain a signal fusion feature;

[0010] Inputting the signal fusion feature into the feature alignment module to obtain an aligned feature, and inputting the aligned feature and a preset instruction feature into the large language model to obtain a situation description;

[0011] Input the situation description into the feature mapper to obtain the target recognition result, and the target recognition result includes the target category and the target location.

[0012] Preferably, the data to be monitored is at least one of ultraviolet light information, visible light information, infrared light information, and radio signal data at the same time and in the same area, and the ultraviolet light information includes ultraviolet light data and ultraviolet light images, the visible light information includes visible light data and visible light images, and the infrared light information includes infrared light data and infrared light images.

[0013] Preferably, the feature extraction module includes a pre-processor and a feature encoding module. For the radio signal data in the data to be monitored, it is first processed by the pre-processor and then input into the feature encoding module. For the ultraviolet light information, visible light information, and infrared light information in the data to be monitored, the images in each information are directly input into the feature encoding module.

[0014] Preferably, in the pre-processor, first, the radio signal data is channelized to obtain a preset number of independent low-rate baseband sub-channel signals, and then each low-rate baseband sub-channel signal is subjected to a fast Fourier transform to obtain a two-dimensional time-frequency image.

[0015] Preferably, the feature encoding module includes an image feature encoder and an original signal encoder;

[0016] For the images and the two-dimensional time-frequency images in each information, they are input into the image feature encoder. And in the image feature encoder, for each image: the image is divided into × grid regions, and all the grid regions pass through the first feature extraction structure to obtain the first image feature representation and the feature weights, where both the first image feature representation and the feature weights are n-dimensional vectors, and each element in the first image feature representation is an integer, with a value range of [1, m], and each element value in the feature weights is between 0 and 1;

[0017] Query the corresponding row vectors in the preset dictionary library according to each element value of the first image feature representation, and the row vectors corresponding to all the element values of the first image feature representation form a query matrix with a dimension of n×d, and the dictionary library is m 1×d vectors;

[0018] Perform a weighted summation calculation on the query matrix through the feature weights to obtain the first signal feature, and then obtain the first signal feature corresponding to each image. The calculation formula of the first signal feature is as follows:

[0019] ;

[0020] where

[0021] ;

[0022] Among them, represents the first signal feature, represents the row vector of the query matrix, represents the feature weight corresponding to the row vector of the query matrix, represents the th element value of the feature weight;

[0023] For the data and radio signal data in each piece of information, input them into the original signal encoder to obtain the second signal features corresponding to each data, and the original signal encoder has a multi-layer perceptron structure;

[0024] The first feature extraction structure is a convolutional neural network.

[0025] Preferably, the feature fusion module includes a shallow feature fuser and a deep feature fuser connected in sequence from the data input to the output;

[0026] In the shallow feature fuser, the first signal features and the second signal features corresponding to each piece of information or radio signal data in the data to be monitored are concatenated according to the feature channel dimension to obtain the first shallow fusion feature, and the corresponding first signal features and second signal features are concatenated according to the feature vector dimension to obtain the second shallow fusion feature;

[0027] The first shallow fusion features are passed through the first mapping module to obtain the first features, the second shallow fusion features are passed through the second mapping module to obtain the second features, then the corresponding first features and second features are added to obtain the third features, and after all the third features are concatenated, they are input into the deep feature fuser to obtain the signal fusion feature;

[0028] Both the first mapping module and the second mapping module are two fully connected layers connected in sequence.

[0029] Preferably, the deep feature fuser is a multi-layer self-attention mechanism structure, and each layer of the self-attention mechanism structure performs an attention mechanism calculation on the input of that layer, and the calculation formula is as follows:

[0030] ;

[0031] Among them, represents the output of the th layer of the self-attention mechanism structure, represents the output of the th layer of the self-attention mechanism structure, represents the dimension, represents function, Represents the query vector of the th layer, Represents the key vector of the th layer, Represents the value vector of the th layer, Represents transpose.

[0032] Preferably, both the feature alignment module and the feature mapper are a plurality of fully connected layers connected in sequence.

[0033] Preferably, when training the stereo monitoring and perception model, the loss function of the situation description and the loss function of the target recognition result are weighted to obtain the total loss, and the parameters of the stereo monitoring and perception model are updated by backpropagation based on the total loss.

[0034] Preferably, the training data set for training the stereo monitoring and perception model includes ultraviolet light information, visible light information, infrared light information, and radio signal data at the same time and in the same area.

[0035] Compared with the prior art, the beneficial effects of the present invention are:

[0036] The stereo monitoring and perception method for rich media data monitors different types of data, effectively realizes the high-dimensional feature fusion of data of different modalities and different source domains, and realizes the deep environment stereo perception;

[0037] During the process of feature extraction by the feature extraction module, features are extracted from both data and images, and shallow and deep fusion are performed by the feature fusion module to minimize information loss to the greatest extent and obtain a more accurate perception result;

[0038] During the process of image feature extraction by the image feature encoder, the corresponding row vector of each element value represented by the first image feature is queried in the preset dictionary library, and then the first signal feature is obtained by weighted summation calculation of the query matrix through the feature weight, so as to form a more effective fusion to obtain a better result. Brief Description of the Drawings

[0039] Figure 1 It is the module block diagram of the stereo monitoring and perception model of the stereo monitoring and perception method for rich media data of the present invention;

[0040] Figure 2 It is the module block diagram of the preprocessor of the present invention;

[0041] Figure 3 It is the module block diagram of the feature encoding module of the present invention;

[0042] Figure 4 It is the module block diagram of the feature fusion module of the present invention;

[0043] Figure 5 This is the module block diagram of the shallow feature fusion device of the present invention. Specific implementation manners

[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0045] In one embodiment, as Figures 1 - 5 shown, a three-dimensional monitoring and sensing method for rich media data is provided, including:

[0046] Step 1: Obtain the data to be monitored, input it into the trained three-dimensional monitoring and sensing model, and obtain a sensing result including a situation description and a target recognition result;

[0047] It should be noted that the data to be monitored is at least one of ultraviolet light information, visible light information, infrared light information, and radio signal data (such as electromagnetic signal data) at the same time and in the same area, and the ultraviolet light information includes ultraviolet light data and ultraviolet light images, the visible light information includes visible light data and visible light images, and the infrared light information includes infrared light data and infrared light images. The process of obtaining the data to be detected can be obtained through corresponding devices (such as cameras, sensors, etc.).

[0048] Step 2: The three-dimensional monitoring and sensing model includes a feature extraction module, a feature fusion module, a feature alignment module, a large language model, and a feature mapper that are sequentially connected from the data input to the output. The data to be monitored is used as the input of the three-dimensional monitoring and sensing model, and the first signal feature and the second signal feature are obtained through the feature extraction module;

[0049] Step 2.1: The feature extraction module includes a pre-processor and a feature encoding module. For the radio signal data in the data to be monitored, it is first processed by the pre-processor and then input into the feature encoding module. For the ultraviolet light information, visible light information, and infrared light information in the data to be monitored, the images in each piece of information are directly input into the feature encoding module.

[0050] Step 2.1: As Figure 2As shown in the figure, in the preprocessor, first, the radio signal data is channelized to obtain a preset number of independent low-rate baseband sub-channel signals (where the channelization process is performed using a preset number of filters. Each time a filter processes the data, a low-rate baseband sub-channel signal is obtained. The number of filters can be determined according to actual needs. For example, if there are two filters, two low-rate baseband sub-channel signals will be obtained). Then, each low-rate baseband sub-channel signal is subjected to a fast Fourier transform to obtain a two-dimensional time-frequency image (by plotting each low-rate baseband sub-channel signal after the fast Fourier transform in the same coordinate system, a two-dimensional time-frequency image is obtained).

[0051] Step 2.2, as Figure 3 shown (where Figure 3 the block diagram of the feature encoding module is shown taking visible light information as an example), the feature encoding module includes an image feature encoder and an original signal encoder.

[0052] Step 2.3, for the images and two-dimensional time-frequency images in each piece of information, they are input into the image feature encoder (the image feature encoder is a ViT network model based on Transformer). And in the image feature encoder, for each image: the image is divided into × grid regions (that is, rows columns, a total of grid regions. For example, if has a value of 16, a picture with a resolution of 224x224 pixels is divided into 256 grid regions). All the grid regions are passed through a first feature extraction structure (where the first feature extraction structure is a convolutional neural network, such as the Resnet18 network) to obtain a first image feature representation and feature weights. Both the first image feature representation and the feature weights are n-dimensional vectors (column vectors), and each element in the first image feature representation is an integer, with a value range of [1, m]. For example, m is 8192 and n is 8. Each element value in the feature weights is between 0 and 1, and the sum of all element values is 1;

[0053] According to each element value in the first image feature representation, the corresponding row vector is queried in a preset dictionary library. And the row vectors corresponding to all element values in the first image feature representation form a query matrix with a dimension of n×d (such as d is 768, that is, an 8×768 query matrix), and the dictionary library is composed of m 1×d vectors; the dictionary library is obtained by pre-training using a large amount of image data. For example, a convolutional neural network (ResNet50, VGG, etc.) or a ViT network is used. Taking any image as the input, the output is a 1×d vector. The vectors obtained for each picture are clustered, and the number of clustering categories is set to 8192. Finally, the 8192 clusters obtained are averaged to obtain 8192 1×d vectors, forming the dictionary library;

[0054] The first signal feature is obtained by weighted summation calculation of the query matrix through feature weights, and then the first signal features corresponding to each image are obtained (i.e., the first signal features corresponding to the ultraviolet image, visible light image, infrared light image, and two-dimensional time-frequency image respectively), and the calculation formula of the first signal feature is as follows:

[0055] ;

[0056] Among them,

[0057] ;

[0058] Among them, represents the first signal feature, represents the th row vector of the query matrix, represents the feature weight corresponding to the th row vector of the query matrix, represents the th element value of the feature weight;

[0059] Step 2.4: For the data and radio signal data in each piece of information, input them into the original signal encoder to obtain the second signal features corresponding to each data (i.e., the second signal features corresponding to the ultraviolet data, visible light data, infrared light data, and radio signal data respectively), and the original signal encoder has a multi-layer perceptron structure (MLP).

[0060] Step 3: Input both the first signal feature and the second signal feature into the feature fusion module to obtain the signal fusion feature.

[0061] Step 3.1: As shown in Figure 4 (where Figure 4 shows the module block diagram of the feature fusion module with visible light information as an example), the feature fusion module includes a shallow feature fusion device and a deep feature fusion device connected in sequence from the data input to the output direction;

[0062] Step 3.2: As shown in Figure 5As shown, in the shallow feature fuser, the first signal features and the second signal features corresponding to each piece of information or radio signal data in the data to be monitored are concatenated according to the feature channel dimension to obtain the first shallow fusion feature, and the corresponding first signal features and the second signal features are concatenated according to the feature vector dimension to obtain the second shallow fusion feature (where the corresponding first signal features and the second signal features mean: the first signal feature obtained from the ultraviolet light image in the ultraviolet light information and the second signal feature obtained from the ultraviolet light data are corresponding, the first signal feature obtained from the visible light image in the visible light information and the second signal feature obtained from the visible light data are corresponding, the first signal feature obtained from the infrared light image in the infrared light information and the second signal feature obtained from the infrared light data are corresponding, the first signal feature obtained from the two-dimensional time-frequency image and the second signal feature obtained from the radio signal data are corresponding, where the specific number of corresponding first signal features and second signal features depends on the types included in the data to be monitored, such as Figure 5 As shown, taking one kind of corresponding first signal feature and second signal feature as an example to show the module block diagram of the shallow feature fuser, and the same is true for the following corresponding first feature and second feature);

[0063] The first shallow fusion features are passed through the first mapping module to obtain the first features, the second shallow fusion features are passed through the second mapping module to obtain the second features, and then the corresponding first features and second features are added (adding features at corresponding positions) to obtain the third features. After all the third features are concatenated (when there are multiple third features, concatenation is required; when there is only one third feature, no concatenation is needed and it is directly input to the deep feature fuser), and then input to the deep feature fuser to obtain the signal fusion feature; both the first mapping module and the second mapping module are two fully connected layers connected in sequence;

[0064] It should be noted that since the data to be monitored is at least one of ultraviolet light information, visible light information, infrared light information, and radio signal data at the same time and in the same area. When the data to be monitored is one of ultraviolet light information, visible light information, infrared light information, and radio signal data at the same time and in the same area, for example, when the data to be monitored is visible light information, after steps 2 - 3.2, the obtained third feature is directly input into the deep feature fusion device; when the data to be monitored is multiple types of ultraviolet light information, visible light information, infrared light information, and radio signal data at the same time and in the same area, for example, these two types of ultraviolet light information and visible light information, after steps 2 - 3.2, the third feature corresponding to the ultraviolet light information and the third feature corresponding to the visible light information are obtained, and then these two third features are concatenated and then input into the deep feature fusion device (that is, when the data to be monitored contains one type, after passing through the shallow feature fusion device, the obtained third feature is directly input into the deep feature fusion device; when the data to be monitored contains multiple types, after passing through the shallow feature fusion device, the obtained multiple third features are first concatenated and then input into the deep feature fusion device).

[0065] Step 3.3: The deep feature fusion device is a multi - layer self - attention mechanism structure. Each layer of the self - attention mechanism structure performs an attention mechanism calculation on the input of that layer, and the calculation formula is as follows:

[0066] ;

[0067] Among them,

[0068] ;

[0069] ;

[0070] ;

[0071] Among them, represents the output of the th layer of the self - attention mechanism structure, represents the output of the th layer of the self - attention mechanism structure. When is equal to 1, , which is the input of the deep feature fusion device, represents the dimension, represents function, represents the query vector of the th layer, represents the key vector of the th layer, represents the value vector of the th layer, represents transpose; Indicates that the dimension of each vector is , Indicates that the output of the -th layer self-attention mechanism structure passes through the first feature mapping to obtain the query vector, Indicates that the output of the -th layer self-attention mechanism structure passes through the second feature mapping to obtain the key vector, Indicates that the output of the -th layer self-attention mechanism structure passes through the third feature mapping to obtain the value vector. Each feature mapping is performed through a fully connected layer, and the dimension is vector. After passing through the fully connected layer with parameters (i.e., the parameter matrix with dimension ), a vector is calculated, and the calculation method is matrix multiplication.

[0072] Step 4: Input the signal fusion feature into the feature alignment module to obtain the aligned feature. The aligned feature and the preset instruction feature are both input into the large language model to obtain the situation description. The feature alignment module is used to map the output of the deep feature fusion device and the input of the large language model to the same high-dimensional space;

[0073] It should be noted that the feature alignment module is multiple sequentially connected fully connected layers, such as two sequentially connected fully connected layers; the preset instruction feature is converted from the task instruction of the text prompt into a token (a token can be a word, a subword (such as a letter, a syllable, or a subword segment), a character, etc.), and then converted into an instruction feature by the method of BPE (Byte Pair Encoding). In this embodiment, the task instruction of the text prompt can be "Please evaluate the current situation", "Give a situation description", etc.; the large language model can be Qwen2.4-7B, Baichuan2-13B, etc.

[0074] Step 5: Input the situation description into the feature mapper to obtain the target recognition result, and the target recognition result includes the target category and the target location. The feature mapper is multiple sequentially connected fully connected layers, such as two sequentially connected fully connected layers, and the number of neurons in each fully connected layer is not less than 2048.

[0075] In another embodiment, during the training of the three-dimensional monitoring and perception model:

[0076] Step I: First, construct a training dataset. The training dataset includes ultraviolet light information, visible light information, infrared light information, and radio signal data at the same time and in the same area, and each piece of information includes corresponding data and images;

[0077] Step II. Annotate the training dataset: Pass the radio signal data through a pre-processor to obtain the corresponding two-dimensional time-frequency image, and annotate the target categories and locations in each image (i.e., ultraviolet image, visible light image, infrared light image, and two-dimensional time-frequency image), as well as describe the situation of the scene in the corresponding area of the data (i.e., ultraviolet data, visible light data, infrared light data, and radio signal data, and the areas corresponding to these four types of data are the same, so there is only one situation description for the scene in this area). When training the stereo monitoring and perception model, each input to the model includes at least two or more of ultraviolet light information, visible light information, infrared light information, and radio signal data;

[0078] Pre-training stage of the stereo monitoring and perception model:

[0079] Step III. Initialize the large language model and the feature encoding module based on the preset pre-training model parameters (the preset pre-training model uses an existing model, such as Tongyi Qianwen model), freeze the parameters of the large language model and the feature encoding module, and randomly initialize the parameters of the feature fusion module, the feature alignment module, and the feature mapper.

[0080] Step IV. Input the annotated dataset into the stereo monitoring and perception model for forward inference to obtain the target categories and locations;

[0081] Step V. Calculate the loss function value between the target categories and locations obtained by inference and the manually annotated labels; update the model parameters through backpropagation according to the loss function value, and repeat Step IV to Step V until the model converges.

[0082] Fine-tuning training stage of the stereo monitoring and perception model:

[0083] Step VI. First, initialize based on the stereo monitoring and perception model parameters obtained in the pre-training stage. Among them, the large language model and the feature encoding module are initialized with the same preset pre-training model parameters as in the pre-training stage, and the large language model is kept in a frozen state and does not participate in parameter updates during the training process;

[0084] Step VII. Input the annotated dataset into the stereo monitoring and perception model for forward inference to obtain the target categories and locations, as well as the situation description;

[0085] Step VIII. Calculate the loss function value between the target categories and locations obtained by inference and the manually annotated labels, and calculate the loss function value between the situation description obtained by inference and the manually made situation description. Weight the loss function of the situation description and the loss function of the target recognition result to obtain the total loss, and update the parameters of the stereo monitoring and perception model through backpropagation based on the total loss until the model converges;

[0086] where the loss function of the target recognition result is represented by and the weight is represented by ; the loss function of the situation description is represented by and the weight is represented by . Then the total loss is: and , , , such as .

[0087] The three-dimensional monitoring and perception method of the rich media data monitors different types of data, effectively realizes the high-dimensional feature fusion of data in different modalities (referring to two modalities of images and data) and different source domains (referring to different source domains of ultraviolet light information, visible light information, infrared light information, and radio signal data), and realizes the three-dimensional perception of the deep environment; in the process of feature extraction by the feature extraction module, features are extracted from both data and images, and shallow and deep fusions are performed through the feature fusion module to minimize information loss to the greatest extent and obtain more accurate perception results; during the process of image feature extraction by the image feature encoder, the row vector corresponding to each element value represented by the first image feature is queried in the preset dictionary library, and then the first signal feature is obtained by weighted summation calculation of the query matrix through the feature weight, so that a more effective fusion is formed to obtain a better result.

[0088] It should be understood that unless there is a clear description in this article, the execution of each step does not have a strict order limit. These steps can be executed in other orders, and the steps can be executed at the same time or at different times.

[0089] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it cannot be understood as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A stereoscopic monitoring perception method for rich media data, characterized by: The stereoscopic monitoring perception method of rich media data includes: Obtain the data to be monitored and input it into the trained stereoscopic monitoring perception model to obtain the perception results including situation description and target recognition results; Wherein, the stereo monitoring perception model includes a feature extraction module, a feature fusion module, a feature alignment module, a large language model and a feature mapper which are sequentially connected from the data input to the output direction; Using the data to be monitored as the input of the stereo monitoring perception model, and obtaining the first signal feature and the second signal feature through the feature extraction module; Inputting the first signal feature and the second signal feature into a feature fusion module to obtain a signal fusion feature; The signal fusion features are input into the feature alignment module to obtain the alignment features. The alignment features and the preset command features are input into the large language model to obtain the situation description. Inputting the situation description into a feature mapper to obtain a target recognition result, wherein the target recognition result includes a target category and a target position; The data to be monitored is at least one of ultraviolet light information, visible light information, infrared light information and radio signal data at the same time and in the same area, and the ultraviolet light information includes ultraviolet light data and ultraviolet light images, the visible light information includes visible light data and visible light images, and the infrared light information includes infrared light data and infrared light images; The feature extraction module includes a preprocessor and a feature encoding module, and the radio signal data is processed by the preprocessor to obtain a two-dimensional time-frequency image; The feature encoding module includes an image feature encoder and an original signal encoder; For each image and two-dimensional time-frequency image in each information, input it to the image feature encoder, and in the image feature encoder, for each image: divide the image into × grid areas, obtaining first image feature representatives and feature weights from all grid areas through a first feature extraction structure, wherein the first image feature representatives and feature weights are both n-dimensional vectors, and each element in the first image feature representative is an integer with a value range of [1, m], and the value of each element in the feature weight is between 0 and 1; According to each element value represented by the first image feature, a corresponding row vector is searched in a preset dictionary library, and the row vectors corresponding to all element values ​​represented by the first image feature constitute a query matrix with a dimension of n×d, and the dictionary library is m 1×d vectors; The first signal feature is calculated by weighted summing the query matrix through the feature weight, and then the first signal feature corresponding to each image is obtained. The calculation formula of the first signal feature is as follows: ; in, ; in, represents the first signal characteristic, represents the query matrix row vector, represents the query matrix The feature weights corresponding to the row vectors, The first element value; The data in each message and the radio signal data are input into an original signal encoder to obtain a second signal feature corresponding to each data, and the original signal encoder has a multi-layer perception structure; The first feature extraction structure is a convolutional neural network.

2. The stereoscopic monitoring perception method of rich media data according to claim 1, characterized in that: In the preprocessor, the radio signal data is firstly channelized to obtain a preset number of independent low-rate baseband sub-channel signals, and then each low-rate baseband sub-channel signal is fast Fourier transformed to obtain a two-dimensional time-frequency image.

3. The stereoscopic monitoring perception method of rich media data according to claim 1, characterized in that: The feature fusion module includes a shallow feature fuser and a deep feature fuser connected in sequence from data input to output direction; In the shallow feature fuser, the first signal feature and the second signal feature corresponding to each information or radio signal data in the monitored data are spliced ​​according to the feature channel dimension to obtain a first shallow fusion feature, and the corresponding first signal feature and the second signal feature are spliced ​​according to the feature vector dimension to obtain a second shallow fusion feature; Pass each first shallow fusion feature through the first mapping module to obtain the first feature, pass each second shallow fusion feature through the second mapping module to obtain the second feature, then add the corresponding first feature and second feature to obtain the third feature, concatenate all the third features, and then input them into the deep feature fuser to obtain the signal fusion feature; The first mapping module and the second mapping module are both two fully connected layers connected in sequence.

4. The stereoscopic monitoring perception method of rich media data according to claim 3, characterized in that: The deep feature fuser is a multi-layer self-attention mechanism structure. Each layer of the self-attention mechanism structure performs attention mechanism calculation on the input of the layer, and the calculation formula is as follows: ; in, Indicates The output of the layer self-attention mechanism structure, Indicates The output of the layer self-attention mechanism structure, Represents the dimension, express function, Indicates The query vector of the layer, Indicates The key vector of the layer, Indicates The value vector of the layer, Indicates transpose.

5. The stereoscopic monitoring perception method of rich media data according to claim 1, characterized in that: The feature alignment module and the feature mapper are both multiple fully connected layers connected in sequence.

6. The stereoscopic monitoring perception method of rich media data according to claim 1, characterized in that: When training the stereo monitoring perception model, the loss function of the situation description and the loss function of the target recognition result are weighted to obtain the total loss, and the parameters of the stereo monitoring perception model are back-propagated and updated based on the total loss.

7. The stereoscopic monitoring perception method of rich media data according to claim 1, characterized in that: The training data set for training the stereo monitoring perception model includes ultraviolet light information, visible light information, infrared light information and radio signal data at the same time and in the same area.

Citation Information

Patent Citations

  • Intelligent cabin analysis method and system based on multi-modal data fusion and analysis

    CN118536069A

  • All-time multi-modal pedestrian re-identification method based on simulation augmentation and prototype learning

    CN118799919A