A multi-source information fusion method based on broadband and narrowband communication
By adopting a wide-narrow band communication method in multi-source information fusion, combined with Bayes sparse representation and attention mechanism, the problem of low accuracy of multi-source information fusion is solved, and high-precision multi-source information fusion and high-quality transmission of semantic information are achieved.
Patent Information
- Application Number
- CN202411937676.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-26
AI Technical Summary
In the prior art, multi-source information fusion accuracy is low, making it difficult to effectively fuse heterogeneous data, forming an accurate understanding of target events or objects and panoramic perception.
Using a multi-source information fusion method based on wide-narrow band communication, the channel is divided adaptively and dynamically allocated, combined with Bayes sparse representation and reconstruction, attention mechanism and other technologies, the deep semantic features of text and image data are extracted and correlated and fusion is performed.
The fusion accuracy of multi-source information is improved, ensuring high-quality transmission of key semantic information and efficient transmission of full data is achieved, significantly reducing fusion errors and improving fusion accuracy.
Smart Images

Figure CN119377892B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data fusion, and in particular to a multi-source information fusion method based on wide- and narrow-band communications. Background Art
[0002] With the rapid development of new generation information technologies such as the Internet of Things, mobile Internet, and social networks, multi-source heterogeneous data such as text, images, audio, and video are growing explosively. These multi-source information from different channels, formats, and modalities contain rich semantic information. How to effectively fuse these heterogeneous data to form an accurate understanding and panoramic perception of the target event or object has become a key issue that needs to be solved in the era of big data. Multi-source information fusion technology has broad application prospects in many fields such as smart cities, autonomous driving, intelligent security, and public opinion analysis.
[0003] However, due to the huge differences in the format, modality, and semantic representation of multi-source data, the uneven data quality, and the different requirements for the credibility, importance, and timeliness of data from different sources, multi-source information fusion has brought great challenges. Traditional multi-source fusion methods, such as weighted average, decision tree, support vector machine, etc., mainly fuse from the shallow feature or decision level, making it difficult to mine deep semantic information and the fusion accuracy is not high. Summary of the invention
[0004] In view of the problem of low multi-source information fusion accuracy in the prior art, the present application provides a multi-source information fusion method based on wide- and narrow-band communications, which improves the multi-source information fusion accuracy by adaptively dividing priorities and dynamically allocating channels, as well as Bayes sparse representation and reconstruction.
[0005] The purpose of this application is achieved through the following technical solutions.
[0006] The present application provides a multi-source information fusion method based on broadband and narrowband communication, comprising: preprocessing the acquired multi-source heterogeneous data; the multi-source heterogeneous data includes text data and image data; using a pre-trained BERT network to extract text feature vectors of the pre-processed text data , use the pre-trained convolutional neural network ResNET to extract the image feature vector of the pre-processed image data ; Use the pre-trained bidirectional gated recurrent network BiGRU to obtain text feature vectors Priority and the image feature vector Priority ; Adaptively set text priority threshold and image priority threshold ; For priority Greater than or Greater than The text data or image data is transmitted through orthogonal frequency division multiple access (OFDMA) via a wideband channel; conversely, it is transmitted through code division multiple access (CDMA) via a narrowband channel; at the receiving end, the text data or image data received via the wideband channel is decoded by OFDMA, and the text data or image data received via the narrowband channel is decoded by CDMA to obtain decoded text data and image data respectively; the decoded text data and image data are sparsely represented using Bayes estimation, and the data are reconstructed by minimizing the L1 norm to obtain reconstructed text data and image data; the restored text data and image data are associated and fused using the attention mechanism to obtain a multi-source information fusion result.
[0007] Further, the pre-trained BERT network is used to extract the text feature vector of the pre-processed text data, and the pre-trained convolutional neural network ResNET is used to extract the image feature vector of the pre-processed image data, including: inputting the pre-processed text data into the pre-trained BERT network, calculating the similarity between each word in the text data through the self-attention mechanism in the BERT network, and extracting the semantic features of the text data; performing feature transformation on the text data through the feedforward neural network in the BERT network, and extracting the context features of the text data; wherein the context features reflect the contextual association relationship of the text data; and using the gating mechanism in the BERT network to fuse the extracted semantic features and context features to obtain the fused text feature vector. ; Input the preprocessed image data into the pre-trained convolutional neural network ResNET, and extract the local features of the image data through the convolutional layer in ResNET; Among them, the local features reflect the texture and edge information of the image data; Use the attention mechanism to calculate the similarity between local features, and perform weighted fusion on the local features to obtain the global features that reflect the overall semantic information of the image data; Through the residual connection, add the local features and the global features element by element to obtain the image feature vector .
[0008] Among them, local features refer to the local descriptors of image data extracted by the convolutional layer in the convolutional neural network ResNET. Through the convolution operation, the convolutional layer can extract detailed information such as local texture, edge, shape, etc. in the image data. These local features reflect the local patterns and structural characteristics of the image data in different regions and scales. In ResNET, each convolutional layer scans the local receptive field of the input image by sliding the window, and uses the convolution kernel to extract features from the local area. The parameters in the convolution kernel are obtained through training and learning, and can capture the local patterns in the image. By stacking multiple convolutional layers layer by layer, ResNET can extract hierarchical local features from the image data, from low-level edges and textures to high-level local structures and components. Global features refer to the feature representation that reflects the overall semantic information of the image data obtained by weighted fusion of local features in the convolutional neural network ResNET. Global features contain high-level semantic concepts and content information of image data, and have stronger abstraction and generalization capabilities. The local features are weighted fused through the attention mechanism to obtain global features. The attention mechanism calculates the similarity between local features and assigns different weights to different local features, highlighting the local features that are more important for overall semantic understanding and suppressing redundant or noisy features. This weighted fusion process can adaptively extract key information from image data and form a high-level representation of the overall content.
[0009] Furthermore, the pre-trained bidirectional gated recurrent network BiGRU is used to obtain the text feature vectors Priority and the image feature vector Priority , including: inputting the text feature vector into the pre-trained gated recurrent network GRU, the gated recurrent network GRU includes an update gate, a reset gate and an output gate; the update gate calculates the weighted sum of the hidden state output by the GRU at the previous moment and the text feature vector at the current moment through the sigmoid function to control the degree of transmission of the hidden state at the previous moment to the current moment; the reset gate calculates the weighted sum of the hidden state output by the GRU at the previous moment and the text feature vector at the current moment through the sigmoid function to control the degree of influence of the hidden state at the previous moment on the current moment; the GRU updates the hidden state according to the output of the update gate and the reset gate to obtain the hidden state at the current moment; using the multi-head attention mechanism, the text feature vector is divided into multiple sub-features, the attention weights between the sub-features and the hidden state at the current moment are calculated respectively, and the hidden state at the current moment is weightedly fused to obtain the multi-head attention vector of the text feature vector; through the output gate, the multi-head attention vector of the text feature vector is spliced with the hidden state at the current moment, and mapped to the low-dimensional space through the fully connected layer to obtain the priority vector of the text feature vector ; Input the image feature vector into the pre-trained gated recurrent network GRU, repeat the process, and obtain the priority vector of the image feature vector ; Calculate the priority vectors of the text feature vectors respectively and the priority vector of the image feature vector The L2 norm of and .
[0010] Among them, the update gate is a gate structure in GRU that controls the degree of transmission of the hidden state of the previous moment to the current moment. It calculates the weighted sum of the hidden state output by the GRU at the previous moment and the text feature vector at the current moment through the sigmoid function to obtain a scalar value between 0 and 1. The reset gate is a gate structure in GRU that controls the degree of influence of the hidden state of the previous moment on the current moment. It calculates the weighted sum of the hidden state output by the GRU at the previous moment and the text feature vector at the current moment through the sigmoid function to obtain a scalar value between 0 and 1. The output gate is a gate structure in GRU used to generate the output at the current moment. It concatenates the multi-headed attention vector of the text feature vector with the hidden state at the current moment, and maps it to a low-dimensional space through a fully connected layer to obtain a priority vector of the text feature vector. Low-dimensional space refers to mapping a high-dimensional feature vector to a relatively low-dimensional representation space through a fully connected layer. In this application, the multi-headed attention vector of the text feature vector is concatenated with the hidden state at the current moment, and then mapped to a low-dimensional space through a fully connected layer to obtain a priority vector of the text feature vector. The purpose of mapping high-dimensional features to low-dimensional space is to reduce the redundancy of feature representation and extract the most critical and discriminative information. The feature vector in low-dimensional space is more compact and efficient, while retaining the main information of the original features. Through dimensionality reduction mapping, the computational complexity of subsequent processing can be reduced and the generalization ability and robustness of the model can be improved.
[0011] Furthermore, adaptively set the text priority threshold and image priority threshold , including: using the extracted text feature vector and the image feature vector , the information entropy of the current text data and image data is calculated as follows and : , ,in, is the text feature vector, is the image feature vector; for The mean vector of for The mean vector of ; is the weight coefficient of the information entropy gain rate, which is used to balance the influence of information entropy and feature difference on semantic importance, and its value range is [0, 1]. Specifically, when calculating information entropy, and When considering not only the eigenvector and The distribution of the feature vector and the difference term between the feature vector and its mean vector are also introduced. Traditional information entropy only measures the distribution of feature vectors, but cannot reflect the differences within the feature vector. After the difference term is introduced, when the variance within the feature vector is large, the information entropy will increase accordingly. This shows that the greater the difference within the feature vector, the more information it contains, and the higher the weight should be given in the semantic importance assessment. By introducing the mean vector, the degree of deviation of each element in the feature vector from the overall mean can be calculated. The greater the degree of deviation, the more the element can reflect the particularity of the feature, the greater its contribution to the semantic distinction, and therefore the higher the semantic importance. Introducing the weight coefficient , the relative importance of information entropy and difference terms can be flexibly adjusted. When is large, the influence of difference on semantic importance is dominant; when When is small, the influence of feature distribution on semantic importance is dominant. The settings can adapt to the characteristics of different data.
[0012] according to and , the semantic importance index is calculated by the following formula and : , ,in, and The calculated and The maximum value of is the smoothing factor of the semantic importance index, which is used to prevent the denominator from being 0, and the value range is (0, 1); specifically, different texts and images may have different information entropies, and direct comparison is difficult to reflect the relative size of semantic importance. By dividing by the maximum value of their respective information entropy and , which can eliminate the influence of numerical scale and convert the semantic importance index and Unified to the interval [0, 1] for easy comparison. A smoothing factor is introduced in the denominator during normalization , takes values between (0, 1). When the information entropy of some data is small, or Near 0 o'clock, This can prevent the denominator from being too small. or Too large. The size of reflects the degree of scaling of the semantic importance index. Will compress and difference.
[0013] Get the generation timestamp of the current text data and image data and ; Calculate the time difference between the generated timestamp of text data and image data and the current system time T respectively and : , ,in, The timestamp for generating text data. is the timestamp of the image data, T is the current system time, and the time unit is seconds; according to the calculated time difference and , calculate the timeliness index and :when Less than the time difference threshold hour, ,on the contrary, ;in, is the decay rate of text data timeliness; when Less than the time difference threshold hour, ,on the contrary, ;in, is the time-dependent decay rate of image data.
[0014] Specifically, a time difference threshold is introduced , the time difference and Compared with the threshold, when the time difference exceeds the threshold, the timeliness index is directly assigned a value of 0. When the difference between the generation time of the data and the current time is too large, its timeliness is already very low, and it is meaningless to continue to assign it a non-zero timeliness index. Introducing a time difference threshold can clearly distinguish whether the data is time-sensitive, avoid invalid calculations, and improve efficiency. Time difference threshold It can be set according to specific application requirements. For applications with high timeliness requirements, such as real-time control, abnormal alarm, etc., For applications with relatively loose timeliness requirements, such as historical data analysis and statistical reports, The threshold is set larger. The adjustability of the threshold improves the flexibility of the timeliness index calculation, making it adaptable to different application scenarios. The exponential function is used to describe the decay law of the timeliness index with the time difference. The exponential function is a common decay function that can well describe the characteristics of the rapid decay of timeliness with the increase of time difference. or When it is small, the timeliness index is close to 1; as the time difference increases, the timeliness index decays exponentially, and the decay speed is fast at first and then slow, which is consistent with the actual situation. and to control the rate of exponential decay. and The larger the value of , the faster the timeliness index decays. This adjustability allows the calculation of timeliness index to adapt to the timeliness characteristics of different types of data. For text or image data with high timeliness requirements, or If the value is set larger, the decay of the timeliness index will be accelerated. For data with relatively loose timeliness requirements, or Set it smaller to slow down the decay of the timeliness indicator.
[0015] According to the semantic importance index and , and timeliness indicators and , the priority index is calculated by the following formula and : , ,in, and is the weight coefficient of semantic importance and timeliness, and its value range is [0, 1]; σ is the priority softening factor;
[0016] Furthermore, the priority softening factor σ is calculated by the following formula: ,in, and are the average values of the semantic importance index and the timeliness index respectively; the value range of σ is (0, 1), which is used to soften the comprehensive priority index to make its distribution more balanced.
[0017] Specifically, different application scenarios place different emphasis on semantic importance and timeliness. By introducing weight coefficients, the proportion of semantic importance and timeliness in the priority index can be flexibly adjusted according to specific application requirements. The priority index obtained directly using weighted fusion may not be distributed evenly, and the priority index of some data may be too high or too low, resulting in extreme transmission strategies. The introduction of a priority softening factor can perform a nonlinear transformation on the priority index to make its distribution more balanced and avoid extreme situations. The priority softening factor σ uses a sigmoid function to map the result of weighted fusion to the interval (0, 1). When When the value of is large, that is, when both the semantic importance index and the timeliness index are higher than the average level, σ is close to 1, and the softening degree of the priority index is small; when When the value of is small, that is, when both the semantic importance index and the timeliness index are lower than the average level, σ is close to 0, and the degree of softening of the priority index is large. This adaptive softening mechanism can reduce the extreme difference of the priority index while maintaining the relative size relationship of the priority index, making the transmission strategy more balanced.
[0018] Furthermore, the priority Greater than or Greater than The text data or image data is transmitted by using a broadband channel through orthogonal frequency division multiple access OFDMA; conversely, it is transmitted by using a narrowband channel through code division multiple access CDMA, including: determining the text feature vector Priority Is it greater than or equal to the text priority threshold? , determine the image feature vector Priority Is it greater than or equal to the image priority threshold? ;if Greater than or equal to or Greater than or equal to , the corresponding text data or image data is divided into the first data, otherwise it is divided into the second data; the first data is transmitted through orthogonal frequency division multiple access OFDMA via a wideband channel; the second data is transmitted through code division multiple access CDMA via a narrowband channel.
[0019] Further, the first data is transmitted using a broadband channel through orthogonal frequency division multiple access OFDMA, including: or As a waterline, the signal-to-noise ratios of the subcarriers in the OFDMA channel are sorted, and the signal-to-noise ratios of the sorted subcarriers are used as the heights of the subcarriers; unassigned subcarriers are selected in order from high to low according to the waterline until the number of selected subcarriers reaches a preset number; wherein the preset number is related to the priority of the first data or Positive correlation; modulate the first data onto the allocated subcarrier, modulate the first data using quadrature amplitude modulation QAM, and the modulation order and priority or Positive correlation; transmitting the modulated first data through a broadband channel.
[0020] Specifically, by mapping the priority to a waterline, the importance of the data can be intuitively indicated. The higher the waterline, the higher the priority of the data, and more resource support should be given when allocating subcarriers. By sorting the signal-to-noise ratios of the subcarriers and using the sorted signal-to-noise ratios as the height of the subcarriers, the channel quality of the subcarriers can be vividly described. The higher the signal-to-noise ratio, the better the channel condition of the subcarrier and the higher the transmission quality. Selecting subcarriers from high to low according to the waterline until the preset number is reached can ensure that high-priority data obtains more high-quality subcarriers and improves transmission reliability; at the same time, the number of subcarriers is positively correlated with the priority, which can provide a larger transmission bandwidth for important data and reduce transmission delay.
[0021] Specifically, by positively correlating the modulation order with the priority, a higher-order modulation method, such as 64QAM or 256QAM, can be provided for important data to improve spectrum utilization and increase transmission rate; at the same time, a relatively low-order modulation method, such as 16QAM or QPSK, can be provided for less important data to improve anti-interference ability and ensure transmission quality. Adaptive modulation technology can dynamically adjust the modulation order according to the importance of the data and the channel conditions, achieve the optimal balance between transmission efficiency and transmission reliability, and improve the adaptability and flexibility of the system. OFDMA technology can make full use of frequency domain resources and improve spectrum utilization by dividing the broadband channel into multiple orthogonal subcarriers. Different users or different data streams can occupy different subcarriers respectively to achieve multiple access in the frequency domain and reduce interference. Compared with traditional frequency division multiple access FDMA, the subcarriers in OFDMA are orthogonal and compactly arranged, with higher spectrum efficiency; compared with time division multiple access TDMA, OFDMA can flexibly allocate subcarriers to different users or data streams according to the priority of the data and the channel conditions, and has stronger adaptability. OFDMA technology combined with adaptive modulation technology can adjust the modulation order according to the channel status at the subcarrier level, further improving spectrum utilization and transmission reliability.
[0022] Further, the second data is transmitted by using a narrowband channel through code division multiple access CDMA, including: according to the priority corresponding to the second data or , determine the spreading factor SF of the code division multiple access CDMA channel, the spreading factor SF and the priority or anti-correlation; according to the spreading factor SF, the second data is spread using the pseudo-random code corresponding to the code channel of the code division multiple access CDMA; wherein the number of chips of the pseudo-random code is equal to the spreading factor SF; the second data after spreading is pseudo-randomly distributed on the code channel; the second data after spreading is power allocated, and the allocated power P is: or ,in, is the maximum transmission power of the code division multiple access CDMA channel; the second data after power allocation is mapped to the code channel of the code division multiple access CDMA; the second data on the code channel is modulated by orthogonal phase shift keying QPSK, and the modulated second data is transmitted through a narrowband channel.
[0023] Specifically, by inversely correlating the spreading factor SF with the priority, a larger spreading factor can be allocated to data with lower importance, thereby improving the data's anti-interference ability and transmission reliability; at the same time, a relatively smaller spreading factor can be allocated to data with higher importance, thereby reducing data transmission delay and improving transmission efficiency. The adaptive adjustment of the spreading factor can achieve the optimal balance between anti-interference and transmission efficiency based on the importance of the data and the channel conditions. This adaptive mechanism improves the system's adaptability to multi-source heterogeneous data and helps meet the service quality requirements of different data. Through the formula or , according to the priority of the data, its transmission power can be dynamically adjusted. The higher the priority, the greater the allocated power, the higher the transmission signal-to-noise ratio, and the stronger the transmission reliability; the lower the priority, the smaller the allocated power, the lower the transmission signal-to-noise ratio, and the relatively weaker the transmission reliability. By allocating orthogonal or low-correlated pseudo-random codes to different users or data streams, multiple data can be transmitted concurrently in the same frequency band and the same time slot. This code division multiplexing feature can significantly improve the spectrum utilization and transmission capacity of the system. Through spread spectrum and pseudo-random code mapping, CDMA technology can expand narrowband data to broadband and improve anti-interference capabilities; at the same time, different data are pseudo-randomly distributed on the code channel, and the interference between each other is small, which can support the concurrent transmission of more users or data streams. Combined with orthogonal phase shift keying QPSK modulation, it can make full use of code domain and phase domain resources to further improve spectrum utilization and data transmission rate.
[0024] Further, the reconstructed text data and image data are obtained, including: dividing the decoded text data into a plurality of text data blocks, dividing the decoded image data into a plurality of image data blocks; calculating the optimal sparse representation coefficient of each text data block and image data block in a preset overcomplete dictionary D : ,in, is the optimal sparse representation coefficient, which is a column vector whose number of elements is equal to the number of atoms in the dictionary D; is the sparse representation coefficient, which is a column vector representing the representation coefficient of the original signal y in the dictionary D; y is the text data block or image data block to be sparsely represented, which is a column vector; λ is the regularization parameter of the sparse representation; the orthogonal matching pursuit OMP algorithm is used to minimize the L1 norm of the sparse representation coefficient α of each data block to obtain the reconstructed sparse representation coefficient ; Using the reconstructed sparse representation coefficients , each text data block and image data block is reconstructed through the preset over-complete dictionary D to obtain the reconstructed text data block and image data blocks : , ; The reconstructed text data block and image data blocks Splice in the original order to get the reconstructed complete text data and image data .
[0025] Among them, the overcomplete dictionary D refers to a matrix D whose number of columns (the number of atoms in the dictionary) is greater than the number of rows (the dimension of the original signal). The overcomplete dictionary D is used to sparsely represent text data blocks and image data blocks. Each column in the overcomplete dictionary D is called an atom, and atoms are the basic building blocks in the dictionary. Atoms are usually some simple signals or basis functions, such as wavelets, Gabor functions, DCT basis, etc. Complex signals can be represented by linear combinations of these atoms. The overcompleteness of the dictionary D means that the number of atoms in the dictionary is much larger than the dimension of the original signal. This redundancy enables the dictionary D to provide more representation freedom, so that the original signal y can be sparsely represented with a small number of atoms in the dictionary D. The overcomplete dictionary D provides a framework for signal representation, so that the signal can have multiple sparse representation methods under this framework. In this application, in the sparse representation process, by minimizing the reconstruction error and the L1 norm, the optimal sparse representation coefficient α* of the original signal y under the overcomplete dictionary D can be found. Among them, the reconstruction error term ensures that the sparse representation Dα is as close to the original signal y as possible, and the L1 norm term makes the representation coefficient α as sparse as possible, that is, only a few atoms are involved in the representation of the signal.
[0026] Furthermore, the reconstructed sparse representation coefficients , solved by the following formula: , constraints: , where β is the optimization variable in the optimization problem, which represents the sparse representation coefficient that minimizes the objective function under the constraints; express The L1 norm distance between and β; represents the L0 norm of β; is the sparsity weight coefficient; is the reconstruction error threshold; Represents the reconstructed data The L2 norm error between the original data y is used to measure the reconstruction quality.
[0027] Specifically, by introducing the L1 norm term, the reconstructed sparse representation coefficient β can be kept as close as possible to the optimal sparse representation coefficient during the reconstruction of the sparse representation coefficient. This helps to maintain the original sparse structure while improving the reconstruction accuracy and reduce the additional noise or distortion introduced in the reconstruction process. The L1 norm is sparse-inducing, that is, it tends to produce sparse solutions during the optimization process. Taking the L1 norm term as part of the objective function can make the reconstructed sparse representation coefficient β inherit the optimal sparse representation coefficient The sparsity of the image can be improved to avoid excessive densification and improve reconstruction efficiency.
[0028] The L0 norm of β, that is, the number of non-zero elements in β, is used to measure the sparsity of the reconstructed sparse representation coefficient β; specifically, compared with the L1 norm, the L0 norm is a true indicator of the number of non-zero elements in the vector. Incorporating the L0 norm term into the objective function can more directly and accurately control the sparsity of the reconstructed sparse representation coefficient β, ensuring the high compressibility of the reconstruction result. is the sparsity weight coefficient, which is a positive real number used to balance the relative importance of the L1 norm term and the L0 norm term, and control the sparsity of the reconstructed sparse representation coefficient β; specifically, the relative importance of the L1 norm term and the L0 norm term can be flexibly adjusted through the sparsity weight coefficient η. When η is large, the reconstruction process will prefer sparsity and produce fewer non-zero elements; when η is small, the reconstruction process will prefer consistency with the initial sparse representation coefficient. This adjustability improves the flexibility of the reconstruction process, enabling it to adapt to different application requirements.
[0029] is the reconstruction error threshold, which is a positive real number, indicating the maximum error allowed between the reconstructed data and the original data; y is the text data block or image data block to be reconstructed, which is a column vector; D is the preset overcomplete dictionary, which is a matrix with more columns than rows, so that the data y can be linearly represented by a small number of atoms in the dictionary D; Represents the reconstructed data The L2 norm error between the reconstructed data and the original data y is used to measure the reconstruction quality. Specifically, through the reconstruction error constraint, we can take into account the reconstruction quality while pursuing sparsity. The constraint condition ensures the fidelity between the reconstructed data and the original data and avoids information loss caused by over-compression. , which means ensuring that the error between the reconstructed data and the original data does not exceed the threshold ε.
[0030] Specifically, the reconstruction error threshold ε can be set according to the fault tolerance requirements of the specific application. For scenarios with high reconstruction quality requirements, ε can be set to a smaller value; for scenarios with relatively loose reconstruction quality requirements, ε can be set to a larger value. The adjustability of the threshold improves the adaptability of the reconstruction process.
[0031] Compared with the prior art, the advantages of this application are:
[0032] Pre-trained BERT and ResNet are used to extract deep semantic features of text and image data, and global feature representation is obtained through attention mechanisms, reducing the semantic loss of fused input.
[0033] The priority threshold is dynamically set according to the semantic importance and timeliness indicators, and the data with large semantic information and high timeliness requirements are classified as high priority, ensuring the transmission quality of key data that has a great impact on the fusion results and avoiding the semantic loss caused by average transmission.
[0034] OFDMA broadband transmission is used for high-priority data, and subcarriers and modulation orders are adjusted adaptively to ensure the transmission quality of key semantic information; CDMA narrowband transmission is used for low-priority data, and the spreading factor and power are adjusted adaptively to balance transmission efficiency and resource utilization. Dynamic channel allocation helps to achieve high-quality transmission of semantically important data and efficient transmission of full data under limited bandwidth.
[0035] By using Bayes estimation and OMP algorithm, the received data is sparsely represented and reconstructed with L1 norm on the overcomplete dictionary, overcoming the data loss and distortion problems during transmission and restoring the original semantic features of the data to the greatest extent. Fusion based on the reconstructed high-quality data can significantly reduce the fusion error and improve the fusion accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present application will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:
[0037] Figure 1 is an exemplary flow chart of a multi-source information fusion method based on wide- and narrow-band communication shown in the present application;
[0038] Figure 2 is an exemplary flow chart of generating priority thresholds as shown in the present application;
[0039] Figure 3 is an exemplary flow chart of data transmission shown in this application;
[0040] Figure 4 is an exemplary flow chart of reconstructing data as shown in this application. DETAILED DESCRIPTION
[0041] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0042] Figure 1 This is an exemplary flow chart of a multi-source information fusion method based on broadband and narrowband communication shown in the present application, including: preprocessing the acquired multi-source heterogeneous data; the multi-source heterogeneous data includes text data and image data; using a pre-trained BERT network to extract text feature vectors of the pre-processed text data , use the pre-trained convolutional neural network ResNET to extract the image feature vector of the pre-processed image data ; Use the pre-trained bidirectional gated recurrent network BiGRU to obtain text feature vectors Priority and the image feature vector Priority ; Adaptively set text priority threshold and image priority threshold ; For priority Greater than or Greater than The text data or image data is transmitted through orthogonal frequency division multiple access (OFDMA) via a wideband channel; conversely, it is transmitted through code division multiple access (CDMA) via a narrowband channel; at the receiving end, the text data or image data received via the wideband channel is decoded by OFDMA, and the text data or image data received via the narrowband channel is decoded by CDMA to obtain decoded text data and image data respectively; the decoded text data and image data are sparsely represented using Bayes estimation, and the data are reconstructed by minimizing the L1 norm to obtain reconstructed text data and image data; the restored text data and image data are associated and fused using the attention mechanism to obtain a multi-source information fusion result.
[0043] Specifically, the acquired multi-source heterogeneous data is preprocessed; the multi-source heterogeneous data includes text data and image data; specifically, the acquired multi-source heterogeneous data set is , where N is the number of data sources, Di represents the data subset obtained by the i-th data source. Among them, the text data subset is represented as , the image data subset is represented as , and are the number of samples for text and image data respectively. , the preprocessing process includes: text cleaning: removing Remove the non-ASCII characters, HTML tags, URL links and other noises in the text to get the cleaned text . Participle: will Split according to the basic vocabulary units of the language (such as English words, Chinese words) to obtain word sequences ,in for Remove stop words: Use the predefined stop word list to remove The high-frequency but meaningless words in the sentence, such as "the" and "a", are used to obtain the word sequence after removing the stop words. . Stemming: Using a stemming algorithm (such as Porter stemmer), The words in are restored to stem form to reduce the impact of word deformation. Get the stem sequence ,in is the number of words after removing stop words.
[0044] For image data The preprocessing process includes: image size normalization: The size of is adjusted to a fixed size, such as 256×256 pixels, to ensure the consistency of subsequent processing. The adjusted image is recorded as . Image normalization: The pixel value is divided by 255, and the pixel value range is mapped to [0, 1] to reduce the difference between images. The normalized image is recorded as . Image enhancement: Perform data enhancement operations such as random flipping, rotation, scaling, and cropping to expand the amount of training data and improve the generalization ability of the model. The enhanced image set is denoted as , where M is the enhancement factor of each image. After preprocessing, a normalized text dataset is obtained and image datasets ,in and They are the number of preprocessed text and image data samples, respectively, which serve as input for subsequent models.
[0045] Use the pre-trained BERT network to extract text feature vectors from pre-processed text data , use the pre-trained convolutional neural network ResNET to extract the image feature vector of the pre-processed image data ; Specifically, for the preprocessed text dataset , for each text sample ,Will Represented as a word embedding matrix ,in is the word embedding dimension. Each word is embedded in the word embedding layer of BERT. Mapped to a dimensional word vector .Will Enter BERT's self-attention layer and calculate The attention weight matrix between word vectors in . Elements represents the attention weight of the lth word to the mth word, reflecting the similarity and dependency between words. Application , and obtain the semantic feature matrix ,in It is the semantic feature dimension. Each line of is a semantic feature vector of a word, integrating the information of other words. Input BERT’s feedforward neural network and extract the context feature matrix through nonlinear transformation ,in is the context feature dimension. Each line of is a context feature vector of a word. and Input the gating unit of BERT, and fuse the semantic features and context features through the gating mechanism to obtain the final word feature matrix ,in is the word feature dimension. The row vectors are pooled (such as average pooling or maximum pooling) to obtain text samples The text feature vector .
[0046] For image datasets Each image sample in ,Will Input the convolution layer of ResNET, and extract the local feature matrix of the image through multi-layer convolution and pooling operations. ,in , and are the height, width and number of channels of the feature map respectively. Input the attention module, and obtain the weighted fusion global feature matrix by calculating the attention weights between local features. ,in is the number of channels of the global feature. and Through residual connection addition, the fused image feature matrix is obtained .right Perform global average pooling to obtain the image feature vector of image sample D'ik Finally, for the preprocessed text dataset , get its text feature vector set , for the preprocessed image dataset , and obtain its image feature vector set . Text feature vector and the image feature vector The high-level semantic information of text data and image data are represented respectively, which can be used for subsequent feature fusion and task learning.
[0047] Use the pre-trained bidirectional gated recurrent network BiGRU to obtain the priority pt of the text feature vector tt and the image feature vector Priority ; Includes: a given text feature vector set And the image feature vector set ,in represents the feature vector of the jth text sample, Represents the feature vector of the kth image sample. Using the pre-trained bidirectional gated recurrent network BiGRU, we process and , get the corresponding priority set and ,in and They represent the priorities of the j-th text sample and the k-th image sample respectively.
[0048] For text feature vector , input it into the forward and backward GRU networks of BiGRU to obtain the hidden state sequence and , where T is the sequence length, and Respectively represent the forward and backward hidden states at the tth moment, is the hidden state dimension. Taking the forward GRU as an example, the hidden state at the tth moment is The calculation process is as follows: Update gate: ,in, , , To update the gate parameters. Reset the gate: ,in , , To reset the gate parameters. Candidate hidden states: ,in , , is the candidate hidden state parameter. Hidden state: .
[0049] In obtaining and Then, the multi-head attention mechanism is used to calculate The attention vector .Will Divided into H sub-features , the dimension of each sub-feature is For the i-th sub-feature , calculate its and The attention weights for each hidden state in : , ,in , is the attention parameter matrix. Then, the attention vector of the i-th sub-feature is calculated: , Finally, the attention vectors of the H sub-features are concatenated to obtain The attention vector . By output keeper It is concatenated with the hidden states hfjT and hbjT at the last moment, and passes through the fully connected layer to obtain the priority vector :,in, , is the output gate parameter, is the priority vector dimension.
[0050] For the image feature vector , repeat the execution to get its priority vector .for and , respectively calculate their L2 norm as the priority: , Through the above steps, the BiGRU network is used to extract the sequence information of text and image features respectively, and combined with the multi-head attention mechanism, the priority vector representing the importance of the feature is obtained. Finally, the priority vector is mapped to a scalar priority value through the L2 norm to obtain the text feature vector set And the image feature vector set The corresponding priority set and The BiGRU network can make full use of the contextual information of the feature sequence, and the multi-head attention mechanism captures the different subspace representations of the feature vector. The combination of the two can better evaluate the importance of features and provide useful prior knowledge for subsequent feature fusion.
[0051] Figure 2 is an exemplary flow chart of generating a priority threshold as shown in the present application, and adaptively setting a text priority threshold and image priority threshold In this embodiment, there is a batch of text data and image data, where the text data contains 1000 text samples and the image data contains 800 image samples. For the jth text sample, extract its 512-dimensional text feature vector ; For the kth image sample, extract its 1024-dimensional image feature vector . Calculate the mean vector of the text feature vector ; Calculate the mean vector of the image feature vector : ; Set the weight coefficient of information entropy gain rate For each text feature vector , calculate its information entropy: For each image feature vector , calculate its information entropy: ; Calculate the maximum value of text information entropy : ; Calculate the maximum value of image information entropy : . Set the smoothing factor for the semantic importance indicator For each text feature vector , calculate its semantic importance index: For each image feature vector , calculate its semantic importance index: .
[0052] Get the generation timestamp of each text sample and the timestamp of each image sample’s generation , in seconds. Get the current system time T, T=1622476800 seconds, which is 0:0:0 on June 1, 2024. For each text sample, calculate the time difference between its generation timestamp and the current system time: For each image sample, calculate the time difference between its generation timestamp and the current system time: ; Set the time difference threshold for the timeliness index to decay to 0 seconds, which is 7 days. Set the text timeliness decay rate , image time-dependent decay rate For each text sample, calculate its timeliness index: , ; For each image sample, calculate its timeliness index: , .
[0053] Set the weight coefficients for semantic importance and timeliness . Calculate the average of the semantic importance index: . Calculate the average value of the timeliness index: For each text sample, calculate its priority softening factor : ; For each image sample, calculate its priority softening factor : For each text sample, calculate its priority index: ; For each image sample, calculate its priority index: ; Calculate the median Pt_median of the text priority index as the text priority threshold Pt: . Calculate the median Pi_median of the image priority index as the image priority threshold: .
[0054] Through the above steps, the text priority threshold is adaptively calculated. and image priority threshold Among them, the semantic importance index and The information entropy and feature difference of the feature vector are comprehensively considered, and the weight coefficient of the information entropy gain rate is used To balance the impact of the two. Timeliness index and The freshness of the data is measured using the time difference threshold (7 days) to determine whether the data is expired, and use the time-based decay rate and To control the decay speed of the index. Finally, by weighted fusion of semantic importance index and timeliness index (weight coefficient , ), and combined with the priority softening factor and , and obtain the comprehensive priority index and Considering that the distribution of priority indicators may be uneven, the median rather than the mean is used as the final priority threshold. and .
[0055] Figure 3 is an exemplary flow chart of data transmission shown in this application, with priority Greater than or Greater than The text data or image data is transmitted by using a broadband channel through orthogonal frequency division multiple access OFDMA; conversely, it is transmitted by using a narrowband channel through code division multiple access CDMA, including: determining the text feature vector Priority Is it greater than or equal to the text priority threshold? , determine the image feature vector Priority Is it greater than or equal to the image priority threshold? ;if Greater than or equal to or Greater than or equal to , the corresponding text data or image data is divided into the first data, otherwise it is divided into the second data.
[0056] For the first data, a broadband channel is used to transmit the first data through orthogonal frequency division multiple access OFDMA; including: obtaining a priority indicator of the first data ( or ), which is used as the waterline. The waterline reflects the priority of the first data. The higher the priority, the higher the waterline. Get the signal-to-noise ratio (SNR) of each subcarrier in the OFDMA channel and sort them. The OFDMA channel has N subcarriers, and their signal-to-noise ratios are . Sort the subcarriers in descending order according to the signal-to-noise ratio: . The signal-to-noise ratio of the sorted subcarriers is used as the height of the subcarrier. The higher the signal-to-noise ratio, the higher the height of the subcarrier. According to the watermark, unallocated subcarriers are selected in sequence from high to low until the number of selected subcarriers reaches the preset number M. The preset number M is positively correlated with the priority index of the first data: if the priority index is higher, M is larger, and more high signal-to-noise ratio subcarriers are selected; if the priority index is lower, M is smaller, and relatively fewer high signal-to-noise ratio subcarriers are selected. The process of selecting subcarriers is as follows: initialize the number of selected subcarriers count=0; from the sorted subcarriers, select the unallocated subcarriers with a height higher than the watermark in turn; each time a subcarrier is selected, count increases by 1; when count reaches the preset number M, stop selecting. The first data is modulated onto the M allocated subcarriers.
[0057] The first data is modulated using quadrature amplitude modulation QAM, and the modulation order L is positively correlated with the priority index: determine the value range of the priority index. Assume that the value range of the priority index is [0, 1], and the larger the value, the higher the priority. Different priority intervals are divided according to the value of the priority index. The value range of the priority index can be divided into three intervals: high priority interval: [0.7, 1], corresponding to high-order modulation (such as 64QAM); medium priority interval: [0.4, 0.7), corresponding to medium-order modulation (such as 32QAM); low priority interval: [0, 0.4), corresponding to low-order modulation (such as 16QAM). Obtain the priority index of the first data and determine the priority interval to which it belongs. Assuming that the priority index of the first data is p, then: if p∈[0.7, 1], the first data is of high priority, and high-order modulation (such as 64QAM) is selected; if p∈[0.4, 0.7), the first data is of medium priority, and medium-order modulation (such as 32QAM) is selected; if p∈[0, 0.4), the first data is of low priority, and low-order modulation (such as 16QAM) is selected. According to the priority index of the first data, the modulation order L of the orthogonal amplitude modulation QAM is determined.
[0058] The modulation order L is positively correlated with the priority index p: if p is higher, L is larger, high-order modulation is used, each symbol carries more bits, and the transmission rate is higher; if p is lower, L is smaller, low-order modulation is used, each symbol carries fewer bits, and the transmission rate is relatively low. The following corresponding relationship can be set: if p∈[0.7, 1], L=6, 64QAM modulation is used; if p∈[0.4, 0.7), L=5, 32QAM modulation is used; if p∈[0, 0.4), L=4, 16QAM modulation is used. Using the determined modulation order L, orthogonal amplitude modulation is performed on the first data. The process of QAM modulation is as follows: the first data is divided into L-bit symbols; according to the Gray mapping rule, each L-bit symbol is mapped into a complex number, and the real part and the imaginary part represent two orthogonal components respectively; using different amplitudes and phases, the mapped complex number is modulated onto the corresponding subcarrier. The modulated first data is transmitted through a broadband channel. Compared with narrowband channels, broadband channels have larger bandwidth and higher transmission rates, and are more suitable for transmitting high-priority data.
[0059] The second data is transmitted by using a narrowband channel through code division multiple access CDMA; including: obtaining a priority indicator corresponding to the second data ( or ). The priority index reflects the importance of the second data. The larger the value, the lower the priority. According to the priority index, the spreading factor SF of the code division multiple access CDMA channel is determined. The spreading factor SF is inversely correlated with the priority index: if the priority index is high, the SF is small and the spreading degree is low; if the priority index is low, the SF is large and the spreading degree is high. The following corresponding relationship can be set: if or , then SF=64; if or , then SF=128; if or , then SF=256. According to the determined spreading factor SF, the second data is spread using the pseudo random code corresponding to the code channel of the code division multiple access CDMA. The number of chips of the pseudo random code is equal to the spreading factor SF, and the value of the chip is {+1,-1}.
[0060] The process of spectrum spreading is as follows: obtain the bit sequence of the second data and the spreading factor SF of the code division multiple access CDMA channel. Assume that the bit sequence of the second data is , the spreading factor is SF. Generate a pseudo-random code sequence of length SF. The pseudo-random code sequence is , where each code chip has a value of {+1, -1}. Different code channels use different pseudo-random code sequences to ensure orthogonality between code channels. Spread spectrum for each bit of the second data. For the i-th bit , multiplied by SF chips of the pseudo-random code sequence: if , then the pseudo-random code sequence takes the original value: ;if , then the pseudo-random code sequence is inverted: . Splice the spread data to form a second spread data sequence. The second spread data sequence has a length of n·SF, and the data energy is spread to a wider frequency band. Map the second spread data to the code channel. Since a pseudo-random code sequence is used for spreading, the second spread data is pseudo-randomly distributed on the code channel.
[0061] Power allocation process: Obtain the priority index corresponding to the second data ( or ). The priority index reflects the importance of the second data, and a larger value indicates a lower priority. According to the priority index, the transmission power P allocated to the second data is calculated. If the priority index is (text data), then ; If the priority index is (image data), then .in, is the maximum transmission power of the code division multiple access CDMA channel, and are the priority thresholds for text data and image data, respectively. According to the above formula, the higher the priority index (the lower the priority), the greater the allocated power; the lower the priority index (the higher the priority), the smaller the allocated power. Multiply the allocated power P by the second data sequence after spread spectrum. For each element d of the second data sequence after spread spectrum, multiply by the power factor : The data sequence after power allocation is . The second data sequence after output power allocation.
[0062] The principle of power allocation is: the higher the priority index, the lower the allocated power; the lower the priority index, the higher the allocated power. The second data after power allocation is mapped to the code channel of code division multiple access CDMA. The number of code channels is equal to the spreading factor SF, and each code channel corresponds to a pseudo-random code. The mapping process is to multiply the second data after power allocation with the corresponding pseudo-random code. The second data on the code channel is modulated by orthogonal phase shift keying QPSK. The process of QPSK modulation is as follows: the data on the code channel is grouped in pairs, with 2 bits in each group; according to the constellation diagram of QPSK, each group of 2 bits is mapped into a complex number, and the real part and the imaginary part represent two orthogonal components respectively; using different phases, the mapped complex number is modulated onto the carrier. The second data after QPSK modulation is transmitted through a narrowband channel. Compared with a broadband channel, a narrowband channel has a narrower bandwidth and a lower transmission rate, but has a stronger anti-interference ability and is more suitable for transmitting low-priority data.
[0063] Figure 4 This is an exemplary flow chart of reconstructed data shown in the present application. At the receiving end, OFDMA decoding is performed on the text data or image data received by the broadband channel, and CDMA decoding is performed on the text data or image data received by the narrowband channel to obtain decoded text data and image data respectively; OFDMA decoding is performed on the text data or image data received by the broadband channel: OFDMA modulated text data or image data transmitted by the broadband channel is received. Assume that the received data is an OFDMA symbol sequence , where N is the number of subcarriers. Perform serial-to-parallel conversion on the received OFDMA symbol sequence. Convert the OFDMA symbol sequence into parallel subcarrier data streams Perform FFT (Fast Fourier Transform) on the data on each subcarrier. Using the FFT algorithm, the subcarrier data stream in the time domain is converted into a data sequence in the frequency domain. . Perform subcarrier demapping on the frequency domain data sequence. According to the OFDMA subcarrier allocation scheme, demap the data on each subcarrier to the corresponding user data. Assume that the subcarrier set occupied by the text data is , the subcarrier set occupied by the image data is . Demodulate the demapped user data. According to the modulation method (such as QAM), demodulate the complex data on each subcarrier into bit data. Obtain the demodulated text data bit sequence and image data bit sequence Channel decode the demodulated bit sequence. Use the corresponding decoding algorithm of channel coding (such as Turbo code or LDPC code) to decode the bit sequence of text data and image data. Obtain the decoded text data and image data.
[0064] CDMA decoding of text data or image data received via a narrowband channel: Receive CDMA modulated text data or image data transmitted via a narrowband channel. Assume that the received data is a CDMA symbol sequence. , where Q is the number of symbols. Despread the received CDMA symbol sequence. Despread each symbol using the CDMA spreading code (pseudo-random code sequence). For each symbol , and the corresponding spreading code Perform inner product operation: ; Get the symbol sequence after despreading . Demodulate the despread symbol sequence. According to the modulation method (such as QPSK), demodulate each symbol into bit data. Obtain the demodulated text data bit sequence and image data bit sequence . Channel decode the demodulated bit sequence. Use the corresponding decoding algorithm of channel coding (such as convolutional code or Turbo code) to decode the bit sequence of text data and image data. Obtain the decoded text data and image data. Divide the decoded text data into N text data blocks according to certain rules The size of the data block can be determined based on factors such as the length of the text and the relevance of the content. For example, it can be divided according to a fixed number of characters or words, or according to the integrity of the semantics. The decoded image data is divided into M image data blocks according to certain rules. The size of the data block can be determined based on factors such as the size of the image, the distribution of local features, etc. For example, the data block can be divided according to a fixed pixel block size or according to a region of interest (ROI).
[0065] Calculate the sparse representation coefficient α of each text data block and image data block in the preset overcomplete dictionary D. For each text data block or image data block to be sparsely represented, represent it as a column vector y. For a text data block, it can be converted into a word frequency vector or a TF-IDF vector; for an image data block, it can be converted into a pixel value vector or a feature descriptor vector. Prepare a preset overcomplete dictionary D. The overcomplete dictionary D is a matrix whose number of columns is greater than the number of rows, so that the data can be represented by a linear combination of a small number of atoms (columns) in the dictionary. For text data, a dictionary based on the bag of words model can be used, such as a dictionary constructed by the K-SVD dictionary learning algorithm; for image data, a dictionary based on patches or convolutional features can be used, such as a DCT dictionary, a Gabor dictionary, etc. Model the sparse representation problem as an optimization problem: .
[0066] The objective function of the optimization problem consists of two parts: Denotes the reconstruction error, that is, the square of the L2 norm of the error of reconstructing the data y using the dictionary D and the sparse representation coefficient α. The goal is to minimize the reconstruction error so that the reconstructed data Dα is as close to the original data y as possible. Part II It represents the L1 norm of the sparse representation coefficient α, that is, the sum of the absolute values of each element of α, which is used to measure the sparsity of α. λ is the regularization parameter of sparse representation, which is used to balance the reconstruction error and sparsity. The larger the value, the sparser the α obtained, but the larger the reconstruction error may be; the smaller the value, the smaller the reconstruction error, but the sparsity of α may be lower.
[0067] The sparse representation problem is an NP-hard problem, and it is impossible to directly find the optimal solution. It is necessary to use an iterative optimization algorithm, such as the orthogonal matching pursuit (OMP) algorithm and the least angle regression (LARS) algorithm, to approximate the solution. This embodiment uses the orthogonal matching pursuit OMP algorithm to minimize the L1 norm of the sparse representation coefficient α of each data block to obtain the reconstructed sparse representation coefficient . For each text data block or image data block, prepare its corresponding initial sparse representation coefficient α, that is, the calculated sparse representation coefficient. Prepare a preset overcomplete dictionary D, which is the same as the dictionary D used. Set the sparsity weight coefficient η and the reconstruction error threshold ε. Objective function: min {∥α-β∥1 + η∥β∥0}. The objective function consists of two parts: L1 norm term ∥α-β∥1: measures the difference between the reconstructed sparse representation coefficient β and the initial sparse representation coefficient α. Minimize this term so that the reconstructed sparse representation coefficient β is as close to the initial sparse representation coefficient α as possible. L0 norm term η∥β∥0: measures the sparsity of the reconstructed sparse representation coefficient β. Minimize this term so that the reconstructed sparse representation coefficient β is as sparse as possible, that is, the number of non-zero elements is as small as possible. Sparsity weight coefficient η: balances the relative importance of the L1 norm term and the L0 norm term, and controls the sparsity of the reconstructed sparse representation coefficient β. Constraint: ∥y - Dβ∥2 ≤ ε; The constraint ensures that the L2 norm error between the reconstructed data Dβ and the original data y does not exceed the reconstruction error threshold ε, ensuring the reconstruction quality.
[0068] The orthogonal matching pursuit (OMP) algorithm is used to solve the optimization problem. In each iteration, the dictionary atom most relevant to the residual is selected, added to the active set, and the residual is updated until the stop condition is met. The main steps of the OMP algorithm are as follows: Initialization: The residual vector r = y represents the current reconstruction error. Active set , represents the index set of the selected dictionary atoms. The number of iterations is t = 0. Iteration process: Calculate the inner product of the residual r and each atom in the dictionary D, and select the atom index with the largest absolute value of the inner product: ;in, represents the jth column atom of dictionary D, and K is the number of atoms in the dictionary. Join the active set : ; Use the least squares method to solve the sparse representation coefficient on the active set Λ : ;in, represents the submatrix consisting of dictionary atoms indexed by the active set Λ. Update the residual: ; Iteration number t=t+1. Stop condition: When the L2 norm of the residual r is less than or equal to the reconstruction error threshold ε, the iteration stops. Or, when the iteration number t reaches the preset maximum number of iterations, the iteration stops. Output the reconstructed sparse representation coefficients : The dimension of is the same as the dimension of the initial sparse representation coefficient α. In the example, the position corresponding to the active set Λ is , and the rest of the positions are 0. For each text data block or image data block, output its corresponding reconstructed sparse representation coefficient . The reconstructed sparse representation coefficients It can be used for subsequent data processing tasks, such as data compression, feature extraction, classification, etc.
[0069] Using the reconstructed sparse representation coefficients , reconstruct each text data block and image data block through the preset overcomplete dictionary D, and splice the reconstructed data blocks in the original order. , reconstruct the text data block. Reconstruction formula: ;in, represents the reconstructed text data block, D represents the preset over-complete dictionary, Represents the reconstructed sparse representation coefficient. Reconstruction process: Each atom (column) of the dictionary D is associated with the corresponding sparse representation coefficient Multiply the elements in to get the different components of the reconstructed text data block. Add these components together to get the final reconstructed text data block . Reconstructed text data block Has the same dimensions as the original text data block.
[0070] For each image data block: use the preset over-complete dictionary D and the reconstructed sparse representation coefficients , reconstruct the image data block. Reconstruction formula: ;in, represents the reconstructed image data block, D represents the preset over-complete dictionary, Represents the reconstructed sparse representation coefficient. Reconstruction process: Each atom (column) of the dictionary D is associated with the corresponding sparse representation coefficient Multiply the elements in to get the different components of the reconstructed image data block. Add these components to get the final reconstructed image data block . Reconstructed image data block The same dimension as the original image data block. The linear combination of the overcomplete dictionary D realizes the reconstruction of text data blocks and image data blocks. The non-zero elements in correspond to the important atoms in the dictionary D, which play a key role in data reconstruction. The overcomplete dictionary D provides a rich set of basic representation elements, so that data can be accurately reconstructed by combining a small number of atoms.
[0071] Concatenate the reconstructed data blocks. For text data: concatenate the reconstructed text data blocks Splice in the original order. Splice process: According to the original division order of the text data blocks, the reconstructed text data blocks are Connect them one by one. After splicing, we get the reconstructed complete text data . The complete text data after reconstruction The length and content of the original complete text data are the same. For image data: the reconstructed image data block Splicing is performed in the original order. Splicing process: According to the original division position of the image data block, the reconstructed image data block is Put it back to the corresponding position. After splicing, the reconstructed complete image data is obtained . The reconstructed complete image data It has the same size and visual content as the original complete image data. By splicing the reconstructed data blocks in the original order, the integrity and continuity of the text data and image data are restored. For text data, the spliced complete text data The semantic and contextual information of the original text is preserved.
[0072] For image data, the complete image data after stitching The visual structure and detail information of the original image are retained. Through sparse representation and dictionary reconstruction, data compression and denoising are achieved while retaining the key features of the data. Through the splicing of data blocks, the integrity and continuity of the data are restored, ensuring the semantic and visual consistency of the reconstructed data with the original data.
[0073] The attention mechanism is used to associate and fuse the restored text data and image data to obtain the multi-source information fusion result, and the reconstructed complete text data is obtained. Perform necessary preprocessing, such as tokenization, word embedding, etc. Represent the preprocessed text data as a matrix or tensor, with each row corresponding to a word or word vector. Perform necessary preprocessing, such as scaling and normalization. Represent the preprocessed image data as a matrix or tensor, where each element corresponds to a pixel value or feature value. The attention mechanism calculates the correlation or importance between different data and adaptively assigns different weights to achieve data association fusion.
[0074] The calculation of attention weights takes into account the content and context of the data, which can highlight key information and suppress secondary information. , the attention weights between different words are calculated through the self-attention mechanism. The steps of the self-attention mechanism are: Through linear transformation, we get the query matrix Q, key matrix K and value matrix V. We calculate the similarity between the query matrix Q and the key matrix K to get the attention weight matrix A: , where d is the dimension of the word vector. Multiply the attention weight matrix A with the value matrix V to get the weighted text representation : Through the self-attention mechanism, each word in the text data is associated with other words, and the weights reflect the relevance and importance between words.
[0075] Attention calculation of image data: For image data , extract the local features of the image through the convolutional neural network, and then calculate the attention weights between different regions through the attention mechanism. Steps of the image attention mechanism: The local features are extracted through the convolutional neural network to obtain the feature map F. The feature map F is linearly transformed to obtain the query matrix , key matrix Sum Matrix . Calculate the query matrix With key matrix The similarity of the attention weight matrix is obtained : , where d is the feature dimension. The attention weight matrix With value matrix Multiply them together to get the weighted image representation : Through the image attention mechanism, different regions in the image are associated, and the weights reflect the relevance and importance between regions.
[0076] Cross-attention calculation between text and image: In order to realize the association and fusion of text data and image data, it is necessary to calculate the cross-attention weight between text and image. Steps of the cross-attention mechanism: Represent the text And image representation The query matrix is obtained by linear transformation , key matrix Sum Matrix . Calculate the text query matrix Matrix with image keys Similarity, get the attention weight matrix from text to image . Calculate the image query matrix Matrix with text keys Similarity, get the attention weight matrix from image to text . Multiply the attention weight matrix by the value matrix to get the representation of the text-image interaction and Through the cross-attention mechanism, text data and image data are associated, and the weight reflects the relevance and importance between the two types of data.
[0077] Multi-source information fusion, text representation , Image Representation And the expression after interaction and The fusion method can be splicing, weighted averaging, gating mechanism and other methods to obtain the final multi-source information fusion representation . Multi-source information fusion representation It integrates the information of text data and image data, and makes full use of the association and complementarity between different data. Output multi-source information fusion representation , as the result of the association fusion of the restored text data and image data. Multi-source information fusion results It can be used for subsequent tasks such as classification, retrieval, generation, etc., providing a more comprehensive and accurate information representation.
Claims
1. A multi-source information fusion method based on broadband and narrowband communication, characterized in that: include: Preprocessing the acquired multi-source heterogeneous data; the multi-source heterogeneous data includes text data and image data; Use the pre-trained BERT network to extract text feature vectors from pre-processed text data , use the pre-trained convolutional neural network ResNET to extract the image feature vector of the pre-processed image data ; Use the pre-trained bidirectional gated recurrent network BiGRU to obtain text feature vectors respectively Priority and the image feature vector Priority ; Adaptively set text priority threshold and image priority threshold ; Priority Greater than or Greater than Text data or image data is transmitted through orthogonal frequency division multiple access (OFDMA) over a wideband channel; conversely, it is transmitted through code division multiple access (CDMA) over a narrowband channel. At the receiving end, OFDMA decoding is performed on the text data or image data received through the broadband channel, and CDMA decoding is performed on the text data or image data received through the narrowband channel to obtain decoded text data and image data respectively; The decoded text data and image data are sparsely represented using Bayes estimation, and the data is reconstructed by minimizing the L1 norm to obtain the reconstructed text data and image data; The attention mechanism is used to associate and fuse the restored text data and image data to obtain the multi-source information fusion result.
2. The multi-source information fusion method based on broadband and narrowband communication according to claim 1 is characterized in that: Use the pre-trained BERT network to extract text feature vectors from pre-processed text data , use the pre-trained convolutional neural network ResNET to extract the image feature vector of the pre-processed image data ,include: The preprocessed text data is input into the pre-trained BERT network. Through the self-attention mechanism in the BERT network, the similarity between each word in the text data is calculated to extract the semantic features of the text data. Through the feedforward neural network in the BERT network, the text data is transformed to extract the contextual features of the text data; the contextual features reflect the contextual association relationship of the text data; The gating mechanism in the BERT network is used to fuse the extracted semantic features and context features to obtain the fused text feature vector ; The preprocessed image data is input into the pre-trained convolutional neural network ResNET, and the local features of the image data are extracted through the convolutional layer in ResNET; wherein the local features reflect the texture and edge information of the image data; By using the attention mechanism, the similarity between local features is calculated and weighted fusion is performed on the local features to obtain the global features that reflect the overall semantic information of the image data. Through residual connection, local features and global features are added element by element to obtain the image feature vector .
3. The multi-source information fusion method based on broadband and narrowband communication according to claim 1 is characterized in that: Use the pre-trained bidirectional gated recurrent network BiGRU to obtain text feature vectors respectively Priority and the image feature vector Priority ,include: Input the text feature vector into the pre-trained gated recurrent network GRU, which contains an update gate, a reset gate, and an output gate; The update gate calculates the weighted sum of the hidden state output by the GRU at the previous moment and the text feature vector at the current moment through the sigmoid function to control the degree of transmission of the hidden state at the previous moment to the current moment; The reset gate calculates the weighted sum of the hidden state output by the GRU at the previous moment and the text feature vector at the current moment through the sigmoid function to control the influence of the hidden state at the previous moment on the current moment; GRU updates the hidden state according to the output of the update gate and the reset gate to obtain the hidden state at the current moment; Using the multi-head attention mechanism, the text feature vector is divided into multiple sub-features, the attention weights between the sub-features and the current hidden state are calculated respectively, and the current hidden state is weightedly fused to obtain the multi-head attention vector of the text feature vector; Through the output gate, the multi-head attention vector of the text feature vector is concatenated with the current hidden state, and mapped to a low-dimensional space through a fully connected layer to obtain the priority vector of the text feature vector. ; Input the image feature vector into the pre-trained gated recurrent network GRU and repeat the process to obtain the priority vector of the image feature vector. ; Calculate the priority vectors of the text feature vectors respectively and the priority vector of the image feature vector The L2 norm of and .
4. The multi-source information fusion method based on broadband and narrowband communication according to claim 1 is characterized in that: Adaptively set text priority threshold and image priority threshold ,include: Using the extracted text feature vector and the image feature vector , the information entropy of the current text data and image data is calculated as follows and : , ; in, is the text feature vector, is the image feature vector; for The mean vector of for The mean vector of ; is the weight coefficient of the information entropy gain rate, which is used to balance the influence of information entropy and feature difference on semantic importance, and its value range is [0, 1]; according to and , the semantic importance index is calculated by the following formula and : , ; in, and The calculated and The maximum value of is the smoothing factor of the semantic importance index, which is used to prevent the denominator from being 0 and has a value range of (0, 1); Get the generation timestamp of the current text data and image data and ; Calculate the time difference between the generated timestamp of text data and image data and the current system time T respectively and : , ; in, The timestamp for generating text data. is the generation timestamp of the image data, T is the current system time, and the time unit is seconds; The time difference calculated and , calculate the timeliness index and : when Less than the time difference threshold hour, ,on the contrary, ;in, is the decay rate of text data timeliness; when Less than the time difference threshold hour, ,on the contrary, ;in, is the time-dependent decay rate of image data; According to the semantic importance index and , and timeliness indicators and , the priority threshold is calculated by the following formula and : , ; in, and is the weight coefficient of semantic importance and timeliness; σ is the priority softening factor.
5. The multi-source information fusion method based on broadband and narrowband communication according to claim 4 is characterized in that: The priority softening factor σ is calculated by the following formula: ; in, and are the average values of the semantic importance index and the timeliness index respectively; the value range of σ is (0, 1), which is used to soften the comprehensive priority index to make its distribution more balanced.
6. The multi-source information fusion method based on broadband and narrowband communication according to claim 1 is characterized in that: Priority Greater than or Greater than The text data or image data is transmitted through orthogonal frequency division multiple access (OFDMA) via a broadband channel. On the contrary, narrowband channels are used for transmission through code division multiple access CDMA, including: Determine text feature vector Priority Is it greater than or equal to the text priority threshold? , determine the image feature vector Priority Is it greater than or equal to the image priority threshold? ; if Greater than or equal to or Greater than or equal to , the corresponding text data or image data is divided into the first data, otherwise it is divided into the second data; For the first data, a broadband channel is used to transmit the data through orthogonal frequency division multiple access (OFDMA); The second data is transmitted through code division multiple access (CDMA) using a narrowband channel.
7. The multi-source information fusion method based on broadband and narrowband communication according to claim 6 is characterized in that: For the first data, a broadband channel is used to transmit the data through orthogonal frequency division multiple access (OFDMA), including: Prioritize the first data or As a waterline, the signal-to-noise ratios of the subcarriers in the OFDMA channel are sorted, and the signal-to-noise ratios of the sorted subcarriers are used as the height of the subcarriers; Select unassigned subcarriers in order from high to low according to the watermark until the number of selected subcarriers reaches a preset number; wherein the preset number is related to the priority of the first data or Positive correlation; The first data is modulated onto the allocated subcarrier, and the first data is modulated using quadrature amplitude modulation (QAM). The modulation order and priority or Positive correlation; The modulated first data is transmitted through a broadband channel.
8. The multi-source information fusion method based on broadband and narrowband communication according to claim 7 is characterized in that: The second data is transmitted by using a narrowband channel through code division multiple access CDMA, including: According to the priority corresponding to the second data or , determine the spreading factor SF of the code division multiple access CDMA channel, the spreading factor SF and the priority or Anti-correlation; According to the spreading factor SF, the second data is spread using a pseudo-random code corresponding to a code channel of code division multiple access CDMA; wherein the number of chips of the pseudo-random code is equal to the spreading factor SF; and the second data after the spreading is pseudo-randomly distributed on the code channel; Power allocation is performed on the second data after spectrum spread, and the allocated power P is: or ; in, is the maximum transmission power of the code division multiple access CDMA channel; Mapping the second number after power allocation to a code channel of code division multiple access CDMA; The second data on the code channel is modulated by quadrature phase shift keying (QPSK), and the modulated second data is transmitted through a narrowband channel.
9. The multi-source information fusion method based on broadband and narrowband communication according to claim 1 is characterized by: The reconstructed text data and image data are obtained, including: Dividing the decoded text data into a plurality of text data blocks, and dividing the decoded image data into a plurality of image data blocks; Calculate the optimal sparse representation coefficient of each text data block and image data block in the preset overcomplete dictionary D : ; in, is the optimal sparse representation coefficient, which is a column vector whose number of elements is equal to the number of atoms in the dictionary D; is the sparse representation coefficient, which is a column vector representing the representation coefficient of the original signal y in the dictionary D; y is the text data block or image data block to be sparsely represented, which is a column vector; λ is the regularization parameter of sparse representation; The orthogonal matching pursuit (OMP) algorithm is used to minimize the L1 norm of the sparse representation coefficient α of each data block to obtain the reconstructed sparse representation coefficient. ; Using the reconstructed sparse representation coefficients , each text data block and image data block is reconstructed through the preset over-complete dictionary D to obtain the reconstructed text data block and image data blocks : , ; The reconstructed text data block and image data blocks Splice in the original order to get the reconstructed complete text data and image data .
10. The multi-source information fusion method based on broadband and narrowband communication according to claim 9 is characterized in that: Reconstructed sparse representation coefficients , solved by the following formula: ; Constraints: ; Among them, β is the optimization variable in the optimization problem, which represents the sparse representation coefficient that minimizes the objective function under the constraints; express The L1 norm distance between and β; represents the L0 norm of β; is the sparsity weight coefficient; is the reconstruction error threshold; Represents the reconstructed data The L2 norm error between the original data y is used to measure the reconstruction quality.
Citation Information
Patent Citations
Text sentiment analysis method based on BERT model and double-channel attention
CN110717334A
Aspect-level sentiment analysis system and method based on multi-channel attention fusion
CN116205222A