A construction method of multi-modal trusted semantic communication for audio-visual event positioning

By employing a multimodal trusted semantic communication method, the data transmission problem of traditional communication systems in complex environments is solved, achieving efficient, reliable, and secure information transmission in audio-visual event localization tasks, thereby improving the overall performance and information fidelity of the system.

CN119011087BActive Publication Date: 2025-11-07JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411225534.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2025-11-07
Estimated Expiration
2044-09-03

AI Technical Summary

Technical Problem

Traditional communication systems struggle to handle data transmission in complex environments during multimodal tasks, leading to information semantic loss or errors, especially in audio-visual event localization tasks, where maintaining efficient and stable communication is particularly difficult under complex channel conditions.

Method used

A multimodal trusted semantic communication method is adopted, which acquires audiovisual perception data through audiovisual sensors to realize a cross-modal audio-guided visual attention mechanism. Combined with trusted channel coding, real-to-complex conversion, channel estimation and signal recovery, a two-layer coding system is designed using public-key cryptography and symmetric encryption technology to ensure the reliability and security of information transmission.

Benefits of technology

It achieves efficient, reliable and secure multimodal information transmission under complex channel conditions, improves the accuracy of audio-visual event localization and system performance, and ensures the absolute fidelity and security of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119011087B_ABST
    Figure CN119011087B_ABST
Patent Text Reader

Abstract

The application relates to a construction method of a multi-modal trusted semantic communication for audio-visual event positioning, and aims to improve the security and reliability of an audio-visual event (AVE) positioning task. The application effectively protects the integrity and privacy of data in the transmission process through advanced semantic coding and channel coding technology. Through simulating real-world channel conditions and adopting Reed-Solomon coding and AES encryption technology, the accuracy and reliability of data transmission are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, and particularly relates to a construction method of multi-modal trusted semantic communication for audio-visual event positioning. BACKGROUND

[0002] With the rapid development of technology, mobile devices and intelligent applications have become increasingly popular, greatly enriching people's daily life and work efficiency. The widespread use of these devices and applications has led to an explosive growth in wireless data traffic. According to data from the International Telecommunication Union (ITU), global mobile data traffic has grown exponentially in the past few years, and this trend is expected to continue in the coming years. This rapid growth in data traffic poses unprecedented challenges to modern communication systems.

[0003] Traditional communication systems mainly rely on bit-level data transmission, that is, converting data into a bit stream at the transmitting end, and then accurately recovering these bits at the receiving end to reconstruct the original data. This process requires the communication system to have high-quality channel conditions and a high signal-to-noise ratio (SNR) to ensure the accuracy and reliability of data transmission. However, as the amount of data increases, these requirements of traditional communication systems become increasingly difficult to meet. In addition, traditional communication systems often struggle to maintain efficiency and stability when dealing with data transmission in complex environments.

[0004] The limitations of traditional communication systems are more apparent in multi-modal tasks, especially in audio-visual event (AVE) positioning tasks. Multi-modal tasks usually involve processing data from different sources (such as audio and video) and integrating these data to obtain richer semantic information. For example, in security monitoring, environmental perception of autonomous vehicles, and intelligent medical diagnosis applications, accurately identifying and positioning audio and visual events is crucial. However, due to the presence of physical noise and channel interference, multi-modal information is easily lost and disturbed during transmission, leading to loss or error of semantic information, which in turn affects the performance of the task and the system.

[0005] To solve these problems, a new communication paradigm is needed, that is, to change from traditional bit accuracy to semantic fidelity. This paradigm emphasizes the direct transmission and retrieval of the semantics of the content, rather than just the bit stream. Semantic communication can better handle data transmission under complex channel conditions, as it focuses on the semantics of the information rather than the specific bit pattern. This shift requires communication systems not only to be able to handle accurate transmission of data, but also to understand and utilize the semantic content of the data, thereby achieving more efficient and reliable communication in various environments.

[0006] Furthermore, the cross-modal information complementarity and semantic richness in multi-modal tasks are crucial for improving the overall performance of the system. For example, in audio-visual event localization, audio and video data can complement each other, providing a more comprehensive understanding of the scene. Audio data can provide temporal information of event occurrence, while video data can provide spatial information. By integrating the information of these two modalities, the system can more accurately locate the position and time of event occurrence. However, to achieve such integration, a communication framework that can effectively process and utilize multi-modal data is needed. SUMMARY

[0007] Modern communication systems face the challenge of rapid growth of wireless data traffic, and new technologies and methods need to be developed to improve the security and reliability of multi-modal semantic communication, especially in key tasks such as audio-visual event localization. This not only requires improvements in channel coding and decoding techniques, but also the development of new semantic coding and decoding methods, as well as signal processing techniques that can adapt to complex communication environments.

[0008] In order to solve the problems existing in the prior art, the present application proposes a construction method of multi-modal trusted semantic communication for audio-visual event localization, comprising the following steps:

[0009] S1: Obtain audio-visual perception data in the same time interval through audio-visual sensors, one user holds video data and one user holds audio data;

[0010] S2: Divide the video and audio time series S into T non-overlapping but continuous segments; respectively represented as video segment sequence and audio segment sequence The duration of each segment is one second, where t is the time segment, ranging from 1 to T, T is a constant representation, V t represents the video data in time segment t, a t represents the audio data in time segment t;

[0011] S3: Implement cross-modal audio-guided visual attention mechanism for semantic communication between two users, realize visual attention under audio guidance;

[0012] S4: Perform trusted channel coding on the video segments in the continuous and synchronous audio-visual segments held by the two users respectively and the corresponding time interval audio segments;

[0013] S5: Real-to-complex conversion module, convert the feature tensor to complex representation to cope with the complexity of data transmission in analog channels;

[0014] S6: Solve the semantic information distortion problem in the communication ecosystem, including additive white Gaussian noise, Rayleigh fading and Rice fading, by using chan_layer;

[0015] S7: Calculate the channel matrix by means of complex domain transformation and channel estimation method, and realize signal recovery;

[0016] S8: The receiver model decodes and decrypts the signal propagated through the channel, and prepares for the final classification or prediction task;

[0017] S9: Calculate the cross-modal similarity and predict the event category, and locate.

[0018] Further, the visual attention mechanism in S3 above includes:

[0019] S3.1 V t and a t Use VGG-19 network to extract initial visual feature V and initial audio feature A as initial vector, and feature representation on the whole time sequence is audio feature and visual feature

[0020] S3.2 Transmit the audio from the audio holding user to the video holding user through the physical channel, and encode it by using the audio semantic encoder:

[0021]

[0022] Where A single is the output result, that is, the single audio segment obtained after enhancement processing; is the audio feature, is a function of processing audio feature, which enhances the signal, and Θ a,1 represents trainable parameters, which consists of a one-dimensional convolution with C channels and 7 kernel sizes, and B convolution blocks;

[0023]

[0024] Each block includes a residual unit and a down-sampling layer with step convolution, where the kernel size K s is twice the step size S, and the residual unit has two convolutions with kernel size 3 and a jump connection;

[0025] S3.3 Integrate all the audio segments obtained after processing to obtain The integration process is as follows.

[0026] Block i = ResUnit°DownSample(Ks

[0027] where DownSample(K s ,S) is a function, representing down-sampling, K s is kernel size, indicating the size of the convolution kernel used in the down-sampling process; S is stride, indicating the step size of moving the convolution kernel in the down-sampling process, Block i is the i-th block after a series of operations, and ResUnit is a residual unit containing several convolution layers and has a skip connection to preserve the information of the original input;

[0028] S3.4 the user holding the video, the received audio feature is accurately recovered in the single-modal task-oriented decoding stage; this recovered feature is used in the cross-modal semantic encoder to enhance the visual feature, and the cross-modal visual semantic encoder is defined as:

[0029]

[0030] where Θ v is the trainable parameter of the encoder, is the visual feature, is a function for processing the visual feature, which enhances the signal;

[0031] S3.5 calculate the attention weight α t :

[0032]

[0033] where σ is the sigmoid function, M v , M a is the projection feature obtained by projecting the visual and audio features to a shared space, W f is a trainable parameter, is the audio feature segment of the audio feature at time t , v t is the video feature segment of the visual feature V at time t

[0034] S3.6 use the attention weight α t to weight the received audio feature and visual feature to obtain the weighted visual feature

[0035] Further, the trusted channel encoding in the above S4 includes:

[0036] S4.1 use public key cryptography to establish a shared session key between the sender and the receiver; ​

[0037] S4.2 symmetric encryption of the transmitted information using the shared session key;

[0038] S4.3 The semantic encoder holds a key pair (pk, sk) of an elliptic curve cryptography algorithm, and the formula is as follows:

[0039] X a = CE a (A1; Φ a ) A1 represents the input data as the weighted visual feature V at obtained by weighting in S3 and the received audio feature

[0040] CE a is used to extract attention-related features, and Φ a is a vector representing the learnable parameters in CE a ;

[0041] S4.4 Set the AES encryption key K from usekey, encrypt the key K using the ECC public key, obtain the ciphertext c1, and process the audio message and the video message according to the data block model respectively;

[0042] S4.5 AES encrypt the processed message m using the key K and the CBC mode to obtain the ciphertext c2, calculate the hash value h of the message m according to the formula h = SHA-3 (m), and obtain the final ciphertext c = (c1, c2, h);

[0043] S4.6 Encode the obtained ciphertext c using Reed-Solomon encoding.

[0044] Further, the real-to-complex conversion module in S5 above includes:

[0045] S5.1 The convolutional neural network layer learns diverse local features from multi-modal semantics, where

[0046] M represents a multi-modal set, including the weighted visual feature V at and the initial audio feature A, and are the time encoding of the enhanced visual and audio features using LSTM at time t, respectively, are the hidden state and cell state of the visual LSTM at the previous time t-1, respectively, and V t is the video feature segment at time t; similarly, the hearing is also enhanced by are the hidden state and cell state of the hearing LSTM at the previous time t-1, respectively, and A tis the audio feature segment at time t;

[0047] S5.2 Reshape the encoded and learned data into a three-dimensional tensor and convert it to a complex form, i.e., f complex = RealToComplex(f cnn ).

[0048] Further, the above S6 solves the semantic information distortion problem in the communication ecosystem, including:

[0049] S6.1 Channel simulation, AWGN channel: generate Gaussian noise for the real and imaginary parts of the input signal respectively, and add the Gaussian noise to the input signal; set the channel matrix I as the unit matrix, representing the undistorted channel; Rayleigh fading channel: generate a complex channel matrix H and apply it to the input signal to form the output signal Y; calculate the conjugate transpose H H and outer product H H H of the channel matrix H;

[0050] S6.2 Rayleigh fading channel simulation: generate a complex channel matrix H and apply it to the input signal to form the output signal Y; calculate the conjugate transpose H H and outer product H H H of the channel matrix H;

[0051] S6.3 Rician fading channel simulation: construct a channel matrix H based on the line-of-sight and non-line-of-sight components and apply it to the signal to produce the output signal Y.

[0052] Further, the above S7 includes:

[0053] S7.1 The received output signal Y, the transmitted signal vector is X, the channel matrix is defined as H, and the additive white Gaussian noise vector is defined as N; the received signal can be represented as: Y = HX + N;

[0054] S7.2 Signal detection and signal recovery: use a zero-forcing minimum mean square error detector to eliminate multipath interference and reduce noise amplification; the purpose of the ZF-LMMSE detector is to find a detection matrix W, so that the output is close to the real transmitted signal The detection of the ZF-LMMSE detector can be shown as:

[0055]

[0056] where H H is the conjugate transpose of the channel matrix H, is the variance of the noise, is the variance of the transmitted signal; I is the identity matrix.

[0057] Further, the above S8 includes the following:

[0058] S8.1 After signal recovery and signal detection, the estimated complex signal is RS-decoded by a trusted decoder, and the data is converted into audio and visual feature information;

[0059] S8.2 The semantic decoder recovers the AES encrypted message c,

[0060] c1 = the first 160 bits of c

[0061] The first 160 bits are taken as c1, the last 128 bits of the ciphertext c are extracted as h, and the remaining bits in the middle of the ciphertext c are extracted as c2; using the elliptic curve cipher (ECC) private key, the decrypted message m is obtained after decryption;

[0062] S8.3 Using the combined audio, visual semantic information, using positive sample propagation network wherein is the audio feature in the decrypted message m, wherein is the video feature in the decrypted message m, φ represents the parameter set of the entire network, and SD represents decoding of the audio and visual features at the same time; by screening samples, negative samples are removed, and the decoded audio and visual features A r and V r are obtained by completing semantic decoding.

[0063] Further, the above S9 includes:

[0064] S9.1 The decoded audio feature A r is processed through two fully connected layers to obtain A e ; the decoded visual feature V r is processed through two Conv2D blocks to obtain V e ; and the cosine similarity is calculated

[0065] S9.2 Feature fusion is performed in a shared fully connected layer manner to complete event prediction

[0066] F = ReLU (W e · [A e ; V e ] + b)

[0067] wherein W e is a weight matrix, b is a bias vector, and F is a prediction result;

[0068] S9.3 locates the position where the event occurs, and determines the position of the audio event in the video frame The positioning formula is:

[0069]

[0070] Wherein L represents the positioning result of the event, P is a set of all possible positions, and the position with the highest cosine similarity score is taken as the occurrence position of the event.

[0071] The beneficial technical effects of the present application are as follows:

[0072] 1. A multi-modal trusted semantic communication construction method for audio-visual event positioning is designed, and reliable information transmission is realized through a reliable channel encoder and decoder.

[0073] 2. A double-layer coding system is designed, and a layer of error correction code is added outside the traditional channel encoder to correct the small weight error that may be introduced by the traditional encoding and decoding, thereby ensuring the absolute fidelity of the output information. With the help of reliable channel encoders and decoders, the information input by the encoder and the information output by the decoder can be ensured to be completely consistent, thereby laying a foundation for the transmission of ciphertext by the sender and subsequent decryption by the receiver.

[0074] 3. The concept of hybrid encryption is adopted to realize the secure transmission of semantic information. First, a shared secure session key is established between the sender and the receiver by using the principle of public key cryptography. Then, the session key is used for symmetric encryption of the information to be transmitted. This method combines the advantages of public key and symmetric key cryptography, the former facilitates key distribution, and the latter provides data encryption efficiency and speed. Finally, a powerful secure communication framework is obtained, which can maintain the confidentiality and integrity of communication in a potentially hostile network. BRIEF DESCRIPTION OF DRAWINGS

[0075] Figure 1 A multi-modal trusted semantic communication construction method for audio-visual event positioning. DETAILED DESCRIPTION

[0076] The present application will be described in detail below in conjunction with the drawings and specific embodiments. As shown in the drawings, the present application is a multi-modal trusted semantic communication construction method for audio-visual event positioning, which comprises the following steps: Figure 1

[0077] S1: Acquire audio-visual perception data in the same time interval through audio-visual sensors, one user holds video data and one user holds audio data;

[0078] ​S2: Segmenting the video and audio time series S (corresponding to the same application scenario but sent to the remote receiver through different sensors and transmission ends) into T non-overlapping but continuous segments; represented as video segment sequence and audio segment sequence The duration of each segment is one second, where t is the time segment, ranging from 1 to T, T is a constant representation, V t represents the video data in time segment t, a t represents the audio data in time segment t;

[0079] S3: Implementing a cross-modal audio-guided visual attention mechanism for semantic communication between two users, realizing visual attention under audio guidance;

[0080] S4: Respectively encoding the video segments in the continuous and synchronous audio-visual segments held by the two users and the corresponding time interval audio segments in the trusted channel;

[0081] S5: Real-to-complex conversion module, converting the feature tensor into a complex representation to cope with the complexity of data transmission in the analog channel;

[0082] S6: Using chan_layer to solve the semantic information distortion problem in the communication ecosystem, which includes additive white Gaussian noise (AWGN), Rayleigh fading and Rice fading;

[0083] S7: Calculating the channel matrix with the help of complex domain transformation and channel estimation method to realize signal recovery;

[0084] S8: The receiver model decodes and decrypts the signals propagated through the channel, preparing for the final classification or prediction task;

[0085] S9: Calculate the cross-modal similarity and predict the event category, and locate.

[0086] In order to realize the cross-modal cooperation, we design an attention mechanism that allows audio features to guide the model's focus on video content. This includes mapping audio features onto the time axis of the video sequence and calculating the correlation or similarity between audio and video features. Based on these calculated similarities, the system dynamically assigns attention weights to video frames, emphasizing those visual parts that are synchronized with audio and have a large amount of information. In this way, the model can more effectively focus on important visual content, improving overall understanding ability.

[0087] As a preferred embodiment of the present application, the visual attention mechanism in S3 includes:

[0088] S3.1 mapping V t and a tThe initial visual feature V and the initial audio feature A are extracted as initial vectors using the VGG-19 network, and the feature representation on the entire time sequence is denoted as audio feature and visual feature

[0089] S3.2 Transmit the audio from the audio-holding user to the video-holding user through the physical channel, and encode it using the audio semantic encoder:

[0090]

[0091] where A single is the output result, i.e., the single audio segment obtained after enhancement processing; is the audio feature, is a function of processing the audio feature, and Θ a,1 represents trainable parameters, which are composed of a one-dimensional convolution with C channels and a kernel size of 7, and B convolution blocks;

[0092]

[0093] Each block includes a residual unit and a down-sampling layer with a stride convolution, where the kernel size K s is twice the stride S, and the residual unit has two convolutions with a kernel size of 3 and a skip connection;

[0094] S3.3 Integrate all the processed audio segments to obtain The integration process is as follows.

[0095]

[0096] where DownSample(K s ,S) is a function representing down-sampling, K s is the kernel size, indicating the size of the convolution kernel used in the down-sampling process; S is the stride, indicating the step size of moving the convolution kernel in the down-sampling process; Block i is the i-th block after a series of operations, and ResUnit is a residual unit containing several convolution layers and having a skip connection to preserve the information of the original input.

[0097] S3.4 The video-holding user receives the audio feature which is accurately recovered in the single-modal task-oriented decoding stage; this recovered feature is used in the cross-modal semantic encoder to enhance the visual feature, and the cross-modal visual semantic encoder is defined as:

[0098]

[0099] where Θ v trainable parameters of the encoder, is the input visual feature, is a function that processes the visual feature, enhancing the signal;

[0100] S3.5 Compute attention weight α t :

[0101]

[0102] where σ is the sigmoid function, M v , M a is the projected feature from the visual and audio features projected into the shared space, W f is a trainable parameter;

[0103] S3.6 Use attention weight α t to weight the audio feature and the visual feature to get the weighted visual feature

[0104] As a preferred embodiment of the present application, the trusted channel encoding in S4 includes:

[0105] S4.1 Use public key cryptography to establish a shared session key between the sender and the receiver;

[0106] S4.2 Use the shared session key to symmetrically encrypt the transmitted information;

[0107] S4.3 The semantic encoder holds the key pair (pk, sk) of the elliptic curve cryptography algorithm, and the calculation formula is as follows:

[0108] X a = CE a (A1; Φ a ) A1 represents the input data, which is the weighted visual feature V at obtained by weighting in S3 and the received audio feature

[0109] CE a is used to extract attention-related features, Φ a is a vector, representing the learnable parameters in CE a ;

[0110] S4.4 Set the AES encryption key K from usekey, encrypt the key K using the ECC public key to get the ciphertext c1, process the audio message and the video message according to the data block model respectively;

[0111] S4.5 uses key K and CBC mode to encrypt the processed message m with AES to obtain ciphertext c2. The hash value h of message m is calculated according to the formula h = SHA-3(m) to obtain the final ciphertext c = (c1, c2, h).

[0112] S4.6 Encodes the obtained ciphertext c using Reed-Solomon encoding, specifying the number of roots (nroots) of the generator polynomial as 1280, and using the extended field GF (2^18).

[0113] To fully utilize the rich information in video data, we innovatively reshape and extract video features through 2D convolutional layers, treating each time point in the sequence as a "height" in image processing, thus enhancing the ability to capture spatiotemporal features. Subsequently, the video features are converted into complex form, leveraging the properties of the complex domain to enhance the model's expressive power and noise resistance. Complex L2 norm normalization ensures the normalization of all feature vectors, maintaining the consistency and efficiency of computation.

[0114] In a preferred embodiment of the present invention, the real-to-complex conversion module in S5 includes:

[0115] The S5.1 convolutional neural network layer learns diverse local features from multimodal semantics, among which... M represents a multimodal set, including weighted visual features V. at And initial audio features A, and These are the temporal encodings of visual and audio features enhanced using LSTM at time t. These are the hidden state and cell state of the visual LSTM at the previous time step t-1, respectively, V t It is a video feature segment at time t; similarly, hearing is also through... These are the hidden state and unit state of the auditory LSTM at the previous time step t-1, A t It is the audio feature segment at time t;

[0116] S5.2 reshapes the encoded and learned data into a three-dimensional tensor and converts it into complex form, i.e., f. complex =RealToComplex(f cnn ).

[0117] Considering the challenges in the actual communication environment, we integrate a set of channel simulation and signal recovery mechanisms, namely chan_layer. It can not only simulate additive white Gaussian noise (AWGN), Rayleigh fading, Rician fading and other channel environments, but also recover the damaged signal under poor channel conditions based on ideal channel estimation theory, such as matrix solving algorithm. By controlling the signal-to-noise ratio (SNR) and channel type, the model can be trained in a simulated complex channel to learn adaptive strategies for different conditions, significantly improving the generalization performance of the model in unknown environments.

[0118] As a preferred embodiment of the present application, S6 solves the problem of semantic information distortion in the communication ecosystem, which includes:

[0119] S6.1 Channel simulation, AWGN (Additive White Gaussian Noise) channel: generate Gaussian noise for the real and imaginary parts of the input signal, and add the Gaussian noise to the input signal; set the channel matrix I to the identity matrix, representing a distortionless channel; Rayleigh fading channel: generate a complex channel matrix H (representing the fading coefficient) and apply it to the input signal to form the output signal Y; calculate the conjugate transpose H H and outer product H H H, Rician fading channel: construct a channel matrix H based on line of sight (LOS) and non-line of sight (NLOS) components and apply it to the signal to produce output Y;

[0120] S6.2 Rayleigh fading channel simulation: generate a complex channel matrix H (representing the fading coefficient) and apply it to the input signal to form the output signal Y; calculate the conjugate transpose H H and outer product H H H;

[0121] S6.3 Rician fading channel simulation: construct a channel matrix H based on line of sight (LOS) and non-line of sight (NLOS) components and apply it to the signal to produce output Y.

[0122] As a preferred embodiment of the present application, S7 includes:

[0123] S7.1 Receive the output signal Y, the transmitted signal vector is defined as X, the channel matrix is defined as H, and the additive white Gaussian noise vector is defined as N; the received signal can be represented as: Y = HX + N;

[0124] S7.2 Signal detection and signal recovery: use a zero-forcing least mean square error (ZF-LMMSE) detector to eliminate multipath interference and reduce noise amplification; the purpose of the ZF-LMMSE detector is to find a detection matrix W, so that the output is close to the true transmitted signal The detection of the ZF-LMMSE detector can be shown as:

[0125]

[0126] where H H is the conjugate transpose of the channel matrix H, is the variance of the noise, is the variance of the transmitted signal; I is the identity matrix.

[0127] As a preferred embodiment of the present application, S8 comprises the following:

[0128] S8.1 After signal recovery and signal detection, the estimated complex signal is subjected to RS decoding by a reliable decoder, and the data is converted into audio and visual feature information;

[0129] S8.2 The semantic decoder recovers the AES-encrypted message c,

[0130] c1 = the first 160 bits of c

[0131] The first 160 bits are taken as c1, the last 128 bits of the ciphertext c are extracted as h, and the remaining bits in the middle of the ciphertext c are extracted as c2; and the decrypted message m is obtained by decrypting the same using an elliptic curve cipher (ECC) private key;

[0132] S8.3 Using the combined audio and visual semantic information, a positive sample propagation network is used wherein is the audio feature in the decrypted message m, wherein is the video feature in the decrypted message m, and φ represents the parameter set of the entire network, and SD indicates that the audio and visual features are decoded simultaneously; by screening the samples, the negative samples are removed, and the decoded audio and visual features A r and V r are obtained.

[0133] As a preferred embodiment of the present application, S9 comprises:

[0134] S9.1 The decoded audio feature A r is processed through two fully connected layers (number of units = 256, activation function = 'ReLU') to obtain A e ; the decoded visual feature V r is processed through two Conv2D blocks (number of units = 512, activation function = 'ReLU') to obtain V e ; and the cosine similarity is calculated

[0135] S9.2 Feature fusion is completed by using a shared fully connected layer to predict the event

[0136] F = ReLU(W e ; V e + b) e

[0137] where W e is a weight matrix, b is a bias vector, and F is the predicted result.

[0138] S9.3 Positioning the location of the event, determining the location of the audio event in the video frame, the positioning formula is:

[0139]

[0140] where L represents the positioning result of the event, P is a set of all possible positions, and the position with the highest cosine similarity score is taken as the occurrence position of the event.​

Claims

1. A method for constructing multi-modal trustworthy semantic communication oriented to audio-visual event localization, characterized in that, The method comprises the following steps: S1: acquiring audio-visual perception data in the same time interval through audio-visual sensors, one user holding video data and one user holding audio data; S2: segmenting the video and audio time series S into T non-overlapping but consecutive segments; denoted as a video segment sequence and an audio segment sequence each segment has a duration of one second, where t is a time segment, ranging from 1 to T, T is a constant denoting the number of segments, V t represents the video data in time segment t, a t represents the audio data in time segment t; S3: realizing an audio-guided visual attention mechanism across modalities for semantic communication between two users, realizing visual attention under audio guidance; S4: respectively encoding video segments in continuous and synchronous audio-visual segments held by two users and audio segments in corresponding time intervals through a trusted channel; S5: a real-to-complex conversion module, converting a feature tensor into a complex representation to cope with the complexity of data transmission in an analog channel; S6: using chan_layer to solve the problem of semantic information distortion in the communication ecosystem, which includes additive white Gaussian noise, Rayleigh fading and Rician fading; S7: calculating a channel matrix by means of complex domain transformation and channel estimation method to realize signal recovery; S8: a receiver model decodes and decrypts signals propagated through a channel, preparing for a final classification or prediction task; S9: calculating cross-modal similarity and predicting event categories for positioning; The visual attention mechanism in S3 comprises: S3.1 extract V from step S2 t and a t extract initial visual features V and initial audio features A using VGG-19 network as initial vectors, the feature representation on the whole time series is audio features S a and visual features S3.2: transmitting audio from the audio-holding user to the video-holding user through a physical channel and encoding it using an audio semantic encoder; where A single is the output result, i.e., the single audio segment after enhancement processing; is the audio feature, is a function of processing the audio feature, Θ a,1 represents trainable parameters, which consists of a one-dimensional convolution with C channels and 7 kernel sizes, with B convolution blocks; Each block includes a residual unit and a down-sampling layer with stride convolution, where the kernel size K s is twice the stride S, the residual unit has two convolutions with kernel size 3 and a skip connection; S3.3 integrate all the processed audio segments to obtain The integration process is as follows; where DownSample(K s ,S) is a function representing down-sampling, K s is the kernel size, which refers to the size of the convolution kernel used in the down-sampling process; S is the stride, which refers to the step size of moving the convolution kernel in the down-sampling process; Block i is the i-th block after a series of operations, and ResUnit is a residual unit containing several convolution layers and having a skip connection to preserve the information of the original input. S3.4 User holding the video, received audio features The single modality task-oriented decoding stage is accurately recovered; this recovered feature is used in the cross-modal semantic encoder to enhance the visual feature, and the cross-modal visual semantic encoder is defined as: wherein Θ v trainable parameters of the encoder, is a visual feature, is a function that processes the visual feature, enhancing the signal; S3.5 Compute attention weight a t : where σ is a sigmoid function, M v , a is a projected feature resulting from projecting the visual and audio features into a shared space, W f is a trainable parameter, is an audio feature segment of the audio feature V at time t, v t is a video feature segment of the visual feature V at time t; S3.6 Using attention weights a t On the received audio features And visual features Weighted to obtain weighted visual features 2. A method of constructing a multi-modal trustworthy semantic communication oriented to audiovisual event localization as claimed in claim 1, characterized in that, The trusted channel encoding in S4 comprises: S4.1: using public key cryptography to establish a shared session key between the sender and the receiver; S4.2: using the shared session key to symmetrically encrypt the transmitted information; S4.3: the semantic encoder holds an elliptic curve cryptography key pair (pk, sk), and the calculation formula is as follows: X a = CE a (A1; Φ a ) A1 represents the weighted visual features V resulting from weighting the input data in S3 at and the received audio features CE a used to extract attention-related features, Φ a is a vector representing the CE a learnable parameters; S4.4: setting an advanced encryption standard (AES) encryption key K from the user key, encrypting the key K using an elliptic curve cryptography (ECC) public key, obtaining ciphertext c1, and processing the audio message and the video message according to a data block model; S4.5: using the key K and the CBC mode to perform AES encryption on the processed message m, obtaining ciphertext c2, calculating the hash value h of the message m according to the formula h = SHA-3(m), and obtaining the final ciphertext c = (c1, c2, h); S4.6: encoding the obtained final ciphertext c using Reed-Solomon encoding.

3. A method of constructing a multi-modal trustworthy semantic communication oriented to audiovisual event localization as claimed in claim 1, characterized in that, The real-to-complex conversion module in S5 comprises: S5.1 A convolutional neural network layer learns diverse local features from the multi-modal semantics, where M represents a multi-modal set, including weighted visual features V at and initial audio features A, and are the temporal encodings of the enhanced visual and audio features using LSTM at time t, respectively, are the hidden and cell states of the visual LSTM at the previous time t-1, respectively, V t is a video feature snippet at time t; similarly, the hearing is also through are the hidden and cell states of the hearing LSTM at the previous time t-1, respectively, A t is an audio feature snippet at time t; S5.2 Reshape the encoded and learned data into a three-dimensional tensor and convert it to complex form, i.e., f complex = RealToComplex(f cnn ).

4. A method of constructing a multi-modal trustworthy semantic communication oriented to audiovisual event localization as claimed in claim 1, characterized in that, Solving the problem of semantic information distortion in the communication ecosystem in S6 comprises: S6.1 Perform channel simulation, AWGN channel: Generate Gaussian noise for the real and imaginary parts of the input signal, add the Gaussian noise to the input signal; Set the channel matrix I to the identity matrix, representing a distortionless channel; Rayleigh fading channel: Generate a complex channel matrix H and apply it to the input signal to form the output signal Y; Calculate the conjugate transpose H H and the outer product H H H, Rician fading channel: Construct a channel matrix H based on the line-of-sight and non-line-of-sight components and apply it to the signal to produce the output Y; S6.2 Rayleigh fading channel simulation: generate a complex channel matrix H and apply it to the input signal to form the output signal Y; compute the conjugate transpose H H and the outer product H H H; S6.3: Rician fading channel simulation: constructing a channel matrix H based on line-of-sight and non-line-of-sight components and applying it to the signal to produce an output signal Y.

5. A method of constructing a multi-modal trustworthy semantic communication oriented to audiovisual event localization as claimed in claim 1, characterized in that, S7 comprises: S7.1: the received output signal Y, the transmitted signal vector X, the channel matrix defined as H, and the additive white Gaussian noise vector defined as N; the received signal can be represented as: Y = HX + N; S7.2 Signal detection and signal recovery: using zero-forcing minimum mean square error detector to eliminate multipath interference and reduce noise amplification; the purpose of ZF-LMMSE detector is to find a detection matrix W, so that the output is close to the real transmitted signal The detection of ZF-LMMSE detector can be shown as: where H H is the conjugate transpose of the channel matrix H, is the variance of the noise, is the variance of the transmitted signal; I is the identity matrix.

6. A method of constructing a multi-modal trustworthy semantic communication oriented to audiovisual event localization as claimed in claim 1, wherein, S8 comprises the following: S8.1: after signal recovery and signal detection, the estimated complex signal is RS-decoded by a trusted decoder, and the data is converted into audio and visual feature information; S8.2: the semantic decoder recovers the AES-encrypted message c, c1 = the first 160 bits of c The first 160 bits are taken as c1, the last 128 bits of the ciphertext c are taken as h, and the remaining bits in the middle of the ciphertext c are taken as c2; and the decrypted message m is obtained by using an elliptic curve cryptography (ECC) private key to decrypt c1, h and c2; S8.3 Utilize the combined audio, visual semantic information, adopt positive sample propagation network wherein is the audio feature in the decrypted message m, wherein is the video feature in the decrypted message m, and φ represents the parameter set of the entire network, and SD indicates decoding of the audio and visual features simultaneously; by screening the samples, the negative samples are removed, and the decoded audio and visual features are obtained respectively as A r and V r .

7. A method of constructing a multi-modal trustworthy semantic communication oriented to audiovisual event localization as claimed in claim 1, wherein, The S9 comprises: S9.1 decoded audio feature A r processed by two fully connected layers to get A e ; decoded visual feature V r processed by two Conv2D blocks to get V e ; cosine similarity is calculated simultaneously S9.2, feature fusion is performed in a shared full connection layer mode to complete event prediction F = ReLU(W e · [A e ; V e ] + b) where W e is a weight matrix, b is a bias vector, and F is the predicted result. S9.3, a position where the event occurs is located, and a position locating formula of the audio event in the video frame is as follows: Wherein, L represents a locating result of the event, P is a set of all possible positions, and a position with a highest cosine similarity score is taken as a position where the event occurs.