A multimodal face semantic communication method, device and medium

Through the combination of a generative semantic distillation encoder and a channel state predictor, efficient semantic reshaping and channel adaptation of face images are achieved, and the problems of insufficient flexibility, coordination and channel adaptability of face semantic communication solutions in the prior art are solved, thereby improving communication efficiency and data transmission quality.

CN120318893BActive Publication Date: 2025-09-02BEIJING MIANBI INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510789542.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-02
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

The existing face semantic communication solutions have shortcomings in flexibility, coordination, communication efficiency and channel adaptability, and cannot adapt to the current users' diverse needs for face transmission services.

Method used

Generative semantic distillation encoder and double regularization constraints are used to convert high-dimensional face images into low-dimensional latent spatial characterization, combining lightweight channel state predictors and decoder parameters fine-tuning to realize end-to-end optimization of semantic reshaping and channel transmission. Through generative adversarial network reverse mapping technology and multimodal attention mechanism, a fine-grained semantic association channel is established to adapt to different channel environments.

Benefits of technology

It significantly improves the flexibility of semantic remodeling, bandwidth utilization efficiency and channel adaptability, ensures stable and efficient data transmission in complex wireless environments, and solves the semantic association and channel adaptability problems of traditional methods in multimodal data communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318893B_ABST
    Figure CN120318893B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal face semantic communication method, device and medium, which belongs to the field of semantic communication technology and is used to solve the technical problem that the current face semantic communication solution has deficiencies in flexibility, coordination, communication efficiency and channel adaptability, and cannot adapt to the current users' diverse needs for face transmission services. The method comprises: semantically transforming the face image in the face video stream to obtain corresponding image semantic representation information; determining the semantic offset based on the semantic features in the language reshaping instruction and the face features in the face image; optimizing the image semantic representation information according to the semantic offset; optimizing the encoding strategy of the transmission channel between the transmitter and the receiver, and fine-tuning the parameters of the decoder at the receiver based on the channel noise; sending the image semantic representation optimization information to the receiver through the optimized transmission channel, and decoding the semantic representation optimization information and reconstructing the image through the decoder after fine-tuning the parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semantic communication technology, and in particular to a multimodal face semantic communication method, device and medium. Background Art

[0002] With the rapid development of mobile internet technology, demand for facial recognition technology in real-time scenarios such as social media and video conferencing is surging. As an emerging communication paradigm, semantic communication significantly reduces bandwidth requirements and improves transmission efficiency by extracting and understanding the deeper meaning of information rather than simply transmitting raw data, providing a more intelligent facial recognition solution. Further combining edge computing with the coordinated optimization of intelligent chips on terminals can achieve precise facial semantic parsing and reconstruction on mobile devices, providing high-fidelity, low-power, real-time facial recognition services for scenarios such as immersive remote interactions and virtual live broadcasts.

[0003] However, the application of existing semantic communication solutions in face transmission tasks still faces the following bottlenecks:

[0004] (1) It only supports the transmission of fixed semantic information and cannot flexibly reshape the user's semantic information according to the user's customized needs during the communication process (for example, modifying the facial expression during the communication process);

[0005] (2) The collaborative capabilities of multimodal data communication are insufficient. Currently, shallow feature interaction modes are often used (such as simple concatenation and fusion of text embedding vectors and image feature maps), which makes it difficult to establish fine-grained semantic association channels. There is a general lack of fine-grained semantic alignment mechanisms, making it impossible to dynamically capture implicit semantic associations in the multimodal data communication process.

[0006] (3) The communication bandwidth is inefficiently utilized. Currently, based on the goal of pixel-level reconstruction, data-oriented communication modes are often used to encode various parts of multimodal data indiscriminately. This cannot adapt to the dynamic needs of semantic reconstruction and easily leads to serious bandwidth waste.

[0007] (4) The adaptability to dynamic channels is insufficient. Currently, it can only perform well under a fixed high signal-to-noise ratio, but it is difficult to adapt to time-varying wireless channels. The performance in real-time face transmission scenarios fluctuates violently.

[0008] Therefore, the current facial semantic communication solutions have deficiencies in flexibility, collaboration, communication efficiency, and channel adaptability, and are unable to meet the diverse needs of current users for facial transmission services. Summary of the Invention

[0009] The embodiments of the present invention provide a multimodal facial semantic communication method, device and medium for solving the following technical problems: the current facial semantic communication solutions have deficiencies in flexibility, collaboration, communication efficiency and channel adaptability, and are no longer able to meet the current users' diverse needs for facial transmission services.

[0010] The embodiment of the present invention adopts the following technical solutions:

[0011] In one aspect, an embodiment of the present invention provides a multimodal face semantic communication method, the method comprising: obtaining a face video stream and a language reshaping instruction from a sending end;

[0012] Performing semantic transformation on the face image in the face video stream by using a generative semantic distillation encoder to obtain corresponding image semantic representation information;

[0013] Determining a semantic offset based on semantic features in the language reshaping instruction and facial features in the facial image; optimizing the image semantic representation information according to the semantic offset to obtain image semantic representation optimization information;

[0014] Optimize the encoding strategy of the transmission channel between the transmitter and the receiver, and fine-tune the parameters of the receiver's decoder based on the channel noise;

[0015] The image semantic representation optimization information is sent to the receiving end through the optimized transmission channel, and the semantic representation optimization information is decoded and the image is reconstructed through a decoder with fine-tuned parameters to realize face data transmission.

[0016] In a feasible implementation, semantic transformation is performed on the face image in the face video stream by a generative semantic distillation encoder to obtain corresponding image semantic representation information, specifically including:

[0017] Based on the inverse feature mapping technology of the pre-trained style generation network StyleGAN-XL, a nonlinear projection relationship model from the high-dimensional face data space to the compact semantic eigenspace is established. Dual regularization constraints are introduced into the nonlinear projection relationship model for optimization to obtain the generative semantic distillation encoder. The dual regularization constraints include pixel-level fidelity constraints and semantic consistency constraints.

[0018] Training the generative semantic distillation encoder using a facial image dataset to convert high-dimensional facial images into low-dimensional latent space representations while satisfying the dual regularization constraints;

[0019] The facial image in the facial video stream is input into a trained generative semantic distillation encoder to obtain a corresponding low-dimensional latent space representation, namely the image semantic representation information.

[0020] In a feasible implementation, the pixel-level fidelity constraint is: ;

[0021] The semantic consistency constraints are: ;

[0022] in, is the number of samples in the training batch, For the Original input face image, For pre-trained style generation network, For the The latent space encoding of samples, is a feature extraction function based on the pre-trained contrastive learning model CLIP, Calculates the cosine similarity function.

[0023] In a feasible implementation, based on the semantic features in the language reshaping instruction and the facial features in the facial image, a semantic offset is determined, and according to the semantic offset, the image semantic representation information is optimized to obtain image semantic representation optimization information, specifically including:

[0024] Extracting semantic features from the language reshaping instructions using a pre-trained language model BERT; wherein the semantic features include at least text embedding features;

[0025] extracting facial features from the facial image using an image feature extraction model;

[0026] Performing bidirectional dynamic attention calculation on the semantic features and the facial features to obtain an attention matrix between the language reshaping instructions and the facial features;

[0027] Generate the semantic offset based on the attention matrix , and the latent space representation of the face image The image semantic representation optimization information is obtained by adding them together.

[0028] In a feasible implementation, performing a bidirectional dynamic attention calculation on the semantic features and the facial features to obtain an attention matrix between the language reshaping instructions and the facial features specifically includes:

[0029] according to , get the attention matrix between the language reshaping instructions and facial features ;

[0030] in, The first Embedding features of text tokens, For the The facial features of the image blocks, is the feature dimension normalization factor, is the normalized exponential function.

[0031] In a feasible implementation, optimizing the coding strategy for the transmission channel between the transmitting end and the receiving end specifically includes:

[0032] Embedding a lightweight channel state predictor at the transmitting end, and estimating state parameters of the transmission channel in real time through the channel state predictor; wherein the state parameters at least include a time-varying signal-to-noise ratio;

[0033] According to the state parameter, a corresponding channel coding strategy is dynamically selected from a preset coding strategy library to optimize the coding strategy for the transmission channel.

[0034] In a feasible implementation, a lightweight channel state predictor is embedded in the transmitting end, and the state parameters of the transmission channel are estimated in real time by the channel state predictor, specifically including:

[0035] In the channel state predictor, according to , real-time estimation of the time-varying signal-to-noise ratio of the transmission channel ;

[0036] in, is the bit signal-to-noise ratio, is the sliding mean of the historical bit error rate, is a multi-layer perceptron predictor.

[0037] In a feasible implementation, fine-tuning parameters of a decoder at the receiving end based on channel noise specifically includes:

[0038] A differential noise simulation layer is constructed in the decoder at the receiving end, and mixed Gaussian noise is injected during the decoder training phase: ;

[0039] in, is the simulated channel noise, is the number of mixed analog noise components, For the The weights of the mixed simulated noise components, For the The variance of the mixed simulated noise components, is the Gaussian distribution function, I is the identity matrix;

[0040] The mixed Gaussian noise is used to enhance the training data of the decoder during the decoder training phase, and a small number of parameters in the decoder are fine-tuned through a parameter isolation fine-tuning method to adapt to the noise environment in the transmission channel.

[0041] On the other hand, an embodiment of the present invention also provides a multimodal facial semantic communication device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor so that the at least one processor can execute the multimodal facial semantic communication method.

[0042] Finally, an embodiment of the present invention also provides a storage medium, which is a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores at least one program, each of which includes instructions. When the instructions are executed by the terminal, the terminal executes the multimodal facial semantic communication method.

[0043] Compared with the prior art, the multimodal facial semantic communication method, device, and medium provided by the embodiments of the present invention have the following beneficial effects:

[0044] This paper proposes a multimodal facial semantic communication method that supports semantic reshaping. This method innovatively combines the inverse mapping technique of a generative adversarial network with a multimodal attention mechanism. By constructing a semantic eigenspace, it achieves efficient and reshapable transmission of facial features. The system uses a generative semantic distillation encoder to map high-dimensional facial images into compact semantic representations. A bidirectional dynamic attention network is designed to establish fine-grained associations between textual commands and visual attributes. This effectively addresses the limitations of traditional methods in terms of semantic reshaping flexibility, bandwidth efficiency, and multimodal association.

[0045] In terms of communication architecture, this system innovatively implements end-to-end joint optimization of semantic reconstruction and channel transmission. Through the feature compensation mechanism of channel noise perception, it adapts to different channel environments and significantly improves dynamic adaptability. In particular, the system introduces a lightweight parameter fine-tuning strategy to quickly adapt to time-varying channel conditions without changing the decoding backbone network structure, ensuring stable data transmission quality in complex wireless environments. This technical solution fundamentally breaks through the technical bottlenecks of existing systems in semantic plasticity, multimodal collaboration, bandwidth efficiency, dynamic adaptability, etc. through semantic-level feature extraction and multimodal joint optimization, providing a new solution for intelligent face transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments described in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0047] Figure 1 A flow chart of a multimodal face semantic communication method provided by an embodiment of the present invention;

[0048] Figure 2 A schematic structural diagram of a multimodal face semantic communication device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0050] The embodiment of the present invention provides a multimodal face semantic communication method, such as Figure 1 As shown, the multimodal face semantic communication method specifically includes steps S101-S104:

[0051] S101, obtaining a facial video stream and a language reshaping instruction from a sending end; performing semantic transformation on a facial image in the facial video stream through a generative semantic distillation encoder to obtain corresponding image semantic representation information.

[0052] Specifically, the system first collects the high-definition facial video stream of the user sending the video, as well as the user's voice-generated language reshaping instructions, and converts them into text. These reshaping instructions, such as natural language instructions like "increase facial brightness, reduce background noise," are used to adjust the image quality of a specific area in the user's current video.

[0053] After acquiring the facial video stream, the team traverses the video frames. For each facial frame, a nonlinear projection relationship model is established from the high-dimensional facial data space to the compact semantic eigenspace based on the inverse feature mapping technology of the pre-trained style generation network StyleGAN-XL. Dual regularization constraints are introduced into the nonlinear projection relationship model for optimization, resulting in a generative semantic distillation encoder. The dual regularization constraints include pixel-level fidelity constraints and semantic consistency constraints.

[0054] Furthermore, a generative semantic distillation encoder is trained using a facial image dataset, enabling it to convert high-dimensional facial images into low-dimensional latent space representations while satisfying dual regularization constraints. The facial images captured in the facial video stream are then sequentially fed into the trained generative semantic distillation encoder to obtain the corresponding low-dimensional latent space representations. This, in turn, yields image semantic representation information for each facial image. This allows for a precise and compact representation of user facial data in the latent space, reducing data transmission volume and effectively improving channel bandwidth utilization.

[0055] As a feasible implementation, the pixel-level fidelity constraint is: ; The semantic consistency constraints are: .

[0056] in, is the number of samples in the training batch, For the Original input face image, For pre-trained style generation network, For the The latent space encoding of samples, is a feature extraction function based on the pre-trained contrastive learning model CLIP, Calculates the cosine similarity function.

[0057] S102: Determine a semantic offset based on semantic features in the language reshaping instruction and facial features in the facial image; and optimize the image semantic representation information according to the semantic offset to obtain image semantic representation optimization information.

[0058] Specifically, the pre-trained language model BERT is used to extract semantic features from language reshaping instructions. These semantic features include at least text embedding features. Meanwhile, an image feature extraction model is used to extract facial features from facial images.

[0059] Furthermore, a bidirectional dynamic attention calculation is performed on the semantic features and facial features to obtain the attention matrix between the language reshaping instructions and the facial features. .in, The first feature in the text embedding Embedding features of text tokens, For the The facial features of the image blocks, is the feature dimension normalization factor, is the normalized exponential function.

[0060] Furthermore, the attention matrix is ​​used as the semantic offset , and the latent space representation of face images By adding them together, we can obtain the optimized information of image semantic representation, thereby realizing progressive semantic reshaping. The attention-guided network dynamically adjusts the offset direction of the latent space representation to reflect the user's language reshaping instruction content in the video image.

[0061] S103: Optimize the encoding strategy of the transmission channel between the transmitting end and the receiving end, and fine-tune the parameters of the decoder at the receiving end based on the channel noise.

[0062] Specifically, a lightweight channel state predictor is embedded at the transmitting end and used to estimate the state parameters of the transmission channel in real time. The state parameters include at least the time-varying signal-to-noise ratio (SNR). Based on the state parameters, a corresponding channel coding strategy is dynamically selected from a preset coding strategy library to optimize the coding strategy for the transmission channel.

[0063] As a feasible implementation, in the channel state predictor, according to , real-time estimation of the time-varying signal-to-noise ratio of the transmission channel .in, is the bit signal-to-noise ratio, is the sliding mean of the historical bit error rate, is a multi-layer perceptron predictor.

[0064] Furthermore, the decoder parameters at the receiving end are fine-tuned based on the channel noise. The specific implementation is as follows:

[0065] A differentiable noise simulation layer is constructed in the decoder at the receiving end, and mixed Gaussian noise is injected during the decoder training phase: ;in, is the simulated channel noise, is the number of mixed analog noise components, For the The weights of the mixed simulated noise components, For the The variance of the mixed simulated noise components, is the Gaussian distribution function, I is the identity matrix. This formula is used to express that the covariance structure of the noise is Gaussian noise with independent and identical distribution in each dimension.

[0066] By mixing Gaussian noise, the decoder training data is enhanced during the decoder training stage. During the training process, a small number of parameters in the decoder are fine-tuned through parameter isolation fine-tuning method to adapt to the noisy environment in the transmission channel. Only a small number of parameters need to be updated to adapt to dynamic channel environments of varying degrees.

[0067] S104 , sending the image semantic representation optimization information to the receiving end through the optimized transmission channel, and decoding the semantic representation optimization information and reconstructing the image through the decoder after parameter fine-tuning to realize face data transmission.

[0068] Specifically, the optimized image semantic representation information is sent to the receiving end through the transmission channel optimized by the coding strategy, and the semantic representation information is decoded by the decoder with fine-tuned parameters at the receiving end, and then the image is reconstructed to quickly reconstruct high-quality video images that meet the user's instruction requirements.

[0069] The above technical solution is further explained below through two embodiments:

[0070] Example 1:

[0071] In cross-border video conferencing scenarios, this invention can provide users in mobile, weak network environments with a real-time face transmission and reconstruction experience. The system's front-end equipment first captures the user's high-definition facial video stream. Then, using generative semantic distillation coding technology, it converts the complex visual information into a highly compressed semantic representation. When the user issues a natural language instruction, such as "increase facial brightness and reduce background noise," the system's built-in intelligent understanding module accurately interprets the text's intent, establishes a deep correlation with facial visual features, and performs feature reconstruction and optimization of the specified area within the semantic space. The system possesses intelligent channel awareness, enabling real-time monitoring and adaptation to the ever-changing wireless transmission environment. It dynamically adjusts encoding and decoding strategies to offset the impact of channel noise on signal quality. At the receiving end, a decoder, injected with mixed high-speed noise, rapidly reconstructs high-quality video footage based on the optimized semantic representation, ensuring clear presentation of key facial features while effectively suppressing background interference, providing users with a superior video conferencing experience. This invention integrates three core technologies: semantic understanding, feature reconstruction, and channel adaptation, improving transmission efficiency while maintaining visual quality.

[0072] Example 2:

[0073] In live streaming applications of virtual digital humans, the system uses high-precision motion capture equipment to capture rich facial expression data from performers in real time and convert it into structured semantic representations. When a performer issues natural language commands such as "exaggerate facial expressions," the system's built-in multimodal understanding engine deeply analyzes the command semantics, intelligently identifying areas of expression that require enhancement within the feature space, and achieving precise reshaping and optimization of facial features. The system intelligently senses changes in network channel conditions and dynamically adjusts transmission strategies to address various interference factors. Even in unstable mobile network environments, the system ensures the complete transmission of facial expression details, enabling real-time driving and rendering of digital humans in the cloud through a lightweight parameter transmission mechanism. This solves the problems of command response delay and expression distortion common in traditional digital human systems, providing a more natural and smooth interactive experience for applications such as virtual live streaming and remote performances. Through intelligent reshaping of the semantic space and adaptive processing of channel noise, the system ensures the precise transmission of facial expression details while improving the real-time and reliability of digital human driving.

[0074] The present invention demonstrates significant performance advantages in low signal-to-noise ratio environments. Under harsh channel conditions, the present invention's FID score (95.80) is 78.9% and 10.4% lower than the traditional JPEG&LDPC scheme (453.36) and DJSCC scheme (106.97), respectively, and the Total Variation loss is reduced by 68.1% and 30.4%, respectively, fully verifying the present invention's breakthrough improvement in image fidelity and detail recovery capabilities. This advantage stems from an innovative end-to-end neural network architecture. Through the coordinated optimization of deep feature compression and adaptive modulation technology, it effectively overcomes the performance degradation of existing methods in low signal-to-noise ratio environments, providing a more robust solution for dynamic semantic reconstruction of facial communication.

[0075] In addition, the embodiment of the present invention also provides a multimodal face semantic communication device, such as Figure 2 As shown, the equipment specifically includes:

[0076] at least one processor; and a memory communicatively connected to the at least one processor; wherein,

[0077] The memory stores instructions executable by at least one processor, so as to enable the at least one processor to perform:

[0078] Obtain the facial video stream and language reshaping instructions from the sender;

[0079] Performing semantic transformation on the face image in the face video stream by using a generative semantic distillation encoder to obtain corresponding image semantic representation information;

[0080] Determining a semantic offset based on semantic features in the language reshaping instruction and facial features in the facial image; optimizing the image semantic representation information according to the semantic offset to obtain image semantic representation optimization information;

[0081] Optimize the encoding strategy of the transmission channel between the transmitter and the receiver, and fine-tune the parameters of the receiver's decoder based on the channel noise;

[0082] The image semantic representation optimization information is sent to the receiving end through the optimized transmission channel, and the semantic representation optimization information is decoded and the image is reconstructed through a decoder with fine-tuned parameters to realize face data transmission.

[0083] Finally, the present invention also provides a storage medium, which is a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores at least one program, each of which includes instructions. When the instructions are executed by a terminal, the terminal executes:

[0084] Obtain the facial video stream and language reshaping instructions from the sender;

[0085] Performing semantic transformation on the face image in the face video stream by using a generative semantic distillation encoder to obtain corresponding image semantic representation information;

[0086] Determining a semantic offset based on semantic features in the language reshaping instruction and facial features in the facial image; optimizing the image semantic representation information according to the semantic offset to obtain image semantic representation optimization information;

[0087] Optimize the encoding strategy of the transmission channel between the transmitter and the receiver, and fine-tune the parameters of the receiver's decoder based on the channel noise;

[0088] The image semantic representation optimization information is sent to the receiving end through the optimized transmission channel, and the semantic representation optimization information is decoded and the image is reconstructed through a decoder with fine-tuned parameters to realize face data transmission.

[0089] The various embodiments of the present invention are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are simplified. For relevant details, refer to the descriptions of the method embodiments.

[0090] The above description of specific embodiments of the present invention is provided. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0091] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations may be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A multimodal face semantic communication method, characterized in that: The method comprises: Obtain the facial video stream and language reshaping instructions from the sender; Performing semantic transformation on the face image in the face video stream by using a generative semantic distillation encoder to obtain corresponding image semantic representation information; Determining a semantic offset based on semantic features in the language reshaping instruction and facial features in the facial image; optimizing the image semantic representation information according to the semantic offset to obtain image semantic representation optimization information; Optimize the encoding strategy for the transmission channel between the transmitter and receiver, and fine-tune the parameters of the receiver's decoder based on the channel noise. This includes: A lightweight channel state predictor is embedded in the transmitting end, and the state parameters of the transmission channel are estimated in real time by the channel state predictor; wherein the state parameters at least include a time-varying signal-to-noise ratio; specifically comprising: in the channel state predictor, according to , real-time estimation of the time-varying signal-to-noise ratio of the transmission channel ;in, is the bit signal-to-noise ratio, is the sliding mean of the historical bit error rate, is a multi-layer perceptron predictor; Dynamically selecting a corresponding channel coding strategy from a preset coding strategy library according to the state parameter to optimize the coding strategy for the transmission channel; A differential noise simulation layer is constructed in the decoder at the receiving end, and mixed Gaussian noise is injected during the decoder training phase: ; in, is the simulated channel noise, is the number of mixed analog noise components, For the The weights of the mixed simulated noise components, For the The variance of the mixed simulated noise components, is the Gaussian distribution function, I is the identity matrix; Performing training data enhancement on the decoder during the decoder training phase by using the mixed Gaussian noise, and fine-tuning a small number of parameters in the decoder by using a parameter isolation fine-tuning method to adapt to the noise environment in the transmission channel; The image semantic representation optimization information is sent to the receiving end through the optimized transmission channel, and the semantic representation optimization information is decoded and the image is reconstructed through a decoder with fine-tuned parameters to realize face data transmission.

2. A multimodal face semantic communication method according to claim 1, characterized in that: The facial image in the facial video stream is semantically transformed by a generative semantic distillation encoder to obtain corresponding image semantic representation information, specifically including: Based on the inverse feature mapping technology of the pre-trained style generation network StyleGAN-XL, a nonlinear projection relationship model from the high-dimensional face data space to the compact semantic eigenspace is established. Dual regularization constraints are introduced into the nonlinear projection relationship model for optimization to obtain the generative semantic distillation encoder. The dual regularization constraints include pixel-level fidelity constraints and semantic consistency constraints. Training the generative semantic distillation encoder using a facial image dataset to convert high-dimensional facial images into low-dimensional latent space representations while satisfying the dual regularization constraints; The facial image in the facial video stream is input into a trained generative semantic distillation encoder to obtain a corresponding low-dimensional latent space representation, namely the image semantic representation information.

3. A multimodal face semantic communication method according to claim 2, characterized in that: The pixel-level fidelity constraint is: ; The semantic consistency constraints are: ; in, is the number of samples in the training batch, For the Original input face images, For pre-trained style generation network, For the The latent space encoding of samples, is a feature extraction function based on the pre-trained contrastive learning model CLIP, Calculates the cosine similarity function.

4. A multimodal face semantic communication method according to claim 1, characterized in that: Determining a semantic offset based on semantic features in the language reshaping instruction and facial features in the facial image, and optimizing the image semantic representation information according to the semantic offset to obtain image semantic representation optimization information, specifically including: Extracting semantic features from the language reshaping instructions using a pre-trained language model BERT; wherein the semantic features include at least text embedding features; extracting facial features from the facial image using an image feature extraction model; Performing bidirectional dynamic attention calculation on the semantic features and the facial features to obtain an attention matrix between the language reshaping instructions and the facial features; Generate the semantic offset based on the attention matrix , and the latent space representation of the face image The image semantic representation optimization information is obtained by adding them together.

5. A multimodal face semantic communication method according to claim 4, characterized in that: Performing a bidirectional dynamic attention calculation on the semantic features and the facial features to obtain an attention matrix between the language reshaping instructions and the facial features, specifically including: according to , get the attention matrix between the language reshaping instructions and facial features ; in, The first Embedding features of text tokens, For the The facial features of the image blocks, is the feature dimension normalization factor, is the normalized exponential function.

6. A multimodal face semantic communication device, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, so that the at least one processor can execute the multimodal face semantic communication method according to any one of claims 1-5.

7. A storage medium, characterized in that: The storage medium is a non-volatile computer-readable storage medium, which stores at least one program, each of which includes instructions. When the instructions are executed by the terminal, the terminal executes a multimodal facial semantic communication method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Stretchable face image coding method and system for man-machine mixed vision

    CN115880762A

  • Multi-modal semantic communication method, system and equipment based on large model and medium

    CN118350416A

  • Channel environment adaptive sampling-semantic-channel coding joint optimization method and system

    CN119561650A