Multi-mode face semantic communication method and device and medium

The multi-modal face semantic communication method addresses flexibility, collaboration, and adaptability issues by transforming face images into compact semantic representations and optimizing channel conditions, ensuring efficient and stable face data transmission.

CN120318893AActive Publication Date: 2025-07-15BEIJING FACE WALL INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510789542.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-15
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

The existing face semantic communication solutions have shortcomings in flexibility, coordination, communication efficiency and channel adaptability, and cannot adapt to users' diverse needs for face transmission services.

Method used

Generative semantic distillation encoder and dual regularization constraint technology are used to convert high-dimensional face images into low-dimensional latent space characterization, and through fine-tuning of lightweight channel state predictors and decoder parameters, the optimization of semantic offsets and dynamic adjustment of channel coding strategies is achieved. Combined with multimodal attention mechanism and channel noise perception, end-to-end semantic reshaping and channel transmission optimization are achieved.

Benefits of technology

It significantly improves the flexibility of semantic remodeling and bandwidth utilization efficiency, enhances the coordination of multimodal data and the dynamic adaptability of channels, and ensures stable and efficient face data transmission in complex wireless environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318893A_ABST
    Figure CN120318893A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode face semantic communication method and device and a medium, belongs to the technical field of semantic communication, and aims to solve the defects of the current face semantic communication scheme in the aspects of flexibility, collaboration, communication efficiency and channel adaptability. And the diversified requirements of the current user on the face transmission service cannot be met. Comprising the steps of performing semantic transformation on a face image in a face video stream to obtain corresponding image semantic representation information; determining a semantic offset based on semantic features in the language remodeling instruction and face features in the face image; optimizing the image semantic representation information according to the semantic offset; encoding strategy optimization is carried out on a transmission channel between a sending end and a receiving end, and parameter fine adjustment is carried out on a decoder of the receiving end based on channel noise; and sending the image semantic representation optimization information to a receiving end through the optimized transmission channel, and carrying out decoding and image reconstruction on the semantic representation optimization information through a decoder after parameter fine tuning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic communication, and particularly to a multimodal face semantic communication method, device and medium. Background Art

[0002] With the rapid development of mobile Internet technology, the demand for face transmission technology in real-time scenarios such as social media and video conferencing has been increasing rapidly. As an emerging communication paradigm, semantic communication significantly reduces bandwidth requirements and improves transmission efficiency by extracting and understanding the deep meaning of information rather than simply transmitting raw data, and can provide a more intelligent face transmission solution. Further combined with the collaborative optimization of edge computing and terminal intelligent chips, accurate facial semantic parsing and reconstruction on the mobile side can be realized, providing high-fidelity, low-power real-time face transmission services for scenarios such as immersive remote interaction and virtual live broadcast.

[0003] However, the current existing semantic communication solutions still face the following bottlenecks in the application of face transmission tasks: (1) Only support fixed semantic information transmission, and cannot effectively reshape the user's semantic information according to the customized needs of users during the communication process (for example, modifying the facial expression during the communication process); (2) The collaborative ability of multimodal data communication is insufficient. Currently, a shallow feature interaction mode is often adopted (such as simply splicing and fusing text embedding vectors and image feature maps), which is difficult to establish a fine-grained semantic association channel, generally lacks a fine-grained semantic alignment mechanism, and cannot dynamically capture the implicit semantic associations during the multimodal data communication process; (3) The utilization efficiency of communication bandwidth is low. Currently, often based on the goal of pixel-level reconstruction, a data-oriented communication mode is used to encode each part of the multimodal data without discrimination, which cannot adapt to the dynamic requirements of semantic reshaping and is prone to serious bandwidth waste; (4) The adaptability to dynamic channels is insufficient. Currently, it often only performs well under a fixed high signal-to-noise ratio, but it is difficult to adapt to time-varying wireless channels, and the performance fluctuates violently in real-time face transmission scenarios.

[0004] Therefore, the current face semantic communication solutions have deficiencies in terms of flexibility, collaboration, communication efficiency, and channel adaptability, and can no longer meet the diverse needs of current users for face transmission services. Summary of the Invention

[0005] Embodiments of the present invention provide a multimodal face semantic communication method, device and medium, which are used to solve the following technical problems: The current face semantic communication solutions have deficiencies in terms of flexibility, collaboration, communication efficiency, and channel adaptability, and can no longer meet the diverse needs of current users for face transmission services.

[0006] The embodiments of the present invention adopt the following technical solutions: On the one hand, the embodiments of the present invention provide a multimodal face semantic communication method, and the method includes: obtaining a face video stream of a sending end and a language reshaping instruction; Semantically transforming the face images in the face video stream through a generative semantic distillation encoder to obtain corresponding image semantic representation information; Determining a semantic offset based on the semantic features in the language reshaping instruction and the face features in the face image; optimizing the image semantic representation information according to the semantic offset to obtain optimized image semantic representation information; Optimizing the encoding strategy of the transmission channel between the sending end and the receiving end, and fine-tuning the parameters of the decoder at the receiving end based on channel noise; Sending the optimized image semantic representation information to the receiving end through the optimized transmission channel, and decoding and image reconstruction of the semantic representation optimization information through the decoder with fine-tuned parameters to achieve face data transmission.

[0007] In a feasible implementation manner, semantically transforming the face images in the face video stream through a generative semantic distillation encoder to obtain corresponding image semantic representation information specifically includes: Based on the inverse feature mapping technology of the pre-trained StyleGAN-XL style generation network, establishing a non-linear projection relationship model from the high-dimensional face data space to the compact semantic eigen space, and introducing double regularization constraints for optimization in the non-linear projection relationship model to obtain the generative semantic distillation encoder; wherein, the double regularization constraints include pixel-level fidelity constraints and semantic consistency constraints; Training the generative semantic distillation encoder through a face image data set to enable it to convert high-dimensional face images into low-dimensional latent space representations on the premise of satisfying the double regularization constraints; Inputting the face images in the face video stream into the trained generative semantic distillation encoder to obtain corresponding low-dimensional latent space representations, that is, the image semantic representation information.

[0008] In a feasible implementation manner, the pixel-level fidelity constraint is: ; The semantic consistency constraint is: ; Wherein, is the number of samples in the training batch, is the th original input face image, is the pre-trained style generation network, is the latent space encoding of the th sample, is the feature extraction function based on the pre-trained contrastive learning model CLIP, is the cosine similarity calculation function.

[0009] In a feasible implementation manner, based on the semantic features in the language reshaping instruction and the face features in the face image, a semantic offset is determined, and according to the semantic offset, the image semantic representation information is optimized to obtain optimized image semantic representation information, which specifically includes: Extract the semantic features in the language reshaping instruction through the pre-trained language model BERT; wherein, the semantic features at least include text embedding features; Extract the face features in the face image through an image feature extraction model; Perform bidirectional dynamic attention calculation on the semantic features and the face features to obtain an attention matrix between the language reshaping instruction and the face features; Generate the semantic offset based on the attention matrix , and add it to the latent space representation of the face image to obtain the optimized image semantic representation information.

[0010] In a feasible implementation manner, performing bidirectional dynamic attention calculation on the semantic features and the face features to obtain an attention matrix between the language reshaping instruction and the face features specifically includes: According to , obtain the attention matrix between the language reshaping instruction and the face features ; wherein, is the embedding feature of the th text token in the text embedding features, is the face feature of the th image patch, is the feature dimension normalization factor, is the normalization exponential function.

[0011] In a feasible implementation manner, optimizing the encoding strategy of the transmission channel between the sender and the receiver specifically includes: Embed a lightweight channel state predictor at the sender, and use the channel state predictor to estimate the state parameters of the transmission channel in real time; wherein, the state parameters at least include time-varying signal-to-noise ratio; According to the state parameters, dynamically select a corresponding channel coding strategy in a preset coding strategy library to optimize the encoding strategy of the transmission channel.

[0012] In a feasible implementation manner, a lightweight channel state predictor is embedded at the sending end, and the state parameters of the transmission channel are estimated in real time through the channel state predictor, specifically including: In the channel state predictor, according to , the time-varying signal-to-noise ratio of the transmission channel is estimated in real time ; wherein, is the bit signal-to-noise ratio, is the moving average of the historical bit error rate, is the multi-layer perceptron predictor.

[0013] In a feasible implementation manner, the parameters of the decoder at the receiving end are fine-tuned based on the channel noise, specifically including: A differential noise simulation layer is constructed in the decoder at the receiving end, and mixed Gaussian noise is injected during the decoder training phase: ; wherein, is the simulated channel noise, is the number of mixed simulation noise components, is the weight of the th mixed simulation noise component, is the variance of the th mixed simulation noise component, is the Gaussian distribution function, I is the identity matrix; Through the mixed Gaussian noise, the training data of the decoder is enhanced during the decoder training phase, and a small number of parameters in the decoder are fine-tuned through the parameter isolation fine-tuning method to adapt to the noise environment in the transmission channel.

[0014] On the other hand, an embodiment of the present invention further provides a multimodal face semantic communication device, the device includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions that can be executed by the at least one processor, so that the at least one processor can execute the multimodal face semantic communication method.

[0015] Finally, an embodiment of the present invention further provides a storage medium, the storage medium is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, each program includes instructions, and when the instructions are executed by a terminal, the terminal executes the multimodal face semantic communication method.

[0016] Compared with the prior art, a multimodal face semantic communication method, device and medium provided by an embodiment of the present invention have the following beneficial effects: This paper proposes a multimodal face semantic communication method that supports semantic reshaping. It innovatively integrates the generative adversarial network inverse mapping technology and the multimodal attention mechanism, and realizes the reshaping and efficient transmission of facial features by constructing a semantic eigenspace. The system uses a generative semantic distillation encoder to map high-dimensional face images into compact semantic representations, and designs a bidirectional dynamic attention network to establish fine-grained associations between text instructions and visual attributes, effectively solving the limitations of traditional methods in terms of semantic reshaping flexibility, bandwidth utilization efficiency, and multimodal association.

[0017] In terms of communication architecture, this system innovatively realizes end-to-end joint optimization of semantic reconstruction and channel transmission. Through the feature compensation mechanism of channel noise perception, it adapts to different channel environments and significantly improves dynamic adaptability. In particular, the system introduces a lightweight parameter fine-tuning strategy to quickly adapt to time-varying channel conditions without changing the decoding backbone network structure, ensuring stable data transmission quality in complex wireless environments. This technical solution fundamentally breaks through the technical bottlenecks of existing systems in semantic plasticity, multimodal collaboration, bandwidth efficiency, dynamic adaptability, etc. through semantic-level feature extraction and multimodal joint optimization, providing a new solution for intelligent face transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 A flow chart of a multimodal face semantic communication method provided by an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a multimodal face semantic communication device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0020] The embodiment of the present invention provides a multimodal face semantic communication method, such as Figure 1As shown in the figure, the multimodal face semantic communication method specifically includes steps S101 - S104: S101. Obtain the face video stream of the sending end and the language reshaping instruction; perform semantic transformation on the face images in the face video stream through the generative semantic distillation encoder to obtain the corresponding image semantic representation information.

[0021] Specifically, first collect the high - definition face video stream of the user at the video sending end, and collect the language reshaping instruction issued by the user through voice and convert it into text form. Among them, the language reshaping instruction refers to a natural language instruction such as "enhance the facial brightness and reduce the background noise", which is used to adjust the picture quality of the specified area in the user's current video.

[0022] Furthermore, after obtaining the face video stream, traverse the video image frames therein. For each face image, based on the inverse feature mapping technology of the pre - trained style generation network StyleGAN - XL, establish a non - linear projection relationship model from the high - dimensional face data space to the compact semantic eigen - space, and introduce double regularization constraints in the non - linear projection relationship model for optimization to obtain the generative semantic distillation encoder. Among them, the double regularization constraints include pixel - level fidelity constraints and semantic consistency constraints.

[0023] Furthermore, train the generative semantic distillation encoder through the face image data set so that it can convert high - dimensional face images into low - dimensional latent space representations while meeting the double regularization constraints. Then, input the face images in the obtained face video stream into the trained generative semantic distillation encoder in sequence to obtain the corresponding low - dimensional latent space representations, that is, obtain the image semantic representation information of each face image, thereby realizing the accurate and compact representation of the user's face data in the latent space, reducing the data transmission volume, and effectively improving the channel bandwidth utilization rate.

[0024] As a feasible implementation, the pixel - level fidelity constraint is: ; the semantic consistency constraint is: .

[0025] Among them, is the number of samples in the training batch, is the th original input face image, is the pre - trained style generation network, is the th latent space encoding of the sample, is the feature extraction function based on the pre - trained contrastive learning model CLIP, is the cosine similarity calculation function.

[0026] S102. Determine a semantic offset based on the semantic features in the language reshaping instruction and the face features in the face image; optimize the image semantic representation information according to the semantic offset to obtain optimized image semantic representation information.

[0027] Specifically, extract the semantic features in the language reshaping instruction through the pre-trained language model BERT. Among them, the semantic features at least include text embedding features. At the same time, extract the face features in the face image through the image feature extraction model.

[0028] Furthermore, perform bidirectional dynamic attention calculation on the semantic features and the face features to obtain an attention matrix between the language reshaping instruction and the face features . Among them, is the embedding feature of the th text token in the text embedding features, is the face feature of the th image patch, is the feature dimension normalization factor, is the normalization exponential function.

[0029] Furthermore, use the attention matrix as the semantic offset , and add it to the latent space representation of the face image to obtain optimized image semantic representation information, thereby realizing progressive semantic reshaping, and dynamically adjusting the offset direction of the latent space representation through the attention guidance network to reflect the content of the user's language reshaping instruction in the video image.

[0030] S103. Optimize the coding strategy of the transmission channel between the sender and the receiver, and fine-tune the parameters of the decoder at the receiver based on the channel noise.

[0031] Specifically, embed a lightweight channel state predictor at the sender, and use the channel state predictor to estimate the state parameters of the transmission channel in real time; among them, the state parameters at least include the time-varying signal-to-noise ratio. According to the state parameters, dynamically select the corresponding channel coding strategy in the preset coding strategy library to optimize the coding strategy of the transmission channel.

[0032] As a feasible implementation, in the channel state predictor, according to , estimate the time-varying signal-to-noise ratio of the transmission channel in real time. Among them, is the bit signal-to-noise ratio, is the moving average of the historical bit error rate, is the multi-layer perceptron predictor.

[0033] Furthermore, fine-tune the parameters of the decoder at the receiver based on the channel noise. The specific implementation method is as follows: A differential noise simulation layer is constructed in the decoder at the receiving end, and mixed Gaussian noise is injected during the decoder training phase: ;in, is the simulated channel noise, is the number of mixed analog noise components, For the The weights of the mixed simulated noise components, For the The variance of the mixed simulated noise components, is the Gaussian distribution function, I is the identity matrix. This formula is used to express that the covariance structure of the noise is Gaussian noise with independent and identical distribution in each dimension.

[0034] By mixing Gaussian noise, the training data of the decoder is enhanced during the decoder training stage, and a small number of parameters in the decoder are fine-tuned through parameter isolation fine-tuning method during the training process to adapt to the noise environment in the transmission channel. Only a small number of parameters need to be updated to adapt to dynamic channel environments of different degrees.

[0035] S104, sending the image semantic representation optimization information to the receiving end through the optimized transmission channel, and decoding the semantic representation optimization information and reconstructing the image through the decoder after parameter fine-tuning to realize face data transmission.

[0036] Specifically, the optimized image semantic representation information is sent to the receiving end through the transmission channel optimized by the coding strategy, and the semantic representation information is decoded by a decoder with fine-tuned parameters at the receiving end, and then the image is reconstructed to quickly reconstruct high-quality video images that meet the user's command requirements.

[0037] The above technical solution is further explained below through two embodiments: Embodiment 1: In the scenario of cross-border video conferencing, the present invention can provide users in a mobile weak network environment with a real-time face transmission and reshaping experience. The front-end device of the system first collects the high-definition face video stream of the user, and then through the generative semantic distillation coding technology, converts the complex visual information into a highly compressed semantic representation. When the user issues a natural language instruction such as "increase the facial brightness and reduce the background noise", the intelligent understanding module built into the system can accurately analyze the text intention, establish a deep association with the facial visual features, and complete the feature reshaping and optimization of the specified area in the semantic space. The system has the intelligent channel perception ability, can monitor and adapt to the constantly changing wireless transmission environment in real time, and offset the influence of channel noise on the signal quality by dynamically adjusting the encoding and decoding strategies. At the receiving end, the decoder injected with the hybrid high-speed noise quickly reconstructs a high-quality video picture that meets the requirements of the user's instructions based on the optimized semantic representation, ensuring both the clear presentation of the key facial features and effectively suppressing the background interference, bringing a better video conferencing experience to the user. The present invention integrates three core technologies of semantic understanding, feature reshaping and channel adaptation, and improves the transmission efficiency while ensuring the visual quality.

[0038] Embodiment 2: In the application scenario of virtual digital human live broadcast, through high-precision motion capture devices, the system can collect rich facial expression data of the performer in real time and convert it into a structured semantic representation. When the performer issues a natural language instruction such as "exaggerate the expression amplitude", the multi-modal understanding engine built into the system can deeply analyze the instruction semantics, intelligently identify the expression area that needs to be strengthened in the feature space, and achieve precise expression feature reshaping and optimization. The system can intelligently sense the changes in the network channel conditions and dynamically adjust the transmission strategy to cope with various interference factors. Even in a mobile network environment with unstable signals, the system can ensure the complete transmission of the expression details and realize the real-time driving and rendering of the cloud digital human through a lightweight parameter transmission mechanism. It solves the problems of instruction response delay and expression distortion existing in the traditional digital human system, and provides a more natural and smooth interaction experience for application scenarios such as virtual live broadcast and remote performance. Through the intelligent reshaping of the semantic space and the adaptive processing of channel noise, the system improves the real-time performance and reliability of digital human driving while ensuring the accurate transmission of expression details.

[0039] The present invention exhibits significant performance advantages in a low signal-to-noise ratio environment. Under harsh channel conditions, the FID score of the present invention (95.80) is reduced by 78.9% and 10.4% respectively compared with the traditional JPEG&LDPC scheme (453.36) and the DJSCC scheme (106.97), and the Total Variation loss is reduced by 68.1% and 30.4% respectively compared with the two, fully verifying the breakthrough improvement of the present invention in image fidelity and detail restoration ability. This advantage stems from the innovative end-to-end neural network architecture, which effectively overcomes the performance degradation problem of existing methods in a low signal-to-noise ratio environment through the collaborative optimization of deep feature compression and adaptive modulation technology, providing a more robust solution for the dynamic semantic reshaping of face communication.

[0040] In addition, an embodiment of the present invention also provides a multimodal face semantic communication device, as Figure 2 shown, the device specifically includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, so that the at least one processor can execute: Obtain the face video stream of the sending end and the language reshaping instruction; Perform semantic transformation on the face images in the face video stream through a generative semantic distillation encoder to obtain corresponding image semantic representation information; Based on the semantic features in the language reshaping instruction and the face features in the face image, determine the semantic offset; according to the semantic offset, optimize the image semantic representation information to obtain optimized image semantic representation information; Optimize the encoding strategy of the transmission channel between the sending end and the receiving end, and fine-tune the parameters of the decoder at the receiving end based on the channel noise; Send the optimized image semantic representation information to the receiving end through the optimized transmission channel, and decode and reconstruct the image from the optimized semantic representation information through the decoder with fine-tuned parameters to achieve face data transmission.

[0041] Finally, the present invention also provides a storage medium, which is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, and each program includes instructions, and when the instructions are executed by a terminal, the terminal is enabled to execute: Obtain the face video stream of the sending end and the language reshaping instruction; Perform semantic transformation on the face images in the face video stream through a generative semantic distillation encoder to obtain corresponding image semantic representation information; Determine a semantic offset based on the semantic features in the language reshaping instruction and the face features in the face image; optimize the image semantic representation information according to the semantic offset to obtain optimized image semantic representation information; Optimize the encoding strategy for the transmission channel between the sender and the receiver, and fine-tune the parameters of the decoder at the receiver based on the channel noise; Send the optimized image semantic representation information to the receiver through the optimized transmission channel, and decode and reconstruct the image from the optimized semantic representation information through the decoder with fine-tuned parameters to achieve face data transmission.

[0042] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.

[0043] The above describes specific embodiments of the present invention. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0044] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-modal face semantic communication method, characterized in that, The method includes: Obtaining a face video stream of a sending end and a language reshaping instruction; Performing semantic transformation on a face image in the face video stream through a generative semantic distillation encoder to obtain corresponding image semantic representation information; Determining a semantic offset based on the semantic features in the language reshaping instruction and the face features in the face image; and optimizing the image semantic representation information according to the semantic offset to obtain optimized image semantic representation information; Optimizing an encoding strategy for a transmission channel between the sending end and the receiving end, and finely tuning parameters of a decoder at the receiving end based on channel noise; Sending the optimized image semantic representation information to the receiving end through the optimized transmission channel, and decoding and reconstructing an image from the optimized semantic representation information through the decoder with finely tuned parameters to implement face data transmission.

2. The multimodal face semantic communication method according to claim 1, characterized in that Performing semantic transformation on a face image in the face video stream through a generative semantic distillation encoder to obtain corresponding image semantic representation information, specifically including: Based on the inverse feature mapping technology of a pre-trained style generation network StyleGAN-XL, establishing a nonlinear projection relationship model from a high-dimensional face data space to a compact semantic eigen space, and introducing double regularization constraints for optimization in the nonlinear projection relationship model to obtain the generative semantic distillation encoder; wherein the double regularization constraints include pixel-level fidelity constraints and semantic consistency constraints; Training the generative semantic distillation encoder through a face image data set to enable it to convert a high-dimensional face image into a low-dimensional latent space representation on the premise of satisfying the double regularization constraints; Inputting the face image in the face video stream into the trained generative semantic distillation encoder to obtain a corresponding low-dimensional latent space representation, i.e., the image semantic representation information.

3. The multimodal face semantic communication method according to claim 2, characterized in that, The pixel-level fidelity constraint is as follows: ; The semantic consistency constraint is as follows: ; Among them, is the number of samples in the training batch, is the th original input face image, is the pre-trained style generation network, is the th sample's latent space encoding, is the feature extraction function based on the pre-trained contrastive learning model CLIP, is the cosine similarity calculation function.

4. A multimodal face semantic communication method according to claim 1, characterized in that Determining a semantic offset based on the semantic features in the language reshaping instruction and the face features in the face image, and optimizing the image semantic representation information according to the semantic offset to obtain optimized image semantic representation information, specifically including: Extracting the semantic features in the language reshaping instruction through a pre-trained language model BERT; wherein the semantic features at least include text embedding features; Extracting the face features in the face image through an image feature extraction model; Performing bidirectional dynamic attention calculation on the semantic features and the face features to obtain an attention matrix between the language reshaping instruction and the face features; Generating the semantic offset based on the attention matrix and adding it to the latent space representation of the face image to obtain the optimized information of the image semantic representation.

5. A multimodal face semantic communication method according to claim 4, characterized in that, Performing bidirectional dynamic attention calculation on the semantic features and the face features to obtain an attention matrix between the language reshaping instruction and the face features, specifically including: According to , an attention matrix between the language reshaping instruction and the face feature is obtained ; Among them, is the embedding feature of the th text token in the text embedding feature, is the face feature of the th image patch, is the feature dimension normalization factor, is the normalization exponential function.

6. A multimodal face semantic communication method according to claim 1, characterized in that Optimizing an encoding strategy for a transmission channel between the sending end and the receiving end, specifically including: Embedding a lightweight channel state predictor at the sending end, and estimating state parameters of the transmission channel in real time through the channel state predictor; wherein the state parameters at least include a time-varying signal-to-noise ratio; Dynamically selecting a corresponding channel coding strategy from a preset encoding strategy library according to the state parameters to optimize the encoding strategy for the transmission channel.

7. A multimodal face semantic communication method according to claim 6, characterized in that Embed a lightweight channel state predictor at the sending end, and use the channel state predictor to estimate the state parameters of the transmission channel in real time, specifically including: In the channel state predictor, according to , the time-varying signal-to-noise ratio of the transmission channel is estimated in real time ; Among them, is the bit signal-to-noise ratio, is the moving average of the historical bit error rate, is the multi-layer perceptron predictor.

8. A multimodal face semantic communication method according to claim 1, characterized in that Fine-tune the parameters of the decoder at the receiving end based on channel noise, specifically including: Construct a differential noise simulation layer in the decoder of the receiving end and inject mixed Gaussian noise during the decoder training phase: ; in, is the simulated channel noise, is the number of mixed analog noise components, For the The weights of the mixed simulated noise components, For the The variance of the mixed simulated noise components, is the Gaussian distribution function, I is the identity matrix; Use the mixture Gaussian noise to perform training data augmentation on the decoder during the decoder training phase, and use the parameter isolation fine-tuning method to fine-tune a small number of parameters in the decoder to adapt to the noise environment in the transmission channel.

9. A multimodal face semantic communication device, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, so that the at least one processor can execute a multimodal face semantic communication method according to any one of claims 1-8.

10. A storage medium, characterized in that, The storage medium is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, each program includes instructions, and when the instructions are executed by the terminal, the terminal executes a multimodal face semantic communication method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Stretchable face image coding method and system for man-machine mixed vision

    CN115880762A

  • Image cognition semantic communication system and method based on multi-modal knowledge graph

    CN118260432A

  • Multi-modal semantic communication method, system and equipment based on large model and medium

    CN118350416A

  • Image editing method based on prior constraint inversion algorithm

    CN119107388A

  • Channel environment adaptive sampling-semantic-channel coding joint optimization method and system

    CN119561650A

Cited By

  • Face identity exchange method, system and equipment

    CN121190619A

  • Generative semantic communication system for 3D content generation

    CN121547151A

  • Image splicing method, evaluation method and system

    CN122453602A