Intelligent image coding method oriented to man-machine mixed vision

By introducing textual semantic information into the image compression process, and designing a text-guided coding network and residual compensation strategy, the problem of insufficient utilization of semantic information in existing technologies is solved, thereby improving the efficiency and quality of human-computer hybrid visual image coding.

CN121864967APending Publication Date: 2026-04-14TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing image compression methods for human-machine hybrid vision do not fully utilize semantic information, resulting in a large amount of task-irrelevant redundant information in machine vision features. This leads to blurred textures and loss of details in semantically salient regions, limiting the improvement of compression performance.

Method used

By introducing textual semantic information as explicit guidance, and constructing a global semantic prior based on text description, we design a text-guided machine vision layer coding network and a human eye vision layer coding network. Combined with residual compensation strategies, we improve the efficiency and quality of image coding.

Benefits of technology

It achieves efficient compression while improving the performance of downstream machine vision tasks and the quality of human visual reconstruction, thereby increasing coding efficiency and the perceptual quality and fidelity of reconstructed images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121864967A_ABST
    Figure CN121864967A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent image coding method oriented to man-machine mixed vision, and the method comprises the steps: constructing global semantic priori generated based on text description, which is used for providing explicit semantic guidance for feature coding and an image reconstruction process; constructing a text-guided machine vision layer coding network, designing a machine vision-oriented text-guided compression normal form, and fusing text information into transform coding, entropy coding and transform decoding to obtain feature representation; a text-guided human eye vision layer coding network is constructed, a semantic-driven reconstruction mechanism is designed, the reconstruction mechanism takes visual features and text description extracted by machine vision branches as conditional priori, and a high-perceptual-quality image is generated through a semantic-driven generation prediction module in a generation stage. And a residual compensation strategy is introduced at the tail end to ensure pixel-level fidelity, and a reconstruction result is obtained. According to the invention, the performance of the man-machine hybrid vision task is effectively improved while the code stream overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep learning and image coding, and in particular to an intelligent image coding method for human-computer hybrid vision. Background Technology

[0002] With the continuous development of artificial intelligence technology in the field of image analysis, image data is no longer solely used for human visual perception but is increasingly being used for intelligent system analysis. In applications such as intelligent surveillance, autonomous driving, and telemedicine, image data needs to possess both good visual quality to meet the perceptual needs of the human eye and rich, accurate semantic information to support machine vision analysis. These human-machine collaborative scenarios often involve the transmission and storage of massive amounts of image data, posing a significant challenge to communication bandwidth and storage resources. Therefore, there is an urgent need to develop efficient intelligent image coding methods for human-machine hybrid vision.

[0003] In recent years, researchers have made initial explorations into intelligent image coding for human-machine hybrid vision. Some methods improve traditional codecs by employing task-adaptive quantization or spatial partitioning strategies to effectively preserve key information required for downstream visual tasks while improving compression efficiency and reconstructed image quality. Other methods focus on extracting general features based on deep learning to simultaneously support high-fidelity visual reconstruction and high-precision machine vision analysis. However, these methods fail to fully exploit the intrinsic correlation between features required by human vision and those required by machine vision, thus limiting further improvements in coding performance. To overcome these limitations, researchers have proposed scalable coding methods for human-machine hybrid vision. These methods typically employ a hierarchical scalable coding architecture, where the base layer focuses on extracting and encoding compact semantic information to support efficient machine analysis; the enhancement layer uses the decoding information from the base layer as prior conditions to further encode image detail information to achieve high-quality reconstruction for human vision. For example, Choi et al. designed an image codec for human-machine hybrid vision that achieves scalable coding by progressively transmitting different parts of the latent representation, thereby supporting efficient machine vision decoding while achieving high-quality image reconstruction for human vision. Li et al. proposed a task-adaptive learnable embedding quantization method, which selects the optimal quantization step size through a learnable step size predictor, effectively mining inter-layer correlations to achieve scalable coding for human-machine hybrid vision.

[0004] While these methods have achieved some performance improvements, existing compression frameworks have not fully incorporated semantic prior guidance, making it difficult to accurately preserve features relevant to both machine vision and human vision tasks. Existing image compression methods for human-machine hybrid vision do not fully utilize semantic information, lacking explicit semantic guidance for feature encoding and image reconstruction processes. This not only results in machine vision features containing a large amount of task-irrelevant redundant information but also leads to blurred textures and loss of detail in semantically salient regions, limiting further improvements in compression performance in human-machine collaborative scenarios. Summary of the Invention

[0005] This invention provides an intelligent image encoding method for human-machine hybrid vision. It utilizes textual semantic information as explicit semantic guidance to achieve efficient compression while improving the performance of human-machine hybrid vision tasks. The invention designs a text-guided compression paradigm for machine vision, integrating textual semantic information into the transform coding, entropy coding, and transform decoding stages to achieve compact feature representation, improve the probability estimation accuracy of the entropy model, and enhance feature expression at the decoding end. This effectively improves the performance of downstream machine vision tasks while reducing bitstream overhead. Furthermore, the invention designs a semantically driven reconstruction mechanism, using visual features extracted from the machine vision branch and textual descriptions as conditional priors. In the generation stage, a semantically driven generation prediction module generates high-perceptual-quality images, and a residual compensation strategy is introduced at the end to ensure pixel-level fidelity, thereby obtaining high-quality reconstruction results. Details are described below. A smart image encoding method for human-computer hybrid vision, the method comprising: Construct a global semantic prior based on text description to provide explicit semantic guidance for feature encoding and image reconstruction processes; A text-guided machine vision layer coding network was constructed, and a text-guided compression paradigm for machine vision was designed. Text information was integrated into transform coding, entropy coding, and transform decoding to obtain feature representations. A text-guided human visual layer encoding network was constructed, and a semantic-driven reconstruction mechanism was designed. The reconstruction mechanism uses visual features extracted from the machine vision branch and text descriptions as conditional priors. In the generation stage, a high-perceptual-quality image is generated through a semantic-driven generation prediction module, and a residual compensation strategy is introduced at the end to ensure pixel-level fidelity and obtain the reconstruction result.

[0006] The method further includes: training an intelligent image coding network for human-computer hybrid vision, and encoding images based on the intelligent image coding network to achieve efficient processing of human-computer hybrid vision tasks.

[0007] The construction of a global semantic prior based on text description, used to provide explicit semantic guidance for feature encoding and image reconstruction processes, is as follows: Using a pre-trained graph-to-text model to process the input image Text description generation is performed to obtain a text description representing the semantics of the image scene. ; The text description is processed using a pre-trained CLIP text encoder to obtain text features at the encoding end. The text description is losslessly compressed, and the generated bitstream is decompressed after being transmitted to the decoding end. ; The reconstructed text description is also processed by the pre-trained CLIP text encoder to obtain the text features at the decoding end. ; Extracted text features and and the corresponding text description and Together, they constitute a global semantic prior generated based on text descriptions, which is used to guide feature encoding and image reconstruction.

[0008] The text-guided transformation encoding is as follows: By leveraging the intrinsic correlation between visual and textual features to obtain a more compact visual feature representation of an image, a text-guided feature compression module is proposed. This module obtains text-guided encoded features by fusing textual features to focus on semantically salient regions within visual features. .

[0009] The entropy encoding of the text guidance is as follows: By treating textual features as additional semantic priors and leveraging the semantic consistency between textual and visual features, the distribution of latent features is estimated. A text-guided prior fusion module is proposed, which integrates decoded super-prior features. Text features at the decoding end Fusion to obtain semantic perception priors .

[0010] The text-guided transformation decoding is as follows: The semantic information contained in the text description is used as prior knowledge to enhance the decoding features. A text-guided feature enhancement module is proposed to reconstruct the visual features. Text features at the decoding end By combining these features, we obtain the decoding characteristics of the text guidance. .

[0011] In the generation stage, the semantically driven generation prediction module adopts a conditional generation architecture based on a diffusion model and integrating ControlNet control branches. It utilizes machine vision features to provide semantically relevant structural information, and text descriptions to provide fine-grained semantic details, ensuring the semantic consistency and perceptual quality of the generated images.

[0012] The method further includes: concatenating the noisy features and machine vision decoding features as input to the ControlNet control branch, and integrating the text embedding into each layer of the ControlNet control branch and the diffusion model through a cross-attention mechanism to achieve dual guidance of structure and semantics.

[0013] The beneficial effects of the technical solution provided by this invention are: 1. This invention improves the performance of human-machine hybrid vision image coding by introducing textual semantic information as explicit guidance. This invention designs an image coding network for human-machine hybrid vision, which explicitly integrates textual semantic information into the image compression process, providing clear semantic guidance for coding, thereby simultaneously enhancing the performance of downstream machine vision tasks and the image reconstruction quality under human visual perception, and effectively improving coding efficiency. 2. This invention designs a text-guided compression paradigm for machine vision, which integrates text information into transform coding, entropy coding and transform decoding to achieve compact representation of features, assist in probability estimation of the entropy model and enhance feature expression at the decoding end. This significantly improves the performance of downstream machine vision tasks while improving compression efficiency. 3. This invention addresses the need for high-quality reconstruction of human vision by designing a semantically driven reconstruction mechanism. It utilizes machine vision features and text descriptions as joint conditional inputs to guide a diffusion-based generative model to predict reconstructed images. Combined with a residual transfer mechanism, this significantly improves the perceptual quality and fidelity of the reconstructed images. Attached Figure Description

[0014] Figure 1 This is a flowchart of an intelligent image encoding method for human-computer hybrid vision. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below.

[0016] Text, as an effective carrier of high-level semantic information, can serve as explicit semantic guidance in compression tasks oriented towards human-computer hybrid vision. Furthermore, transmitting text information requires less bitrate compared to transmitting image data. Based on this, this invention proposes an intelligent image encoding method for human-computer hybrid vision, which aims to achieve efficient compression while improving the perceptual quality of reconstructed images and the performance of visual tasks by introducing text information as explicit semantic guidance.

[0017] The following examples illustrate the specific implementation of the intelligent image encoding method for human-computer hybrid vision in this invention.

[0018] I. Constructing a global semantic prior based on text description generation We construct a global semantic prior based on text descriptions to provide explicit semantic guidance for subsequent feature encoding and image reconstruction processes.

[0019] Specifically, embodiments of the present invention first utilize a pre-trained graph-to-text model to process the input image. Perform text description generation to obtain text descriptions that can represent the semantics of the image scene. This text describes It includes key information such as the main targets, attributes, and semantic relationships in the scene. Subsequently, a pre-trained CLIP text encoder is used to process the text description. Processing is performed to obtain the text features at the encoding end. Meanwhile, the text description The lossless compressed bitstream is then transmitted to the decoding end and decompressed into... The reconstructed text description is also processed by a pre-trained CLIP text encoder to obtain the text features at the decoding end. Ultimately, the extracted text features ( and ) and the corresponding text description ( and Together, these constitute a global semantic prior generated based on text descriptions, providing explicit semantic guidance for subsequent text-guided machine vision coding networks and human eye vision coding networks, which is used to guide feature encoding and image reconstruction.

[0020] II. Constructing a Text-Guided Machine Vision Layer Encoding Network Considering that the rich semantic information contained in text descriptions can guide compression models or networks to focus more on semantically key regions, enabling them to retain the information most relevant to downstream tasks while suppressing low-level visual details that are less important to machine analysis, a machine vision-oriented text-guided compression paradigm was designed in the text-guided machine vision layer encoding network. This paradigm integrates text information into transform coding, entropy coding, and transform decoding to obtain compact and semantically rich feature representations, thereby improving compression efficiency and task performance.

[0021] (1) Text-guided transformation coding Text descriptions contain semantic information such as the target object and scene context, providing a valuable reference for transforming and encoding machine vision features. Therefore, leveraging the inherent correlation between visual and textual features can yield a more compact visual feature representation of the image. Specifically, the input image... First, visual features are obtained through an encoder. The formula is expressed as follows:

[0022] in, This indicates the encoder.

[0023] To achieve a more compact feature representation, this invention proposes a text-guided feature compression module. This module fuses text features to focus on semantically salient regions in visual features, thereby obtaining text-guided encoded features. The formula is expressed as follows:

[0024] in, This indicates a text-guided feature compression module.

[0025] (2) Text-guided entropy coding In machine vision layer coding networks, the features to be encoded contain a large amount of high-level semantic information related to downstream tasks, and text features also carry rich semantic context. Therefore, text features can be used as additional semantic priors to improve the accuracy of latent feature distribution estimation.

[0026] Specifically, text-guided coding features First, the hyper-prior features are obtained through a hyper-prior encoder. Super-prior features After transmission, the decoded prior features are reconstructed by the prior decoder at the decoding end. To effectively utilize the semantic consistency between textual and visual features and more accurately estimate the distribution of latent features, this invention proposes a text-guided prior fusion module. This prior fusion module integrates decoded super-prior features... Text features at the decoding end Fusion, thereby obtaining semantic perception priors The formula is expressed as follows:

[0027] in, This represents the prior fusion module guided by text.

[0028] Subsequently, semantic-aware priors are combined with an autoregressive entropy model to model the probability distribution:

[0029] in, express The first in One element, express semantic prior knowledge express The autoregressive prior.

[0030] Furthermore, Gaussian conditional models are used to assess probabilities. Perform parametric modeling as follows:

[0031] Where N( ) represents a Gaussian distribution. and These respectively represent the distribution with respect to the first... element The mean and standard deviation.

[0032] The parameters of the Gaussian model are obtained from prior estimates:

[0033] in, Let represent the parameter estimation function of the Gaussian model, and be composed of stacked 1 1. Convolution implementation.

[0034] (3) Text-guided transformation decoding Due to quantization operations, decoded visual features are often affected by semantic information degradation. Semantic information contained in text descriptions can be used as prior knowledge to enhance decoded features.

[0035] Specifically, this invention proposes a text-guided feature enhancement module that reconstructs visual features. Text features at the decoding end By combining these features, we obtain the decoding characteristics of the text guidance. The formula is expressed as follows:

[0036] in, This represents a feature enhancement module for text-guided presentations.

[0037] Then, the text-guided decoding features will be used. After decoding, a reconstructed image is generated for machine analysis. The formula is expressed as follows:

[0038] in, This is the decoder. Finally, the image will be reconstructed. Input the visual task network for downstream task processing.

[0039] III. Constructing a Text-Guided Human Visual Layer Encoding Network Considering that text descriptions can compensate for fine-grained semantic details that may be damaged or lost during compression, thereby improving the perceptual quality of reconstructed images, a semantic-driven reconstruction mechanism was designed in a text-guided human visual layer encoding network. This mechanism uses visual features extracted from the machine vision branch and text descriptions as conditional priors. In the generation stage, a semantically driven generation prediction module generates high-perceptual-quality images, and a residual compensation strategy is introduced at the end to ensure pixel-level fidelity, thus obtaining high-quality reconstruction results.

[0040] During the generation phase, the semantically driven generation prediction module adopts a conditional generation architecture based on a diffusion model and integrating ControlNet control branches. It utilizes machine vision features to provide semantically relevant structural information, and text descriptions to provide fine-grained semantic details, thereby ensuring the semantic consistency and perceptual quality of the generated images.

[0041] Specifically, the original image First, the encoder is mapped from the pre-trained variational autoencoder to the latent representation. Simultaneously, a pre-trained text encoder is used to describe the text. Processing is performed to obtain text features. The formula is expressed as follows:

[0042]

[0043] in, This represents the encoder of the pre-trained variational autoencoder. This represents a pre-trained text encoder.

[0044] Subsequently, in the potential space Applying a forward diffusion process, following the Markov chain rule, Gaussian noise is gradually added to obtain the latent features of the noise. For any time step Latent features after adding noise It can be obtained directly through sampling using the following formula:

[0045] in, This represents noise sampled from a standard normal distribution. For noise scheduling parameters The cumulative term, namely:

[0046] in, For the first The signal retention coefficient of the step.

[0047] In the reverse denoising process, the network predicts the noise injected at each time step and gradually recovers the latent features. Specifically, ControlNet reuses the encoder and intermediate blocks of the pre-trained diffusion model U-Net as the backbone network, and its output is injected into the decoder and intermediate blocks of the backbone U-Net through zero convolutional layers.

[0048] To achieve precise conditional control, noisy features and machine vision decoding features are concatenated as input to ControlNet. Simultaneously, text embeddings are integrated into each layer of ControlNet and the diffusion model via a cross-attention mechanism, thus achieving dual guidance in both structure and semantics. The network optimizes its parameters by minimizing the difference between predicted and actual noise. The noise prediction process of the denoising network can be represented as:

[0049] in, Represents a noise prediction network. and These represent the noise prediction subnetworks corresponding to the backbone U-Net and ControlNet branches, respectively. Indicates zero convolution operation, These are visual features decoded by the machine vision layer network.

[0050] At each time step, the network removes the predicted noise from the current latent representation. The latent feature representation of the previous time step is obtained. This iterative denoising process is represented as follows:

[0051] in, To control the variance term for sampling randomness, It is random noise.

[0052] After multiple iterations, the network finally recovers the initial latent representation. The decoder of the pre-trained variational autoencoder decodes the image, generating a high-quality predicted image that simultaneously meets the constraints of text description and visual features. The formula is expressed as follows:

[0053] in, This represents the decoder of the pre-trained variational autoencoder.

[0054] Although the generated images perform well in terms of perceptual quality and semantic consistency, they may still deviate from the original images in terms of fine texture and local structure, resulting in insufficient pixel-level fidelity. To compensate for this gap and further improve pixel-level reconstruction quality, the semantically driven reconstruction mechanism integrates a residual compensation strategy at the end.

[0055] At the encoder end, the residual between the original image and the predicted image is first calculated:

[0056] Finally, a lightweight codec is used to process the residuals. Compress and transmit.

[0057] At the decoding end, the reconstructed residual Add it back to the predicted image to obtain the final reconstructed image. , represented as:

[0058] IV. Training an intelligent image coding network for human-computer hybrid vision The embodiments of the present invention employ a two-stage training approach to optimize human-machine hybrid visual coding performance.

[0059] In the first stage, the rate-distortion loss function is used. End-to-end joint optimization is performed on the machine vision layer encoding network. Specifically, This includes a distortion term and a bit rate consumption term. Specifically, the distortion term is constructed as a linear combination of the visual task loss and the reconstruction distortion calculated from the mean squared error. This ensures the effectiveness of the encoded features in downstream machine vision tasks and subsequent image reconstruction. The formula is expressed as follows:

[0060] in, Indicates the distortion term. Indicates visual task loss, This indicates the reconstruction distortion calculated based on the mean square error.

[0061] Finally, the rate-distortion loss function The calculation formula is:

[0062] in, Indicates the bit rate consumption of feature transmission. It is a hyperparameter that controls the balance between bit rate and visual task performance.

[0063] The second stage involves training the human vision layer encoding network. This is achieved by first optimizing the ControlNet by minimizing the difference between predicted noise and actual noise in the latent space.

[0064] in, It is the noise prediction loss. yes Noise that is actually added to the latent variables at any given moment. Indicates at time step Noisy latent variables, Indicates the time step. Represents text information. These are visual features decoded by the machine vision layer network. It is noise in the model prediction. This represents the expectation of the joint distribution of the noisy latent variable, the time step, and the corresponding condition variable.

[0065] Subsequently, while keeping the weights of the machine vision layer encoding network and the semantically driven generative prediction module unchanged, the rate-distortion loss function was utilized. The residual compensation network is trained. The loss function for this stage is expressed as:

[0066] in, This represents the bit rate of residual transmission in the human visual layer coding network. It is a hyperparameter that balances bit rate and reconstruction quality in human vision.

[0067] V. Encoding images based on a trained intelligent image coding network The image to be encoded first obtains a text description through a pre-trained graph-to-text model. This description is then directly processed by a pre-trained CLIP text encoder to obtain text features at the encoding end. Simultaneously, after lossless compression, the text description is transmitted to the decoding end and then processed again by a pre-trained CLIP text encoder to obtain text features at the decoding end, thereby obtaining global semantic priors.

[0068] For a text-guided machine vision encoding network, the image to be encoded first undergoes an encoder to extract visual features. These visual features, along with the text features from the encoding end, are then processed by a text-guided feature compression module to obtain text-guided encoded features. These features are then decoded by a super-prior network to obtain super-prior features, which are subsequently fused with the text features from the decoding end through a text-guided prior fusion module to obtain semantic-aware priors. The semantic-aware priors and autoregressive priors are used together to estimate the entropy model parameters of the text-guided encoded features, achieving efficient entropy encoding and decoding, thereby recovering the reconstructed visual features. The reconstructed visual features are then combined with the text features from the decoding end through a text-guided feature enhancement module to obtain text-guided decoded features. These features are then used by a decoder to generate a reconstructed image for machine analysis, which is finally input into the downstream vision task network for task inference.

[0069] For a text-guided human vision layer encoding network, the text-guided decoding features obtained from the machine vision layer are first input into a semantically driven generative prediction module along with the text description at the decoding end to generate a predicted image. Then, the residual between the original image and the predicted image is calculated, and the residual features are extracted by the encoder and compressed into a bitstream for transmission. The decoding end restores the residual features and reconstructs the residual image, which is then added to the predicted image to obtain the final reconstructed image for human viewing.

[0070] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0071] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An intelligent image coding method for human-computer hybrid vision, characterized in that, The method includes: Construct a global semantic prior based on text description to provide explicit semantic guidance for feature encoding and image reconstruction processes; A text-guided machine vision layer coding network was constructed, and a text-guided compression paradigm for machine vision was designed. Text information was integrated into transform coding, entropy coding, and transform decoding to obtain feature representations. A text-guided human visual layer encoding network was constructed, and a semantic-driven reconstruction mechanism was designed. The reconstruction mechanism uses visual features extracted from the machine vision branch and text descriptions as conditional priors. In the generation stage, a high-perceptual-quality image is generated through a semantic-driven generation prediction module, and a residual compensation strategy is introduced at the end to ensure pixel-level fidelity and obtain the reconstruction result.

2. The intelligent image encoding method for human-computer hybrid vision according to claim 1, characterized in that, The method further includes: training an intelligent image coding network for human-computer hybrid vision, and encoding images based on the intelligent image coding network to achieve efficient processing of human-computer hybrid vision tasks.

3. The intelligent image encoding method for human-computer hybrid vision according to claim 1, characterized in that, The construction of a global semantic prior based on text description, used to provide explicit semantic guidance for feature encoding and image reconstruction processes, is as follows: Using a pre-trained graph-to-text model to process the input image Text description generation is performed to obtain a text description representing the semantics of the image scene. ; The text description is processed using a pre-trained CLIP text encoder to obtain text features at the encoding end. The text description is losslessly compressed, and the generated bitstream is decompressed after being transmitted to the decoding end. ; The reconstructed text description is also processed by the pre-trained CLIP text encoder to obtain the text features at the decoding end. ; Extracted text features and and the corresponding text description and Together, they constitute a global semantic prior generated based on text descriptions, which is used to guide feature encoding and image reconstruction.

4. The intelligent image encoding method for human-computer hybrid vision according to claim 1, characterized in that, The text-guided transformation encoding is as follows: By leveraging the intrinsic correlation between visual and textual features to obtain a more compact visual feature representation of an image, a text-guided feature compression module is proposed. This module obtains text-guided encoded features by fusing textual features to focus on semantically salient regions within visual features. .

5. The intelligent image encoding method for human-computer hybrid vision according to claim 1, characterized in that, The entropy encoding of the text guidance is as follows: By treating textual features as additional semantic priors and leveraging the semantic consistency between textual and visual features, the distribution of latent features is estimated. A text-guided prior fusion module is proposed, which integrates decoded super-prior features. Text features at the decoding end Fusion to obtain semantic perception priors .

6. The intelligent image encoding method for human-computer hybrid vision according to claim 1, characterized in that, The text-guided transformation decoding is as follows: The semantic information contained in the text description is used as prior knowledge to enhance the decoding features. A text-guided feature enhancement module is proposed to reconstruct the visual features. Text features at the decoding end By combining these features, we obtain the decoding characteristics of the text guidance. .

7. The intelligent image encoding method for human-computer hybrid vision according to claim 1, characterized in that, In the generation stage, the semantically driven generation prediction module adopts a conditional generation architecture based on a diffusion model and integrating ControlNet control branches. It utilizes machine vision features to provide semantically relevant structural information, and text descriptions to provide fine-grained semantic details, ensuring the semantic consistency and perceptual quality of the generated images.

8. The intelligent image encoding method for human-computer hybrid vision according to claim 7, characterized in that, The method further includes: concatenating the noisy features and machine vision decoding features as input to the ControlNet control branch, and integrating the text embedding into each layer of the ControlNet control branch and the diffusion model through a cross-attention mechanism to achieve dual guidance of structure and semantics.