Image processing method and device based on artificial intelligence, computer equipment and medium
Through the combination of hierarchical encoders, alignment modules and cross-modal adapters, the problem of inaccurate posture control in text-to-image generation is solved, and high-quality target images are generated to meet the needs of high-precision application scenarios.
Patent Information
- Application Number
- CN202510724258.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing text-to-image generation technologies have difficulty in achieving precise control of human posture, and the generated images have low quality, which cannot meet the needs of high-precision application scenarios.
A combined method of artificial intelligence-based hierarchical encoder, hierarchical alignment module, cross-modal adapter and posture condition generator is adopted to generate target images that meet the posture conditions through feature extraction, semantic alignment and optimization processing.
Precise control of human body posture is achieved, and the quality of generated images is significantly improved, which can meet the needs of high-precision application scenarios.
Smart Images

Figure CN120765490A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and can be applied to the fields of financial technology and digital medical treatment, and particularly relates to an image processing method and device based on artificial intelligence, a computer device and a storage medium. BACKGROUND
[0002] In traditional text-to-image generation technology, the content of the generated image mainly depends on the semantic information of the text description, but the precise control ability of specific attributes (such as human posture) is limited. Existing methods usually adopt unsupervised or weakly supervised learning strategies, for example, by self-supervised learning to infer posture information from unannotated images, but such methods lack fine-grained modeling ability for the posture-text relationship, making it difficult to achieve precise control of the posture. Specifically, traditional methods usually regard posture information as implicit features, and generate images through global semantic matching, resulting in insufficient matching between the human posture in the generated image and the text description, and problems such as posture ambiguity and structural distortion. This rough generation method cannot meet the needs of high-precision application scenarios, and the quality of the generated image is low.
[0003] For example, in the remote insurance claim settlement scenario in the financial field, the customer may describe the situation of accidental injury (such as "fell down and supported right hand on the ground") through text, and expect to generate an image to assist in loss assessment. However, traditional text-to-image generation methods cannot accurately control the human posture (such as joint angle and body tilt direction when supporting the right hand on the ground) in the image, resulting in a generated image that does not match the actual injury posture, affecting the efficiency and accuracy of claim settlement. This technical defect not only increases the cost of manual verification, but also may cause claim disputes.
[0004] Therefore, it is urgent to provide a new text-to-image generation method to achieve fine-grained control of specific attributes such as human posture, improve the quality of the generated image, and meet the needs of related fields for precise image generation. SUMMARY
[0005] The purpose of the embodiments of the present application is to provide an image processing method and device based on artificial intelligence, a computer device and a storage medium, to solve the technical problem that existing text-to-image generation technology cannot achieve precise control of posture and the quality of the generated image is low.
[0006] In a first aspect, an image processing method based on artificial intelligence is provided, comprising:
[0007] obtaining an input text description and a reference image;
[0008] performing feature extraction on the text description and the reference image based on a preset hierarchical encoder, to obtain a text embedding corresponding to the text description and a posture embedding corresponding to the reference image;
[0009] Performing semantic alignment processing on the text embedding and the posture embedding based on a preset hierarchical alignment module to obtain corresponding alignment features;
[0010] Performing semantic optimization processing on the alignment features based on a preset cross-modal adapter to obtain corresponding target features;
[0011] Executing image generation processing corresponding to the target feature based on a preset posture condition generator to obtain a corresponding target image;
[0012] Output processing is performed on the target image.
[0013] In a second aspect, an image processing device based on artificial intelligence is provided, comprising:
[0014] An acquisition module is used to obtain input text descriptions and reference images;
[0015] an extraction module, configured to perform feature extraction on the text description and the reference image based on a preset hierarchical encoder to obtain a text embedding corresponding to the text description and a posture embedding corresponding to the reference image;
[0016] an alignment module, configured to perform semantic alignment processing on the text embedding and the posture embedding based on a preset hierarchical alignment module to obtain corresponding alignment features;
[0017] An optimization module, configured to perform semantic optimization processing on the alignment features based on a preset cross-modal adapter to obtain corresponding target features;
[0018] A generation module, configured to perform image generation processing corresponding to the target feature based on a preset posture condition generator to obtain a corresponding target image;
[0019] An output module is used to perform output processing on the target image.
[0020] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned artificial intelligence-based image processing method when executing the computer program.
[0021] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned artificial intelligence-based image processing method are implemented.
[0022] In the scheme implemented by the above-mentioned artificial intelligence-based image processing method, device, computer equipment and storage medium, first, the input text description and reference image are obtained; then, based on the preset hierarchical encoder, feature extraction is performed on the text description and the reference image to obtain a text embedding corresponding to the text description, and a posture embedding corresponding to the reference image; then, based on the preset hierarchical alignment module, semantic alignment processing is performed on the text embedding and the posture embedding to obtain corresponding alignment features; subsequently, semantic optimization processing is performed on the alignment features based on the preset cross-modal adapter to obtain corresponding target features; further, image generation processing corresponding to the target features is performed based on the preset posture condition generator to obtain the corresponding target image; finally, the target image is output. This application extracts features from the input text description and the reference image based on the use of a hierarchical encoder to obtain corresponding text embedding and posture embedding, then semantic alignment processing is performed on the text embedding and posture embedding based on the use of a hierarchical alignment module to obtain alignment features, then semantic optimization processing is performed on the alignment features based on the use of a cross-modal adapter to obtain target features, and then image generation processing corresponding to the target features is performed based on the use of a posture condition generator to obtain the target image and output it. This application processes the input text description and reference image through the combined use of a hierarchical encoder, a hierarchical alignment module, a cross-modal adapter and a posture condition generator, which can achieve precise posture control from text to image, thereby efficiently and accurately generating a target image that meets the posture conditions, thereby improving the quality of the generated target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0025] Figure 2 is a flowchart of an embodiment of an artificial intelligence-based image processing method according to the present application;
[0026] Figure 3 is a structural diagram of an embodiment of an artificial intelligence-based image processing device according to the present application;
[0027] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting upon the application; the terms "comprising," "including," and "having," and variations thereof, as used in enrolling and claims herein, are intended to be open-ended and to mean including, but not limited to; the terms "first," "second," and the like, as used in the description herein, are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. Unless otherwise indicated, the terms "plurality" and "a plurality" as used herein have the same meaning, and refer to two or more.
[0029] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase that in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of other embodiments. It is expressly understood that the embodiments described herein are merely possible examples of implementations, and are thus not a limitation on the scope of the application.
[0030] In order to make the technical personnel in the art better understand the scheme of the application, the technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings.
[0031] As shown in Figure 1 The system architecture 100 can include a terminal device 101, a network 102, and a server 103, the terminal device 101 can be a notebook computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0032] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0033] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0034] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0035] It should be noted that the artificial intelligence-based image processing method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the artificial intelligence-based image processing device is generally set in the server / terminal device.
[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0037] Continue to refer Figure 2 , shows a flow chart of an embodiment of an artificial intelligence-based image processing method according to the present application. According to different needs, the order of the steps in the flow chart can be changed, and some steps can be omitted. The artificial intelligence-based image processing method provided in the embodiment of the present application can be applied to any scenario where image generation is required, and the artificial intelligence-based image processing method can be applied to products in these scenarios, such as cultural image scenarios in the financial and medical fields. The artificial intelligence-based image processing method comprises the following steps:
[0038] Step S201: Obtain input text description and reference image.
[0039] In this embodiment, the image processing method based on artificial intelligence is executed on the electronic device (eg Figure 1The server / terminal device shown in the figure) can obtain the input text description and reference image through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future. The executive subject of the present application is an image generation system, which may be referred to as the system for short. The executive subject of the present application is specifically an image generation system, which may be referred to as the system for short. The above-mentioned text description may be text description information matched with the required generated image input according to the actual needs of the user. The above-mentioned reference image is an image provided with initial posture information, and the reference image can be used to extract the initial posture.
[0040] Among them, the system of the present application includes four main modules: a hierarchical encoder, a hierarchical alignment module, a cross-modal adapter and a posture condition generator. The system input is natural language text and a reference image (used to extract the initial posture), and the output is a generated image that conforms to the text description and has controllable posture. Specifically, the hierarchical encoder first processes the text input through a pre-trained language model, and uses an open source posture estimator to extract 2D key points from the reference image. After feature extraction, the data of these two modalities enter the hierarchical alignment module. The cross-modal adapter receives the output of the hierarchical alignment module and optimizes the feature space through comparative learning. Finally, the posture condition generator (based on a diffusion model or a GAN architecture) uses the adapted multimodal features as conditional input to gradually synthesize the target image.
[0041] In addition, the present application can be applied to business scenarios of Wenshengtu in fields such as finance and insurance and medical fields. For example, in the field of finance and insurance, the content described in the above text may include: through gesture recognition technology, analyzing the customer's hand movements (such as pen holding angle, signature trajectory stability) and body posture (such as sitting uprightness, head tilt angle) when signing an electronic contract, to evaluate whether there are abnormalities or risks in their signing behavior (such as forced signing, lack of concentration), thereby assisting in insurance claim fraud detection or contract validity verification. Explanation: Gesture recognition is used to analyze the body and hand movements of customers when signing a contract. Goal: Detect abnormal behavior (such as forced signing), assist in insurance fraud detection or contract validity verification.
[0042] In the medical field, the content described in the above text may include: using posture recognition technology to monitor the movement posture of elderly patients during home rehabilitation training in real time (such as joint flexion angle, body balance, and movement repetition accuracy), and automatically generate personalized rehabilitation suggestions (such as adjusting the range of motion, increasing training frequency) by comparing with the standard rehabilitation movement library to improve rehabilitation effects and reduce the risk of secondary injuries. Explanation: Posture recognition is used to monitor the patient's rehabilitation movements (such as joint angles and balance). Objective: Generate personalized rehabilitation suggestions by comparing with the standard movement library to improve rehabilitation effects.
[0043] Step S202 : performing feature extraction on the text description and the reference image based on a preset hierarchical encoder to obtain a text embedding corresponding to the text description and a posture embedding corresponding to the reference image.
[0044] In this embodiment, the above-mentioned specific implementation process of extracting features from the text description and the reference image based on the preset hierarchical encoder to obtain text embedding corresponding to the text description and posture embedding corresponding to the reference image will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0045] Step S203 : performing semantic alignment processing on the text embedding and the posture embedding based on a preset hierarchical alignment module to obtain corresponding alignment features.
[0046] In this embodiment, the above-mentioned specific implementation process of semantically aligning the text embedding and the posture embedding based on the preset hierarchical alignment module to obtain the corresponding alignment features will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0047] Step S204: semantically optimize the alignment features based on a preset cross-modal adapter to obtain corresponding target features.
[0048] In this embodiment, the above-mentioned specific implementation process of semantically optimizing the alignment features based on the preset cross-modal adapter to obtain the corresponding target features will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0049] Step S205 : performing image generation processing corresponding to the target feature based on a preset posture condition generator to obtain a corresponding target image.
[0050] In this embodiment, the above-mentioned specific implementation process of the preset posture condition generator performing image generation processing corresponding to the target feature to obtain the corresponding target image will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0051] Step S206: output the target image.
[0052] In this embodiment, the specific implementation process of the output processing of the target image will be further described in detail in subsequent specific embodiments of the present application, and will not be elaborated on here.
[0053] This application first obtains an input text description and a reference image; then, based on a preset hierarchical encoder, it performs feature extraction on the text description and the reference image to obtain a text embedding corresponding to the text description, and a posture embedding corresponding to the reference image; then, based on a preset hierarchical alignment module, it performs semantic alignment processing on the text embedding and the posture embedding to obtain corresponding alignment features; subsequently, based on a preset cross-modal adapter, it performs semantic optimization processing on the alignment features to obtain corresponding target features; further, based on a preset posture condition generator, it performs image generation processing corresponding to the target features to obtain a corresponding target image; finally, it outputs the target image. This application performs feature extraction on the input text description and the reference image based on the use of a hierarchical encoder to obtain corresponding text embedding and posture embedding, then, based on the use of a hierarchical alignment module, it performs semantic alignment processing on the text embedding and the posture embedding to obtain alignment features, then, based on the use of a cross-modal adapter, it performs semantic optimization processing on the alignment features to obtain target features, and then, based on the use of a posture condition generator, it performs image generation processing corresponding to the target features to obtain a target image and outputs it. This application processes the input text description and reference image through the combined use of a hierarchical encoder, a hierarchical alignment module, a cross-modal adapter and a posture condition generator, which can achieve precise posture control from text to image, thereby efficiently and accurately generating a target image that meets the posture conditions, thereby improving the quality of the generated target image.
[0054] In some optional implementations, step S203 includes the following steps:
[0055] The text embedding and the posture embedding are globally aligned based on the hierarchical alignment module to obtain corresponding global features.
[0056] In this embodiment, the hierarchical alignment module includes a global alignment layer, which aims to capture the overall semantic correspondence between text and gesture. Based on the global alignment layer, the global alignment of the text embedding and gesture embedding can be performed. The specific implementation process includes:
[0057] Input: text embedding E t , posture embedding E p .
[0058] Operation: 1. Flattening operation: E t and E p Flattened to a one-dimensional vector:
[0059] 2. Splice the flattened features: Splice the flattened text header and posture features:
[0060] E falt =[Flatten(E t ); Flatten(E p )]
[0061] 3. Capturing global semantics through Transformer encoder:
[0062] E global =Transformer([Flatten(E t ); Flatten(E p )])
[0063] Among them, E global It is a global alignment feature (referred to as global feature) used to capture the overall meaning of the action, such as "dancing", "running", etc.
[0064] Part-level alignment processing is performed on the text embedding and the posture embedding to obtain corresponding part-level features.
[0065] In this embodiment, the implementation process of the part-level alignment process includes:
[0066] Input: text embedding E t , posture embedding E p .
[0067] Operation: 1. Divide the human body into K predefined regions (such as upper body, left arm, etc.). For the kth part, construct a binary mask M k ∈{0,1}^{n×m}, used to identify related text words and joints:
[0068] 2. Establish local correspondence through attention mechanism:
[0069]
[0070] Among them, Attention is the scaled dot product attention, which is used to calculate the text word and i represents a joint, j represents a text word. k It is a part-level feature used to establish the correspondence between "waving the right hand" and the right arm joint.
[0071] Perform joint-level alignment processing on the text embedding and the posture embedding to obtain corresponding joint-level features.
[0072] In this embodiment, the implementation process of the joint-level alignment process includes:
[0073] Input: text embedding E t , posture embedding E p .
[0074] Operation: 1. Through the learnable association matrix Implementing fine-grained matching:
[0075]
[0076] Among them, σ is the sigmoid function, W a is a learnable parameter, A ij Represents the correlation between joint i and text word j.
[0077] 2. Calculate joint-word pair features:
[0078]
[0079] Here, || represents vector concatenation, which is used to fuse the fine-grained features of joints and text.
[0080] The global features, the part-level features, and the joint-level features are integrated to obtain corresponding integrated features.
[0081] In this embodiment, the integrated features are a feature set including the global features, part-level features, and joint-level features.
[0082] The integration feature is used as the alignment feature.
[0083] In this embodiment, the alignment features obtained based on the hierarchical alignment process are used as input for subsequent cross-modal optimization, so that contrastive learning can optimize the embedding space of text and gesture at different semantic levels.
[0084] The application obtains corresponding global features by globally aligning the text embedding and the pose embedding based on the hierarchical alignment module, obtains corresponding part-level features by part-level alignment of the text embedding and the pose embedding, and obtains corresponding joint-level features by joint-level alignment of the text embedding and the pose embedding. Then, the global features, the part-level features and the joint-level features are integrated to obtain corresponding integrated features. The integrated features are used as the alignment features subsequently. By using the hierarchical alignment module to globally align, part-level align and joint-level align the text embedding and the pose embedding, the application can efficiently and accurately complete semantic alignment of the text embedding and the pose embedding, gradually establish semantic correspondence between the text and the pose from global to local and from coarse granularity to fine granularity, improve the richness and accuracy of the generated alignment features, and provide multi-level feature support for subsequent cross-modal optimization and image generation.
[0085] In some optional implementations of the embodiment, step S204 includes the following steps:
[0086] A preset contrast learning strategy is obtained.
[0087] In the embodiment, the contrast learning strategy is a strategy based on a momentum contrast learning framework. The content of the strategy includes: making the cross-modal optimization through contrast learning, and forcing text and pose features with similar semantics to be close in the embedding space, thereby enhancing the semantic consistency of hierarchical alignment.
[0088] Based on the contrast learning strategy, the cross-modal adapter is used to perform semantic consistency enhancement processing on the alignment features, to obtain corresponding processing features.
[0089] In the embodiment, the optimization process corresponding to the semantic consistency enhancement processing of the alignment features based on the cross-modal adapter includes:
[0090] 1. Contrast learning framework
[0091] Objective: Optimize the embedding space of text and pose through contrast learning, and make features with similar semantics close.
[0092] Input: Sample pairs (z p i ,z t i ), where z p i and z t i are positive sample pairs (from the same instance). z p i: The pose embedding of the i-th sample. t i : The text embedding of the i-th sample.
[0093] operate:
[0094] Contrastive loss calculation: Use the InfoNCE loss function to force positive sample pairs to be close in the embedding space and negative sample pairs to be far away:
[0095]
[0096] Where τ is a temperature hyperparameter that controls the sharpness of the similarity distribution in the contrastive loss. N is the batch size, i.e., the number of sample pairs. exp is an exponential function used to calculate the weight of the similarity.
[0097] 2. Layered Sampling Strategy: Global Alignment Layer: Computes contrastive loss using complete sample pairs (all N samples). Part-level and joint-level: Sample negative samples within corresponding partitions (e.g., only within the “right arm” partition) to enforce local semantic consistency.
[0098] Cross-modal optimization enforces semantic consistency of these features in the embedding space through contrastive learning, thereby improving the accuracy and robustness of text-to-pose generation. A layered sampling strategy further enhances the consistency of local semantics (such as the correspondence between the text "right arm" and joints).
[0099] The processed feature is used as the target feature.
[0100] In this embodiment, image generation uses a latent diffusion model and a cross-attention mechanism to inject these semantically aligned features into the denoising process, generating qualified images. A gating mechanism dynamically integrates features from different levels, achieving coarse-to-fine control of conditions, further improving the quality of generated images.
[0101] The present application obtains a preset contrastive learning strategy; then based on the contrastive learning strategy, uses the cross-modal adapter to enhance the semantic consistency of the alignment features to obtain corresponding processing features; and subsequently uses the processing features as the target features. The present application uses a cross-modal adapter to enhance the semantic consistency of the alignment features based on the obtained contrastive learning strategy, forces the text features and posture features to be semantically consistent in the embedding space, and can accurately generate the adapted multi-modal target features, effectively improving the accuracy and robustness of text-to-gesture generation. In addition, the generated target features are used to provide conditional input for subsequent image generation processing, so that the generated target image can simultaneously meet the text description and posture constraints, thereby improving the quality of the generated target image.
[0102] In some optional implementations, step S205 includes the following steps:
[0103] The preset cross-attention mechanism is obtained.
[0104] In this embodiment, the posture condition generator can specifically adopt a diffusion model. The posture condition generator aims to generate an image meeting the posture condition according to the adapted multi-modal feature (target feature).
[0105] Architecture of the posture condition generator (which can be referred to as a generator for short): taking a diffusion model as an example, an image is gradually generated through a denoising process.
[0106] Condition injection: in the l-th layer network, hierarchical features are injected into the denoising process through cross-attention:
[0107]
[0108] wherein Q l is the query feature of the current layer (the feature from the denoising process of the generator), K c and V c are the key and value from the output of the hierarchical alignment module (i.e., the features output after the hierarchical alignment processing), d is the feature dimension, and is used to scale the dot product attention.
[0109] Gating mechanism: used for dynamically fusing different hierarchical features (global, part-level, and joint-level), and the formula is as follows:
[0110]
[0111] wherein H l is the hidden state of the current layer, g k is the gating weight predicted by the current noise level, which controls the fusion proportion of different hierarchical features. Attention k (H l ) is the cross-attention output of the k-th layer feature, and H l ' is the updated hidden state.
[0112] Output: generate an image meeting the text description and posture condition.
[0113] Based on the cross-attention mechanism, the target feature is injected as condition information into the denoising process in the posture condition generator.
[0114] In this embodiment, the process of injecting the target feature as condition information into the denoising process in the posture condition generator can be completed based on the above-mentioned condition injection processing process.
[0115] Based on the posture condition generator injected with condition information, image generation processing is performed according to the target features and a corresponding generated image is obtained.
[0116] In this embodiment, after the conditional information injection is completed, the posture condition generator is used to gradually generate an image through a denoising process, and a generated image that meets the posture conditions is generated based on the adapted multimodal features (target features) and is used as the final target image.
[0117] The generated image is used as the target image.
[0118] In this embodiment, image generation uses a latent diffusion model and a cross-attention mechanism to inject these semantically aligned features into the denoising process, generating images that meet pose requirements. A gating mechanism dynamically integrates features from different levels, achieving coarse-to-fine conditional control and further improving the quality of generated images.
[0119] This application obtains a preset cross-attention mechanism; then, based on the cross-attention mechanism, injects the target features as conditional information into the denoising process in the posture condition generator; then, based on the posture condition generator with the conditional information injected, performs image generation processing according to the target features and obtains a corresponding generated image; and subsequently uses the generated image as the target image. This application uses a cross-attention mechanism to inject the target features as conditional information into the denoising process in the posture condition generator, and then, based on the use of the posture condition generator, performs image generation processing according to the target features, thereby generating a target image that meets the posture conditions, effectively improving the quality of the generated target image.
[0120] In some optional implementations, step S206 includes the following steps:
[0121] The target image is subjected to quality optimization processing based on a preset quality optimization strategy to obtain a corresponding optimized image.
[0122] In this embodiment, the above-mentioned specific implementation process of performing quality optimization processing on the target image based on the preset quality optimization strategy to obtain the corresponding optimized image will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.
[0123] The optimized image is semantically verified based on a preset semantic verification strategy.
[0124] In this embodiment, the process of performing semantic verification on the optimized image using the above-mentioned semantic verification strategy includes: using a pre-trained semantic segmentation model (such as DeepLabv3+) to segment the optimized image, checking whether it matches the posture conditions of the input text description and the reference image, so as to complete the semantic verification of the optimized image and generate a corresponding semantic verification result. If it is detected that the optimized image matches the posture conditions of the input text description and the reference image, the optimized image is determined to have passed the semantic verification; otherwise, the optimized image is determined to have failed the semantic verification.
[0125] If the optimized image passes the semantic verification, a target output mode corresponding to the optimized image is obtained.
[0126] In this embodiment, the user's user information can be obtained and the user's preferred target data push method can be found through this user information, and this target data push method can be used as the target output method. The user can be the user who entered the text description. The user can be tagged with a data push method that matches their preferred method based on their registration options, historical behavior data, or a questionnaire. The data push methods can include at least pop-up notifications, interface displays, emails, and internal messages.
[0127] The optimized image is output-processed based on the target output mode.
[0128] In this embodiment, the generated optimized image may be sent to the user according to the determined target output mode to complete the output processing of the optimized image.
[0129] The present application performs quality optimization processing on the target image based on a preset quality optimization strategy to obtain a corresponding optimized image; then performs semantic verification on the optimized image based on a preset semantic verification strategy; if the optimized image passes the semantic verification, the target output mode corresponding to the optimized image is obtained; and the optimized image is subsequently output-processed based on the target output mode. The present application performs quality optimization processing on the target image based on the use of a quality optimization strategy to obtain an optimized image; then performs semantic verification on the optimized image based on the use of a semantic verification strategy, and after detecting that the optimized image passes the semantic verification, obtains the target output mode corresponding to the optimized image, and then performs output processing on the optimized image based on the use of the target output mode, thereby effectively ensuring the quality and accuracy of the sent optimized image, improving the output intelligence of the optimized image, and being conducive to improving the user experience.
[0130] In some optional implementations of this embodiment, performing quality optimization processing on the target image based on a preset quality optimization strategy to obtain a corresponding optimized image includes the following steps:
[0131] Noise suppression is performed on the target image to obtain a corresponding first processed image.
[0132] In this embodiment, the noise suppression process aims to suppress noise and artifacts in the generated image, making it smoother and more natural. This process includes using a non-local means (NLM) algorithm to suppress noise and performing adaptive thresholding on high-frequency components (such as detail coefficients after wavelet transform).
[0133] Performing style adjustment processing on the first processed image to obtain a corresponding second processed image.
[0134] In this embodiment, the goal of the above-mentioned style adjustment process is to transfer the style of an image to a generated image to change its visual effect. The implementation process of the style adjustment process includes: selecting a style image that contains the desired visual style; using the generated target image as the initial image; using a pre-trained style transfer model (such as Neural Style Transfer, AdaIN, etc.) to extract features of the initial image and the style image respectively; integrating the style features into the initial image through an optimization process, maintaining the content structure while applying the style features; and outputting a style-transferred image with the structure of the initial image and the visual effect of the style image.
[0135] Perform detail enhancement processing on the second processed image to obtain a corresponding third processed image.
[0136] In this embodiment, the goal of the detail enhancement process is to enhance the details (e.g., texture and edge clarity) of the generated image, making it more realistic and refined. The detail enhancement process includes: using a lightweight super-resolution network to enhance the details of the generated image; and using a Laplacian operator or an unsharp mask to enhance image edges.
[0137] The third processed image is used as the optimized image.
[0138] This application performs noise suppression on the target image to obtain a corresponding first processed image; then performs style adjustment on the first processed image to obtain a corresponding second processed image; then performs detail enhancement on the second processed image to obtain a corresponding third processed image; and subsequently uses the third processed image as the optimized image. This application automatically and accurately optimizes the quality of the target image by performing noise suppression, style adjustment, and detail enhancement on the target image, resulting in a significantly improved optimized image with respect to noise, style, and detail, thereby effectively improving the visual quality, smoothness, and naturalness of the generated optimized image.
[0139] In some optional implementations of this embodiment, the hierarchical encoder includes a text encoder and a pose estimator; step S202 includes the following steps:
[0140] The text description is encoded based on the text encoder to obtain word embedding, and the word embedding is used as the text embedding.
[0141] In this embodiment, the hierarchical encoder uses a dual-stream architecture (text encoder and pose estimator) to process text input and pose input. The text encoder can specifically use a pre-trained language model (such as BERT). Specifically, for a text sequence (text description) of length n T = {t1, t2, ..., t n}, first obtain word embeddings through the pre-trained language model:
[0142]
[0143] Among them, d is the embedding dimension, which represents the semantic features of each word.
[0144] Performing pose estimation on the reference image based on the pose estimator to obtain corresponding coordinate data.
[0145] In this embodiment, the above-mentioned posture estimator can specifically adopt an open source posture estimator. A reference image can be input into the posture estimator, so that the posture estimator can estimate the posture of the reference image and output the joint coordinates matching the reference image. Represents the two-dimensional coordinates of m joints, where m is the number of joints. The obtained joint coordinates can be used as the above coordinate data.
[0146] Performing a linear projection on the coordinate data to obtain a posture embedding corresponding to the reference image.
[0147] In this embodiment, the coordinate data, i.e., the joint coordinates, can be linearly projected to the same dimension as the text embedding to obtain the corresponding posture embedding. The specific formula is as follows:
[0148]
[0149] Among them, W p and b p are learnable parameters used to align the embedding spaces of pose and text.
[0150] The present application obtains word embeddings by encoding the text description based on the text encoder, and uses the word embeddings as the text embeddings; then performs pose estimation on the reference image based on the pose estimator to obtain corresponding coordinate data; and subsequently linearly projects the coordinate data to obtain a pose embedding corresponding to the reference image. The present application obtains word embeddings by encoding the text description based on the use of a text encoder and uses the text embeddings, and performs pose estimation on the reference image based on the use of a pose estimator to obtain coordinate data, and linearly projects the coordinate data to obtain a pose embedding corresponding to the reference image, thereby achieving efficient and accurate feature extraction processing for the text description and the reference image based on the use of a hierarchical encoder, improving the processing efficiency of feature extraction and ensuring the accuracy of the extracted text embeddings and pose embeddings.
[0151] In some optional implementations, the user information obtained is obtained with the user's consent and complies with relevant laws and policies.
[0152] In addition, any software tools or components not provided by our company that appear in the embodiments of this application are merely examples and do not represent actual use.
[0153] In addition, the technological breakthroughs of this application are mainly reflected in three dimensions: first, the designed hierarchical encoder realizes posture feature extraction from macro to micro, which can simultaneously capture the overall action intention and local joint details; second, the innovative cross-modal adapter optimizes the embedding space of text and posture features through contrastive learning, enabling the model to understand the relationship between action descriptions such as "raising both hands" and specific joint positions; finally, the proposed dynamic injection mechanism ensures the synergy between posture information and text semantics during the generation process, which not only maintains image quality but also achieves precise control.
[0154] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0155] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned target image, the above-mentioned target image can also be stored in a node of a blockchain.
[0156] The blockchain referred to in this application is a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.
[0157] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0158] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0159] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0160] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0161] Further references Figure 3, as a response to the above Figure 2 The present application provides an embodiment of an image processing device based on artificial intelligence. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0162] like Figure 3 As shown, the artificial intelligence-based image processing device 300 of this embodiment includes: an acquisition module 301, an extraction module 302, an alignment module 303, an optimization module 304, a generation module 305, and an output module 306.
[0163] An acquisition module 301 is used to acquire an input text description and a reference image;
[0164] An extraction module 302 is configured to perform feature extraction on the text description and the reference image based on a preset hierarchical encoder to obtain a text embedding corresponding to the text description and a posture embedding corresponding to the reference image;
[0165] An alignment module 303 is configured to perform semantic alignment processing on the text embedding and the posture embedding based on a preset hierarchical alignment module to obtain corresponding alignment features;
[0166] An optimization module 304 is configured to perform semantic optimization processing on the alignment features based on a preset cross-modal adapter to obtain corresponding target features;
[0167] A generating module 305 is configured to perform image generation processing corresponding to the target feature based on a preset posture condition generator to obtain a corresponding target image;
[0168] The output module 306 is configured to perform output processing on the target image.
[0169] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based image processing method in the aforementioned embodiment, and are not repeated here.
[0170] In some optional implementations of this embodiment, the alignment module 303 includes:
[0171] A first processing submodule is configured to perform global alignment processing on the text embedding and the posture embedding based on the hierarchical alignment module to obtain corresponding global features;
[0172] A second processing submodule is configured to perform part-level alignment processing on the text embedding and the posture embedding to obtain corresponding part-level features;
[0173] A third processing submodule is configured to perform joint-level alignment processing on the text embedding and the posture embedding to obtain corresponding joint-level features;
[0174] An integration submodule, configured to integrate the global features, the part-level features, and the joint-level features to obtain corresponding integrated features;
[0175] The first determining submodule is configured to use the integration feature as the alignment feature.
[0176] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based image processing method in the aforementioned embodiment, and are not repeated here.
[0177] In some optional implementations of this embodiment, the optimization module 304 includes:
[0178] A first acquisition submodule is used to acquire a preset contrastive learning strategy;
[0179] an enhancement submodule, configured to enhance the semantic consistency of the alignment features using the cross-modal adapter based on the contrastive learning strategy to obtain corresponding processed features;
[0180] The second determining submodule is configured to use the processed feature as the target feature.
[0181] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based image processing method in the aforementioned embodiment, and are not repeated here.
[0182] In some optional implementations of this embodiment, the generating module 305 includes:
[0183] The second acquisition submodule is used to obtain the preset cross attention mechanism;
[0184] an injection submodule, configured to inject the target features as conditional information into the denoising process in the pose condition generator based on the cross-attention mechanism;
[0185] An execution submodule, configured to perform image generation processing according to the target features and obtain a corresponding generated image based on the posture condition generator injected with the condition information;
[0186] The third determining submodule is configured to use the generated image as the target image.
[0187] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based image processing method in the aforementioned embodiment, and are not repeated here.
[0188] In some optional implementations of this embodiment, the output module 306 includes:
[0189] An optimization submodule, configured to perform quality optimization processing on the target image based on a preset quality optimization strategy to obtain a corresponding optimized image;
[0190] A verification submodule, configured to perform semantic verification on the optimized image based on a preset semantic verification strategy;
[0191] A third acquisition submodule is configured to acquire a target output mode corresponding to the optimized image if the optimized image passes the semantic verification;
[0192] An output submodule is used to perform output processing on the optimized image based on the target output mode.
[0193] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based image processing method in the aforementioned embodiment, and are not repeated here.
[0194] In some optional implementations of this embodiment, the optimization submodule includes:
[0195] a first processing unit, configured to perform noise suppression processing on the target image to obtain a corresponding first processed image;
[0196] a second processing unit, configured to perform style adjustment processing on the first processed image to obtain a corresponding second processed image;
[0197] a third processing unit, configured to perform detail enhancement processing on the second processed image to obtain a corresponding third processed image;
[0198] A determining unit is configured to use the third processed image as the optimized image.
[0199] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based image processing method in the aforementioned embodiment, and are not repeated here.
[0200] In some optional implementations of this embodiment, the hierarchical encoder includes a text encoder and a pose estimator; the extraction module 302 includes:
[0201] an encoding submodule, configured to encode the text description based on the text encoder to obtain a word embedding, and use the word embedding as the text embedding;
[0202] a fourth processing submodule, configured to perform pose estimation on the reference image based on the pose estimator to obtain corresponding coordinate data;
[0203] The projection submodule is used to perform linear projection on the coordinate data to obtain a posture embedding corresponding to the reference image.
[0204] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based image processing method in the aforementioned embodiment, and are not repeated here.
[0205] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0206] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0207] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0208] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for an artificial intelligence-based image processing method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.
[0209] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for executing the artificial intelligence-based image processing method.
[0210] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0211] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0212] In an embodiment of the present application, first, an input text description and a reference image are obtained; then, based on a preset hierarchical encoder, feature extraction is performed on the text description and the reference image to obtain a text embedding corresponding to the text description, and a posture embedding corresponding to the reference image; then, based on a preset hierarchical alignment module, semantic alignment processing is performed on the text embedding and the posture embedding to obtain corresponding alignment features; subsequently, semantic optimization processing is performed on the alignment features based on a preset cross-modal adapter to obtain corresponding target features; further, image generation processing corresponding to the target features is performed based on a preset posture condition generator to obtain a corresponding target image; finally, output processing is performed on the target image. The present application extracts features from the input text description and the reference image based on the use of a hierarchical encoder to obtain corresponding text embedding and posture embedding, then, based on the use of a hierarchical alignment module, semantic alignment processing is performed on the text embedding and the posture embedding to obtain alignment features, then, based on the use of a cross-modal adapter, semantic optimization processing is performed on the alignment features to obtain target features, and then, based on the use of a posture condition generator, image generation processing corresponding to the target features is performed to obtain a target image and output it. This application processes the input text description and reference image through the combined use of a hierarchical encoder, a hierarchical alignment module, a cross-modal adapter and a posture condition generator, which can achieve precise posture control from text to image, thereby efficiently and accurately generating a target image that meets the posture conditions, thereby improving the quality of the generated target image.
[0213] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned artificial intelligence-based image processing method.
[0214] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0215] In an embodiment of the present application, first, an input text description and a reference image are obtained; then, based on a preset hierarchical encoder, feature extraction is performed on the text description and the reference image to obtain a text embedding corresponding to the text description, and a posture embedding corresponding to the reference image; then, based on a preset hierarchical alignment module, semantic alignment processing is performed on the text embedding and the posture embedding to obtain corresponding alignment features; subsequently, semantic optimization processing is performed on the alignment features based on a preset cross-modal adapter to obtain corresponding target features; further, image generation processing corresponding to the target features is performed based on a preset posture condition generator to obtain a corresponding target image; finally, output processing is performed on the target image. The present application extracts features from the input text description and the reference image based on the use of a hierarchical encoder to obtain corresponding text embedding and posture embedding, then, based on the use of a hierarchical alignment module, semantic alignment processing is performed on the text embedding and the posture embedding to obtain alignment features, then, based on the use of a cross-modal adapter, semantic optimization processing is performed on the alignment features to obtain target features, and then, based on the use of a posture condition generator, image generation processing corresponding to the target features is performed to obtain a target image and output it. This application processes the input text description and reference image through the combined use of a hierarchical encoder, a hierarchical alignment module, a cross-modal adapter and a posture condition generator, which can achieve precise posture control from text to image, thereby efficiently and accurately generating a target image that meets the posture conditions, thereby improving the quality of the generated target image.
[0216] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0217] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. An image processing method based on artificial intelligence, characterized in that: The steps include: Get the input text description and reference image; Performing feature extraction on the text description and the reference image based on a preset hierarchical encoder to obtain a text embedding corresponding to the text description and a posture embedding corresponding to the reference image; Performing semantic alignment processing on the text embedding and the posture embedding based on a preset hierarchical alignment module to obtain corresponding alignment features; Performing semantic optimization processing on the alignment features based on a preset cross-modal adapter to obtain corresponding target features; Executing image generation processing corresponding to the target feature based on a preset posture condition generator to obtain a corresponding target image; Output processing is performed on the target image.
2. The image processing method based on artificial intelligence according to claim 1, characterized in that: The step of performing semantic alignment processing on the text embedding and the posture embedding based on a preset hierarchical alignment module to obtain corresponding alignment features specifically includes: Performing global alignment processing on the text embedding and the posture embedding based on the hierarchical alignment module to obtain corresponding global features; Performing part-level alignment processing on the text embedding and the posture embedding to obtain corresponding part-level features; Performing joint-level alignment processing on the text embedding and the posture embedding to obtain corresponding joint-level features; Integrating the global features, the part-level features, and the joint-level features to obtain corresponding integrated features; The integration feature is used as the alignment feature.
3. The image processing method based on artificial intelligence according to claim 1, characterized in that: The step of performing semantic optimization processing on the alignment features based on a preset cross-modal adapter to obtain corresponding target features specifically includes: Get the preset contrastive learning strategy; Based on the contrastive learning strategy, using the cross-modal adapter to perform semantic consistency enhancement processing on the aligned features to obtain corresponding processed features; The processed feature is used as the target feature.
4. The image processing method based on artificial intelligence according to claim 1, characterized in that: The step of performing image generation processing corresponding to the target feature based on the preset posture condition generator to obtain the corresponding target image specifically includes: Get the preset cross-attention mechanism; Based on the cross attention mechanism, the target features are injected as conditional information into the denoising process in the pose condition generator; Based on the posture condition generator injected with condition information, perform image generation processing according to the target features and obtain a corresponding generated image; The generated image is used as the target image.
5. The image processing method based on artificial intelligence according to claim 1, characterized in that: The step of outputting the target image specifically includes: Performing quality optimization processing on the target image based on a preset quality optimization strategy to obtain a corresponding optimized image; Performing semantic verification on the optimized image based on a preset semantic verification strategy; If the optimized image passes the semantic verification, obtaining a target output mode corresponding to the optimized image; The optimized image is output-processed based on the target output mode.
6. The image processing method based on artificial intelligence according to claim 5, characterized in that: The step of performing quality optimization processing on the target image based on a preset quality optimization strategy to obtain a corresponding optimized image specifically includes: Performing noise suppression processing on the target image to obtain a corresponding first processed image; performing style adjustment processing on the first processed image to obtain a corresponding second processed image; performing detail enhancement processing on the second processed image to obtain a corresponding third processed image; The third processed image is used as the optimized image.
7. The image processing method based on artificial intelligence according to claim 1, characterized in that: The hierarchical encoder includes a text encoder and a pose estimator; the step of extracting features from the text description and the reference image based on the preset hierarchical encoder to obtain a text embedding corresponding to the text description and a pose embedding corresponding to the reference image specifically includes: Encoding the text description based on the text encoder to obtain word embedding, and using the word embedding as the text embedding; performing pose estimation on the reference image based on the pose estimator to obtain corresponding coordinate data; Performing a linear projection on the coordinate data to obtain a posture embedding corresponding to the reference image.
8. An image processing device based on artificial intelligence, characterized in that: include: An acquisition module is used to obtain input text descriptions and reference images; an extraction module, configured to perform feature extraction on the text description and the reference image based on a preset hierarchical encoder to obtain a text embedding corresponding to the text description and a posture embedding corresponding to the reference image; an alignment module, configured to perform semantic alignment processing on the text embedding and the posture embedding based on a preset hierarchical alignment module to obtain corresponding alignment features; An optimization module, configured to perform semantic optimization processing on the alignment features based on a preset cross-modal adapter to obtain corresponding target features; A generation module, configured to perform image generation processing corresponding to the target feature based on a preset posture condition generator to obtain a corresponding target image; An output module is used to perform output processing on the target image.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the artificial intelligence-based image processing method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based image processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data alignment processing method and device for multi-modal fusion
CN118761027A
Cited By
Controllable image generation method and device, equipment and storage medium
CN121414875A