Image processing method and device based on artificial intelligence, computer equipment and medium
By constructing and iteratively optimizing the perceptual map through a self-attention mechanism, the visual illusion problem of visual language models is solved, generating accurate text response data and improving the application effect in the fields of finance, insurance and healthcare.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-01
AI Technical Summary
Existing visual language models suffer from visual illusion problems, resulting in low accuracy in generating text responses and impacting practical applications in fields such as finance, insurance, and healthcare.
An initial perceptual map is constructed based on a self-attention mechanism, iteratively refined, and optimized. Nonlinear resampling is then performed to generate the target image. Finally, a visual language model is used for inference to generate text response data.
It effectively avoids visual illusion problems, improves the accuracy of generated text responses, and ensures reliable applications in the financial, insurance, and medical fields.
Smart Images

Figure CN121961933A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as fintech and digital healthcare, particularly to image processing methods, devices, computer equipment, and storage media based on artificial intelligence. Background Technology
[0002] With the rapid development of visual language models, these models have gained the ability to generate coherent text descriptions based on image input, demonstrating enormous application potential in many fields. However, these models generally suffer from the "visual hallucination" problem, where the generated responses may contain objects that are not present in the image or provide incorrect attribute descriptions. This significantly reduces the accuracy of the generated text responses and severely restricts the effective application of these models in numerous real-world scenarios.
[0003] In the financial insurance sector, taking insurance claims as an example, when a customer submits a claim application containing images of the accident scene, the visual-language model should accurately describe the accident situation based on the images, such as the damaged parts of the vehicle and the extent of the damage. However, due to "visual illusions," the model may generate text containing incorrect information, such as describing a minor scratch as a serious collision, or fabricating damaged parts that do not exist in the image. This not only leads to claims reviewers misjudging the true situation of the accident, affecting the efficiency and fairness of the claims process, but may also cause disputes between customers and insurance companies, damaging the reputation of insurance companies.
[0004] In the medical field, taking medical imaging diagnostic assistance as an example, doctors use visual-language models to analyze patients' medical images (such as X-rays and CT images). The model should accurately describe the location, size, and shape of lesions in the images. However, the "visual illusion" problem can cause the model to generate descriptions that do not match reality, such as misclassifying benign lesions as malignant or incorrectly reporting the number and location of lesions. This can seriously interfere with doctors' diagnoses, leading to misdiagnosis or missed diagnosis, delaying treatment, and causing irreparable damage to the patient's health.
[0005] Therefore, in order to solve the "visual illusion" problem of visual language models, improve the accuracy of generated text responses, and enable them to be applied more reliably in key areas such as finance, insurance, and healthcare, there is an urgent need for an effective technical solution to optimize and correct the text generated by the model. Summary of the Invention
[0006] The purpose of this application is to propose an image processing method, apparatus, computer device, and storage medium based on artificial intelligence, in order to solve the technical problem that existing visual language models suffer from visual illusions, resulting in low accuracy of generated text responses.
[0007] Firstly, an image processing method based on artificial intelligence is provided, including: Receive the raw input image; During the decoding process of the original image using a preset visual language model, a corresponding initial perceptual map is constructed based on a preset self-attention mechanism. Based on the initial perception map, iterative refinement is performed to obtain the corresponding first perception map; The first perceptual map is optimized based on a preset post-processing strategy to obtain the corresponding second perceptual map. Based on the second perceptual map, the original image is subjected to nonlinear resampling processing to generate the corresponding target image; Based on the visual language model, the target image is subjected to inference processing to generate corresponding text response data; The text response data is then processed for output.
[0008] Secondly, an image processing device based on artificial intelligence is provided, comprising: The receiving module is used to receive the input raw image; The construction module is used to construct the corresponding initial perceptual map based on a preset self-attention mechanism during the process of decoding the original image using a preset visual language model. The refinement module is used to perform iterative refinement processing based on the initial perception map to obtain the corresponding first perception map; An optimization module is used to optimize the first perceptual map based on a preset post-processing strategy to obtain a corresponding second perceptual map. The sampling module is used to perform non-linear resampling processing on the original image based on the second perceptual map to generate a corresponding target image; The generation module is used to perform inference processing on the target image based on the visual language model to generate corresponding text response data. The output module is used to process the text response data.
[0009] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based image processing method.
[0010] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described artificial intelligence-based image processing method.
[0011] In the above-described scheme implemented by the image processing method, apparatus, computer equipment, and storage medium based on artificial intelligence, the input raw image is first received; then, during the decoding process of the raw image using a preset visual language model, a corresponding initial perceptual map is constructed based on a preset self-attention mechanism; subsequently, iterative refinement processing is performed based on the initial perceptual map to obtain a corresponding first perceptual map; and the first perceptual map is optimized based on a preset post-processing strategy to obtain a corresponding second perceptual map; subsequently, the raw image is subjected to nonlinear resampling processing based on the second perceptual map to generate a corresponding target image; further, inference processing is performed on the target image based on the visual language model to generate corresponding text response data; finally, the text response data is output. Based on the above automated processing flow, in the process of decoding the input original image using the visual language model in this application, an initial perceptual map is constructed by using a self-attention mechanism. Then, the initial perceptual map is iteratively refined and optimized to obtain a second perceptual map. Based on the obtained second perceptual map, the original image is nonlinearly resampled to generate a target image. Then, the target image is inferred based on the visual language model, thereby realizing decoding using clearer visual details, effectively avoiding visual illusion problems, and improving the accuracy of the generated text response data. Attached Figure Description
[0012] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the artificial intelligence-based image processing method according to this application; Figure 3 This is a schematic diagram of a structure of an embodiment of an artificial intelligence-based image processing apparatus according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0015] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0017] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0018] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0019] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0020] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0021] It should be noted that the AI-based image processing method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the AI-based image processing device is generally located in the server / terminal device.
[0022] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0023] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the AI-based image processing method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The AI-based image processing method provided in this application can be applied to any scenario requiring product recommendation, and thus can be applied to products in these scenarios, such as product recommendations in the financial insurance field. The AI-based image processing method includes the following steps: Step S201: Receive the input raw image.
[0024] In this embodiment, the artificial intelligence-based image processing method runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire the input raw image via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods. The executing entity of this application is specifically an image processing system, which can be simply referred to as the system. The aforementioned raw image can be image data input by the user according to their actual personal processing needs, to be used to generate a response description. This application can be applied to image processing scenarios in the fields of fintech and digital healthcare. For example, in the financial insurance field, the input raw image can be an insurance product promotional poster. The corresponding image content includes: a brightly colored insurance product promotional poster with a large insurance product name in the center, such as "Health Guardian Lifetime Critical Illness Insurance," in eye-catching font with a 3D effect. Surrounding it are elements related to health and protection, such as green leaves and cartoon characters symbolizing health. The poster details the coverage of the insurance product, such as the types of critical illnesses covered, the payout ratio, and some purchase conditions and contact information. Alternatively, the input image can be an introductory image for a bank wealth management product. The image content includes: an introductory image of the bank's wealth management product, with the bank's logo and name at the top. The middle section uses a combination of charts and text to display the product's returns, such as expected yield curves for different investment periods, and comparisons with similar products in the market. Below is a description of the product's investment direction, such as primarily investing in the bond market and money market, along with risk warnings and product features.
[0025] Alternatively, in the field of digital healthcare, the input raw image can be a medical imaging diagnostic report. The corresponding image content includes: an image containing medical images (such as X-rays, CT scans, etc.) and a diagnostic report. The medical imaging section shows a part of the body, such as a CT scan image of the lungs, showing the lung structure and possible lesion areas. The diagnostic report section provides a detailed description of the features observed in the image, such as the size, shape, and location of the lesions, as well as preliminary diagnostic conclusions and recommendations, such as whether further examination or treatment is needed. Alternatively, the input raw image can also be a screenshot of an electronic medical record page. The corresponding image content includes: a screenshot of an electronic medical record page containing the patient's basic information, such as name, age, and gender; the patient's medical history, including past illnesses, surgical history, allergies, etc.; and a description of the current condition, such as symptoms, signs, and examination results. The page may also include some doctor's diagnostic opinions and treatment plans.
[0026] In step S202, during the process of decoding the original image using a preset visual language model, a corresponding initial perceptual map is constructed based on a preset self-attention mechanism.
[0027] In this embodiment, the selection of the aforementioned visual language model is not specifically limited and can be determined according to actual business needs. For example, models such as LLaVA and Qwen-VL can be used. The specific implementation process of constructing the corresponding initial perceptual map based on the preset self-attention mechanism will be further described in detail in subsequent embodiments of this application, and will not be elaborated upon here.
[0028] Step S203: Based on the initial perception map, perform iterative refinement processing to obtain the corresponding first perception map.
[0029] In this embodiment, the specific implementation process of iteratively refining the initial perception map to obtain the corresponding first perception map will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0030] Step S204: Optimize the first perceptual map based on a preset post-processing strategy to obtain the corresponding second perceptual map.
[0031] In this embodiment, the specific implementation process of optimizing the first perceptual map based on the preset post-processing strategy to obtain the corresponding second perceptual map will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0032] Step S205: Perform nonlinear resampling processing on the original image based on the second perceptual map to generate the corresponding target image.
[0033] In this embodiment, the specific implementation process of performing nonlinear resampling processing on the original image based on the second perceptual map to generate the corresponding target image will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0034] Step S206: Based on the visual language model, perform inference processing on the target image to generate corresponding text response data.
[0035] In this embodiment, the specific implementation process of reasoning on the target image based on the visual language model to generate corresponding text response data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0036] Step S207: Output the text response data.
[0037] In this embodiment, the generated text response data can be sent to the relevant user via email, message, or interface display to complete the output processing of the text response data.
[0038] This application first receives an input raw image; then, during the decoding process of the raw image using a preset visual language model, a corresponding initial perceptual map is constructed based on a preset self-attention mechanism; subsequently, iterative refinement processing is performed based on the initial perceptual map to obtain a corresponding first perceptual map; and the first perceptual map is optimized based on a preset post-processing strategy to obtain a corresponding second perceptual map; subsequently, the raw image is subjected to nonlinear resampling processing based on the second perceptual map to generate a corresponding target image; further, inference processing is performed on the target image based on the visual language model to generate corresponding text response data; finally, the text response data is output. Based on the above automated processing flow, in the process of decoding the input original image using the visual language model in this application, an initial perceptual map is constructed by using a self-attention mechanism. Then, the initial perceptual map is iteratively refined and optimized to obtain a second perceptual map. Based on the obtained second perceptual map, the original image is nonlinearly resampled to generate a target image. Then, the target image is inferred based on the visual language model, thereby realizing decoding using clearer visual details, effectively avoiding visual illusion problems, and improving the accuracy of the generated text response data.
[0039] In some alternative implementations, step S202 includes the following steps: A hierarchical analysis is performed on the visual language model to determine a specified intermediate layer from the visual language model.
[0040] In this embodiment, when using a Visual Language Model (VLM, or simply the model) to generate corresponding text descriptions for the input original image, tokens (which can be understood as basic language units such as words or phrases) are generated one by one in a certain order. In the decoding step of generating each token, the VLM's Transformer layer performs self-attention calculations, generating a self-attention matrix. This self-attention matrix reflects the different levels of attention the model pays to different parts of the input image and the already generated text when generating the current token. The self-attention matrix can be extracted from these Transformer layers as the basic data for subsequently constructing the perceptual map. An example of a Chinese description is as follows: Suppose the input is an original image of a poster displaying a financial insurance product, and the model begins generating text descriptions. In the decoding step of generating the first token (e.g., "Welcome"), the self-attention matrix is extracted from the VLM's Transformer layer. This matrix reflects the model's attention to different areas of the poster image (such as the poster title, patterns, etc.) and possible starting markers when generating the word "Welcome".
[0041] In the context of Visual Language Modeling (VLM), a token can be understood as the smallest unit processed by the model. In text, this can be a word, root word, affix, etc.; in images, after feature extraction and other processing, images are also divided into units, which, like tokens in text, represent local feature information of the image during model processing. For example, in a face image, after processing, the features of the eyes, nose, and mouth may each correspond to different tokens.
[0042] Furthermore, extensive experimental verification shows that in the Transformer structure of VLM, the attention weights of the intermediate layers more accurately capture visual object information than the last layer. This is because the intermediate layers are at the middle stage of information processing, receiving basic feature information from the shallow layers and providing a foundation for deeper, higher-level feature processing. Assuming the Transformer structure has 12 layers, layers 4-8 can be selected as the aforementioned designated intermediate layers. This application further selects attention weights from several intermediate layers.
[0043] Obtain the attention weights corresponding to the specified intermediate layer.
[0044] In this embodiment, after determining the designated intermediate layer, attention weights from the designated intermediate layer are selected. For example, assuming the Transformer structure has 12 layers, layers 4-8 can be selected as the designated intermediate layer, and the attention weights corresponding to the designated intermediate layer can be obtained. By obtaining the attention weights, it is possible to understand the model's attention to different parts of the image and text when generating each token, providing a basis for subsequently determining key regions in the image.
[0045] Get the preset pooling aggregation strategy.
[0046] In this embodiment, the pooling aggregation strategy includes the following: In the selected intermediate layers, each layer has multi-head attention, with each head focusing on elements in the sequence from different perspectives. Max pooling is performed along the head dimension, meaning that for each token, the maximum attention score from all heads is taken, thus obtaining the maximum attention level for each token in each intermediate layer. Then, these maximum values are summed along the layers to obtain the basic token-level heatmap, i.e., the basic heatmap. For example, assuming there are 5 intermediate layers, and each token has a maximum attention score in each intermediate layer, adding these 5 scores gives the value of that token in the basic heatmap.
[0047] The attention weights are aggregated based on the pooling aggregation strategy to obtain the corresponding basic heatmap.
[0048] In this embodiment, the aggregation processing of the attention weights can be performed based on the strategy content of the pooling aggregation strategy, and the resulting basic heatmap can be used as the corresponding initial perception map.
[0049] The basic heatmap is used as the initial sensing map.
[0050] This application performs hierarchical analysis on a visual language model to identify a specified intermediate layer. Then, it obtains the attention weights corresponding to the specified intermediate layer. Next, it acquires a preset pooling aggregation strategy and aggregates the attention weights based on this strategy to obtain a corresponding base heatmap. This base heatmap is then used as the initial perceptual map. Based on this processing flow, this application performs hierarchical analysis on a visual language model to identify a specified intermediate layer and obtain its attention weights. Then, it aggregates these attention weights using a pooling aggregation strategy and uses the resulting base heatmap as the corresponding initial perceptual map. This achieves efficient and accurate integration of attention information from multiple heads, highlighting the most significant attention given to each token in each intermediate layer. Through hierarchical summation, it further integrates information from different layers to obtain an initial perceptual map reflecting the importance of each token in the image.
[0051] In some optional implementations of this embodiment, step S203 includes the following steps: Obtain a pre-defined iterative attention refinement strategy.
[0052] In this embodiment, the strategy content of the above-mentioned iterative attention refinement strategy includes: 1. Initialization. Set the maximum number of iterations and the attention threshold: The maximum number of iterations limits the maximum number of executions in the entire iteration process to prevent infinite loops; the attention threshold is used to determine when to stop iterating. When the sum of the current heatmaps is less than the threshold, it means that the model's attention to the image is already very low, and there is no need to continue iterating. Initialize the mask to all 1s: The mask acts as a "switch" here. All 1s mean that all regions of the image are initially open, and the model can freely focus on different parts of the image.
[0053] 2. Iterative Loop. Using the current mask for model forward propagation: At the beginning of each iteration, the current mask (representing the current iteration number) is applied to the visual language model. During forward propagation (i.e., the process of generating text descriptions), the model adjusts its attention to different regions of the original image based on the mask. Regions with a mask value of 1 are normally attended to by the model; regions with a mask value of 0 are "guided" to receive less attention from the model. Calculating the attention heatmap for the current iteration: During model forward propagation, the attention heatmap for different regions of the image at the current iteration is calculated according to the method in step one. This heatmap reflects the model's attention to different parts of the image when generating the current token under the current mask conditions.
[0054] In the forward propagation step of iterative attention refinement, an attention heatmap is generated after masking the original image with the current mask. Although the generation of this attention heatmap is affected by the mask, the model's basic attention mechanism and initial attention bias are based on the information reflected in the generated preliminary perception map. The preliminary perception map has already given the model a preliminary understanding of the regions in the image related to the current token. Based on this, the mask further guides the model to explore other potentially related regions.
[0055] 3. Clustering Identification. Clustering algorithms are used to categorize tokens into high-attention and low-attention groups: Clustering algorithms (such as variants of K-Means) analyze each location in the attention heatmap (corresponding to different regions of the image) and classify them according to their attention values. Regions with higher attention values are assigned to the high-attention group; these regions are typically where the model focuses on salient objects or important information in the current iteration. Regions with lower attention values are assigned to the low-attention group, which may contain secondary objects or background information. Identifying the set of high-attention tokens: The set of high-attention tokens is explicitly marked from the clustering results; these tokens correspond to the parts of the image that require focused attention.
[0056] The clustering identification step uses K-means to categorize tokens in the attention heatmap into high-attention and low-attention groups. This classification is based on the initial attention distribution reflected in the preliminary perception map. The preliminary perception map has already identified the main attention regions, and the attention heatmap generated during the iteration process further refines the level of attention based on this.
[0057] 4. Update the mask. Set the corresponding positions to 0 in the next round of masking (masking): In the next iteration, to force the model to stop over-focusing on already identified high-interest regions, set the corresponding positions of these regions in the mask to 0. This way, during the next round of forward propagation, the model will be more inclined to search for other under-focused regions in the image, thus achieving a more comprehensive extraction of image information.
[0058] The mask update step forces the model to focus on other regions in the next iteration. These high-interest regions are related to the main relevant regions reflected in the initial perception map. The initial perception map identifies the primary targets of interest; during iteration, mask updates gradually eliminate these primary regions, guiding the model to discover secondary relevant regions.
[0059] Based on the initial perception map, the original image is processed by the iterative attention refinement strategy to generate heatmaps, resulting in multiple corresponding heatmaps.
[0060] In this embodiment, based on the initial perceptual map described above, the original image can be subjected to iterative heatmap generation processing using the strategy content of the iterative attention refinement strategy described above, so as to obtain multiple generated heatmaps.
[0061] Determine whether the current iteration meets the preset termination condition.
[0062] In this embodiment, the sum of the heatmaps generated at the moment can be calculated. If the sum is less than the attention threshold or the maximum number of iterations is reached, it indicates that the iteration termination condition is met and the iteration stops. Otherwise, it is determined that the iteration termination condition is not met and the corresponding iterative cycle heatmap generation process will continue to be executed until the iteration termination condition is met.
[0063] If so, all the heatmaps are summed to obtain the corresponding first heatmap.
[0064] In this embodiment, during the iterative attention refinement process, each iteration generates a corresponding attention heatmap. These heatmaps are then collected to form a heatmap set. Assuming N iterations are performed, N distinct heatmaps are obtained. Each heatmap is a matrix, and the elements in the matrix represent the degree of attention the model pays to a specific region of the image in the corresponding iteration.
[0065] Then, all the collected heatmaps are accumulated. Specifically, for each location in the image (i.e., each element in the heatmap matrix), its corresponding element value in all N heatmaps is added together to obtain the accumulated heatmap matrix, i.e., the first heatmap, which integrates the attention paid to different regions of the image by the model in all iteration rounds.
[0066] The first heatmap is normalized to obtain the corresponding second heatmap.
[0067] In this embodiment, the goal of the fusion process, which involves accumulating and normalizing the heatmaps, is to obtain a perceptual map that more comprehensively reflects the regions in the image relevant to the currently generated token. The initial perceptual map provides an initial heatmap focusing on the main relevant regions, while the heatmaps generated during the iterative process gradually supplement information on the secondary relevant regions.
[0068] In this embodiment, the range of element values in the accumulated heatmap matrix may be large. To facilitate subsequent analysis and use, it is necessary to normalize the values so that they fall within a specific range, such as the [0, 1] interval. Then, a normalization method, such as linear normalization, can be used to find the maximum and minimum values in the accumulated heatmap matrix. Next, for each element in the matrix, normalization is performed according to the linear normalization formula, and the normalized second heatmap is used as the corresponding first perceptual map. This first perceptual map more comprehensively and accurately reflects the importance of each region of the image; a larger value indicates a higher level of attention and importance for that region during the model's text description generation process.
[0069] The second heat map is used as the first sensing map.
[0070] This application obtains a preset iterative attention refinement strategy; based on the initial perceptual map, it uses the iterative attention refinement strategy to generate heatmaps from the original image, resulting in multiple corresponding heatmaps; then it determines whether the current condition meets the preset iteration termination condition; if so, it accumulates all heatmaps to obtain the corresponding first heatmap; then it normalizes the first heatmap to obtain the corresponding second heatmap; subsequently, it uses the second heatmap as the first perceptual map. Based on the above processing flow, this application, by using the iterative attention refinement strategy to iteratively refine the initial perceptual map, can intelligently and accurately construct a more complete association system between image regions and currently generated tokens. Through multiple iterations, the model can gain a deeper understanding of the relationship between various regions in the image and tokens, thereby generating a first perceptual map that accurately reflects this complex relationship, ensuring the accuracy of the obtained first perceptual map.
[0071] In some alternative implementations, step S204 includes the following steps: The first perceptual map is normalized to obtain the corresponding first generated perceptual map.
[0072] In this embodiment, the normalization process includes scaling the values of the obtained first perceptual map to the range of [0, 1], which is to make the values of the first perceptual map have a uniform range, so as to facilitate subsequent processing.
[0073] The first generated perceptual map is subjected to variance amplification processing based on a preset objective function to obtain the corresponding second generated perceptual map.
[0074] In this embodiment, the objective function can specifically be the Sigmoid function. By multiplying the normalized first generated perceptual map by an enhancement coefficient (e.g., 10), this enhancement coefficient can amplify the difference between high and low attention regions in the perceptual map. Finally, through processing with the Sigmoid function, the Sigmoid function can map the input value to the (0, 1) interval and further enhance the contrast between high and low attention regions. For example, for a token in the perceptual map, its original value is 0.3, which remains 0.3 after normalization. Multiplying it by the enhancement coefficient 10 makes it 3. After processing with the Sigmoid function, its value will be closer to 1 (if the original value is large) or closer to 0 (if the original value is small), thereby highlighting the difference between tokens with high attention and tokens with low attention.
[0075] The second generated perceptual map is smoothed using a preset mean filter to obtain the corresponding third generated perceptual map.
[0076] In this embodiment, the second generated perceptual map can be smoothed using a mean filter with a specific kernel size (e.g., 3×3). The mean filter eliminates noise and local outliers in the perceptual map by averaging the values of each pixel and its surrounding pixels.
[0077] The third generated perceptual map is upsampled to obtain a fourth generated perceptual map with a resolution that matches that of the original image.
[0078] In this embodiment, the smoothed third generated perceptual map can be upsampled to the original image resolution using bilinear interpolation to obtain the corresponding fourth generated perceptual map. Bilinear interpolation is a commonly used image interpolation algorithm that calculates the value of a new pixel based on the values of surrounding known pixels through linear interpolation, thereby restoring the perceptual map to the same resolution as the original image while maintaining the overall image structure.
[0079] The fourth generated perceptual map is used as the second perceptual map.
[0080] This application obtains a first generated perceptual map by normalizing a first perceptual map; then, it performs variance amplification on the first generated perceptual map based on a preset objective function to obtain a second generated perceptual map; subsequently, it smooths the second generated perceptual map based on a preset mean filter to obtain a third generated perceptual map; and then it upsamples the third generated perceptual map to obtain a fourth generated perceptual map with a resolution matching the original image; finally, the fourth generated perceptual map is used as the second perceptual map. Based on the above processing flow, this application performs normalization, variance amplification, smoothing, and upsampling on the first perceptual map. Normalization ensures that the perceptual map values have a uniform range, the objective function further enhances the contrast between key and non-key regions in the perceptual map, enabling subsequent processing to more accurately identify key regions in the image. Smoothing eliminates noise in the perceptual map, making the identification of key regions more accurate. Upsampling to the original image resolution facilitates the subsequent mapping of the perceptual map to the original image, enabling the localization of key regions in the image, thereby effectively improving the quality of the generated second perceptual map.
[0081] In some alternative implementations, step S205 includes the following steps: Treating the second perceptual map as a probability mass function, the cumulative distribution function is calculated along the horizontal and vertical directions to obtain the corresponding generation function.
[0082] In this embodiment, the generated second perceptual map is treated as a probability mass function, and the cumulative distribution function (CDF) is calculated along the horizontal (X) and vertical (Y) directions, respectively. For the horizontal direction, each row of pixel values in the second perceptual map is processed, and the pixel values are accumulated sequentially from left to right to obtain the horizontal CDF (first cumulative distribution function, or first generating function). For the vertical direction, each column of pixel values in the perceptual map is processed, and the pixel values are accumulated sequentially from top to bottom to obtain the vertical CDF (second cumulative distribution function, or second generating function). Then, the horizontal and vertical CDFs are integrated to obtain the corresponding generating function.
[0083] In this process, edge distribution decomposition breaks down the two-dimensional perceptual map into two one-dimensional CDFs, facilitating subsequent inverse transform sampling. By calculating the CDFs, the probability distribution of the perceptual map in the horizontal and vertical directions can be understood.
[0084] The original image is subjected to inverse transform sampling based on the generating function to obtain the corresponding target coordinate information.
[0085] In this embodiment, the specific implementation process of performing inverse transform sampling on the original image based on the generation function to obtain the corresponding target coordinate information will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0086] The target coordinate information is processed using a bilinear interpolation algorithm to generate the corresponding image data.
[0087] In this embodiment, based on the original image coordinates obtained from inverse transform sampling (i.e., the target coordinate information mentioned above), the value of each pixel in the output image is calculated using a bilinear interpolation algorithm. Bilinear interpolation can calculate the value of a new pixel by linear interpolation based on the values of surrounding known pixels, thereby generating magnified image data. During the magnification process, because the sampling points in key areas are more dense, these areas will occupy more pixels in the output image, achieving a magnification effect, while the background and overall topology are also preserved.
[0088] The image data is used as the target image.
[0089] This application treats the second perceptual image as a probability quality function, calculates the cumulative distribution function along the horizontal and vertical directions to obtain the corresponding generation function; then, based on the generation function, performs inverse transform sampling processing on the original image to obtain the corresponding target coordinate information; subsequently, it performs image generation processing on the target coordinate information based on the bilinear interpolation algorithm to obtain the corresponding image data; finally, it uses the image data as the target image. Based on the above image generation steps, this application effectively ensures the quality of the generated target image by converting the coordinate information obtained from inverse transform sampling into actual image data and by using the bilinear interpolation algorithm, so that the magnified image highlights the key areas while maintaining the overall coherence of the image.
[0090] In some optional implementations of this embodiment, the step of performing inverse transform sampling processing on the original image based on the generation function to obtain the corresponding target coordinate information includes the following steps: Obtain the inverse function corresponding to the generating function.
[0091] In this embodiment, the inverse function of the cumulative distribution function (CDF) is: given a probability value (within the interval [0, 1]), the inverse function of the CDF can find the corresponding original variable value. In image processing, this means finding the corresponding pixel position based on the given probability value. The aforementioned inverse function includes both the horizontal and vertical inverse CDF functions.
[0092] Get uniformly distributed random numbers generated within a specified interval.
[0093] In this embodiment, the specified interval is specifically [0, 1]. A uniformly distributed random number can be generated within the [0, 1] interval, and this random number represents a probability value. Within the [0, 1] interval, each number has an equal probability of being drawn. For example, a randomly generated number could be 0.2, 0.56, 0.99, etc., and these numbers are uniformly distributed within this interval.
[0094] The random number is mapped to the coordinates of the original image based on the inverse function to obtain the corresponding coordinate information.
[0095] In this embodiment, the inverse CDF function in the horizontal direction is used to map the generated random number onto the horizontal coordinates of the original image to obtain the corresponding first coordinate information. Since the CDF is monotonically increasing, the corresponding pixel position can be found through lookup or interpolation. For example, if the CDF curve value at pixel position x is 0.3, then this random number 0.3 corresponds to pixel position x. The same operation is performed in the vertical direction: a uniformly distributed random number in the interval [0, 1] is generated, and the corresponding vertical pixel position is found by using the inverse CDF function in the vertical direction to map the generated uniformly distributed random number onto the vertical coordinates of the original image to obtain the corresponding second coordinate information. Then, the obtained first coordinate information and second coordinate information are integrated to obtain the corresponding coordinate information.
[0096] In areas with high perceptual map values, the sampling points are more densely packed due to the faster increase in CDF, thus occupying more pixels in the output image (i.e., being magnified); conversely, in areas with low perceptual map values, the sampling points are compressed.
[0097] The coordinate information is used as the target coordinate information.
[0098] This application obtains the inverse function corresponding to the generating function; then, it obtains uniformly distributed random numbers generated within a specified interval; subsequently, it maps the random numbers to the coordinates of the original image based on the inverse function to obtain the corresponding coordinate information; and finally, it uses the coordinate information as the target coordinate information. Based on the above processing flow, inverse transform sampling is a sampling method based on probability distribution. This application uses inverse transform sampling to perform nonlinear resampling on the original image based on the attention information in the perceived image, so that the key areas are magnified in the output image while preserving the overall structure of the image, ensuring the accuracy of the obtained target coordinate information.
[0099] In some optional implementations of this embodiment, step S206 includes the following steps: Get the preset text context.
[0100] In this embodiment, the above-mentioned text context refers to the text context of the current stage. If it is an intermediate process of generating a piece of text, the existing text is the context; if it is generated from scratch, the text context is some starting markers or preset guiding information.
[0101] The visual language model is used to analyze and predict the target image and the text context to generate corresponding probability distribution information.
[0102] In this embodiment, the above analysis and prediction processing is the process of generating a probability distribution. Specifically, this includes: after receiving the magnified target image, the visual language model, combined with the current text context, generates the probability distribution of the next token through its complex internal neural network calculation process, typically presented in the form of Logits. Logits is a vector where each element corresponds to a possible token, and its value indicates the probability of that token being generated. For example, in a financial insurance scenario, the model is generating text about an insurance product introduction. The current text context is "This insurance product provides comprehensive coverage, including...". The model might generate a Logits vector where tokens related to the insurance coverage type, such as "medical," "accident," and "property," have larger element values, indicating a higher probability of these tokens being generated.
[0103] Based on a preset target sampling strategy, the probability distribution information and the text context are processed to generate text until a preset stopping condition is met and the corresponding text output is obtained.
[0104] In this embodiment, the target sampling strategy described above can employ either random sampling or greedy sampling. Specifically, random sampling includes randomly selecting based on the probability value corresponding to each element in the Logits vector. For example, if the probability of "medical" in the Logits vector is 0.3, the probability of "accident" is 0.4, and the probability of "property" is 0.3, then there is a 30% probability of selecting "medical," a 40% probability of selecting "accident," and a 30% probability of selecting "property." Finally, a token is randomly selected as the next generated token. Greedy sampling includes directly selecting the token with the highest probability value in the Logits vector as the next generated token. For example, in the above example, "accident" has the highest probability, so "accident" is directly selected as the next token.
[0105] Then, after sampling and generating a token, it is added to the existing text context to form a new text context. The process of generating probability distributions and sampling to generate tokens is then repeated, continuously generating the next token, until a preset stopping condition is met (such as reaching a specified length, encountering a specific end marker, etc.), finally obtaining the complete text output, which serves as the corresponding text response data.
[0106] The text output is used as the text response data.
[0107] This application obtains a preset text context; then, based on a visual language model, it analyzes and predicts the target image and text context to generate corresponding probability distribution information; subsequently, based on a preset target sampling strategy, it performs text generation processing on the probability distribution information and text context until a preset stopping condition is met and the corresponding text output is obtained; finally, the text output is used as text response data. Based on the above processing flow, this application uses a visual language model to analyze and predict the target image and text context to generate probability distribution information, and then uses a target sampling strategy to sample from the probability distribution information to obtain the final text response data, thereby utilizing clearer visual details for decoding and effectively improving the accuracy of the generated text response data.
[0108] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0109] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0110] Furthermore, this application proposes a decoding framework called a "Perception Magnifier (PM)". Its core innovation lies in: Adaptive perceptual magnification decoding: At each step of the generation process, the high-interest regions in the image are dynamically "magnified" (resampled) rather than cropped, thereby significantly improving the effective resolution of key regions while preserving the overall structure of the image.
[0111] Iterative Refinement: This proposes an iterative masking mechanism that forces the model to focus on visual regions that were ignored in the first round but are still relevant through multiple rounds of forward propagation, thus solving the problem of overly singular attention focus.
[0112] Structure-preserving resampling based on probability distribution: This method constructs a probability quality function using an attention heatmap and performs non-linear deformation and magnification of the image through inverse transform sampling. This approach avoids context loss caused by cropping and achieves pixel-level smooth transitions.
[0113] Gradient-free perception map construction: Key regions are located by simply aggregating the Attention weights during the model inference process, without the need for backpropagation to calculate gradients, resulting in higher computational efficiency and easier integration.
[0114] Furthermore, this application has the following significant advantages: 1. Significantly reduces hallucination rate: Experiments have shown that in authoritative hallucination benchmark tests such as POPE and MME, this solution (PM) outperforms existing greedy decoding, VCD, OPERA and other solutions in both object existence judgment and attribute recognition.
[0115] 2. Preserving Model Inference Ability: This is a "lossless" cognitive approach. Since PM does not forcibly distort the Logits distribution through contrastive decoding, nor does it sever the image context like ViCrop, this approach does not exhibit the performance degradation common to other approaches in the MME Cognition subset, and even outperforms the baseline.
[0116] 3. Enhanced fine-grained detail perception: Enables the model to "see" small objects or details (such as bottles or specific text) that were originally blurred due to resolution compression. It provides richer information through physical resampling, rather than just adjusting the weights.
[0117] 4. Preservation of structure and spatial consistency: The scaling algorithm used is a type of elastic deformation (warping), which, compared to direct cropping, preserves the relative positions between objects and the background context, making the model accurate when answering questions involving spatial relationships.
[0118] 5. Plug and play and versatility: This solution does not require retraining or fine-tuning of the model (training-free), is applicable to most interleaved VLMs based on the Transformer structure, and has wide applicability.
[0119] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0120] It should be emphasized that, to further ensure the privacy and security of the aforementioned text response data, the text response data can also be stored in a node of a blockchain.
[0121] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0122] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0123] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0124] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0125] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an image processing device based on artificial intelligence, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0126] like Figure 3 As shown, the AI-based image processing device 300 described in this embodiment includes: a receiving module 301, a construction module 302, a thinning module 303, an optimization module 304, a sampling module 305, a generation module 306, and an output module 307. Wherein: The receiving module 301 is used to receive the input raw image; The construction module 302 is used to construct a corresponding initial perceptual map based on a preset self-attention mechanism during the process of decoding the original image using a preset visual language model. The refinement module 303 is used to perform iterative refinement processing based on the initial perception map to obtain the corresponding first perception map; The optimization module 304 is used to optimize the first perceptual map based on a preset post-processing strategy to obtain the corresponding second perceptual map. The sampling module 305 is used to perform nonlinear resampling processing on the original image based on the second perceptual map to generate a corresponding target image; The generation module 306 is used to perform reasoning processing on the target image based on the visual language model to generate corresponding text response data; The output module 307 is used to output the text response data.
[0127] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based image processing method in the aforementioned embodiments, and will not be repeated here.
[0128] In some optional implementations of this embodiment, the construction module 302 includes: The first determining submodule is used to perform hierarchical analysis on the visual language model in order to determine a specified intermediate layer from the visual language model; The first acquisition submodule is used to acquire the attention weights corresponding to the specified intermediate layer; The second acquisition submodule is used to acquire the preset pooling aggregation strategy; The aggregation submodule is used to aggregate the attention weights based on the pooling aggregation strategy to obtain the corresponding basic heatmap. The second determining submodule is used to use the basic heat map as the initial sensing map.
[0129] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based image processing method in the aforementioned embodiments, and will not be repeated here.
[0130] In some optional implementations of this embodiment, the refinement module 303 includes: The third acquisition submodule is used to acquire the preset iterative attention refinement strategy; The first generation submodule is used to perform heatmap generation processing on the original image based on the initial perceptual map and using the iterative attention refinement strategy to obtain multiple corresponding heatmaps. The judgment submodule is used to determine whether the current iteration meets the preset termination condition; The accumulation submodule is used to accumulate all the heatmaps if the condition is met, to obtain the corresponding first heatmap. The normalization submodule is used to normalize the first heatmap to obtain the corresponding second heatmap. The third determining submodule is used to use the second heat map as the first sensing map.
[0131] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based image processing method in the aforementioned embodiments, and will not be repeated here.
[0132] In some optional implementations of this embodiment, the optimization module 304 includes: The first processing submodule is used to normalize the first perception map to obtain the corresponding first generated perception map. The second processing submodule is used to perform variance amplification processing on the first generated perception map based on a preset objective function to obtain the corresponding second generated perception map. The third processing submodule is used to smooth the second generated perception map based on a preset mean filter to obtain the corresponding third generated perception map. The fourth processing submodule is used to perform upsampling processing on the third generated perceptual map to obtain a fourth generated perceptual map with a resolution matching the original image; The fourth determining submodule is used to use the fourth generated perception map as the second perception map.
[0133] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based image processing method in the aforementioned embodiments, and will not be repeated here.
[0134] In some optional implementations of this embodiment, the sampling module 305 includes: The calculation submodule is used to treat the second perception map as a probability mass function and calculate the cumulative distribution function along the horizontal and vertical directions to obtain the corresponding generation function; The fifth processing submodule is used to perform inverse transform sampling processing on the original image based on the generation function to obtain the corresponding target coordinate information; The second generation submodule is used to perform image generation processing on the target coordinate information based on the bilinear interpolation algorithm to obtain the corresponding image data. The fifth determining submodule is used to use the image data as the target image.
[0135] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based image processing method in the aforementioned embodiments, and will not be repeated here. In some optional implementations of this embodiment, the fifth processing submodule includes: The first acquisition unit is used to acquire the inverse function corresponding to the generating function; The second acquisition unit is used to acquire uniformly distributed random numbers generated within a specified interval; A mapping unit is used to map the random number to the coordinates of the original image based on the inverse function to obtain the corresponding coordinate information; A determining unit is used to use the coordinate information as the target coordinate information.
[0136] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based image processing method in the aforementioned embodiments, and will not be repeated here.
[0137] In some optional implementations of this embodiment, the generation module 306 includes: The fourth submodule is used to obtain the preset text context; The sixth processing submodule is used to analyze and predict the target image and the text context based on the visual language model, and generate corresponding probability distribution information; The third generation submodule is used to perform text generation processing on the probability distribution information and the text context based on a preset target sampling strategy until a preset stopping condition is met and the corresponding text output is obtained. The sixth determining submodule is used to use the text output as the text response data.
[0138] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based image processing method in the aforementioned embodiments, and will not be repeated here. To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0139] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0140] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0141] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for image processing methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0142] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the artificial intelligence-based image processing method.
[0143] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0144] Compared with the prior art, the embodiments of this application have the following beneficial effects: In this embodiment, an input original image is first received; then, during the decoding process of the original image using a preset visual language model, a corresponding initial perceptual map is constructed based on a preset self-attention mechanism; subsequently, iterative refinement processing is performed based on the initial perceptual map to obtain a corresponding first perceptual map; and the first perceptual map is optimized based on a preset post-processing strategy to obtain a corresponding second perceptual map; subsequently, nonlinear resampling processing is performed on the original image based on the second perceptual map to generate a corresponding target image; further, inference processing is performed on the target image based on the visual language model to generate corresponding text response data; finally, the text response data is output. Based on the above automated processing flow, in the process of decoding the input original image using the visual language model in this application, an initial perceptual map is constructed by using a self-attention mechanism. Then, the initial perceptual map is iteratively refined and optimized to obtain a second perceptual map. Based on the obtained second perceptual map, the original image is nonlinearly resampled to generate a target image. Then, the target image is inferred based on the visual language model, thereby realizing decoding using clearer visual details, effectively avoiding visual illusion problems, and improving the accuracy of the generated text response data.
[0145] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based image processing method described above.
[0146] Compared with the prior art, the embodiments of this application have the following main advantages: In this embodiment, an input original image is first received; then, during the decoding process of the original image using a preset visual language model, a corresponding initial perceptual map is constructed based on a preset self-attention mechanism; subsequently, iterative refinement processing is performed based on the initial perceptual map to obtain a corresponding first perceptual map; and the first perceptual map is optimized based on a preset post-processing strategy to obtain a corresponding second perceptual map; subsequently, nonlinear resampling processing is performed on the original image based on the second perceptual map to generate a corresponding target image; further, inference processing is performed on the target image based on the visual language model to generate corresponding text response data; finally, the text response data is output. Based on the above automated processing flow, in the process of decoding the input original image using the visual language model in this application, an initial perceptual map is constructed by using a self-attention mechanism. Then, the initial perceptual map is iteratively refined and optimized to obtain a second perceptual map. Based on the obtained second perceptual map, the original image is nonlinearly resampled to generate a target image. Then, the target image is inferred based on the visual language model, thereby realizing decoding using clearer visual details, effectively avoiding visual illusion problems, and improving the accuracy of the generated text response data.
[0147] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0148] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. An image processing method based on artificial intelligence, characterized in that, Includes the following steps: Receive the raw input image; During the decoding process of the original image using a preset visual language model, a corresponding initial perceptual map is constructed based on a preset self-attention mechanism. Based on the initial perception map, iterative refinement is performed to obtain the corresponding first perception map; The first perceptual map is optimized based on a preset post-processing strategy to obtain the corresponding second perceptual map. Based on the second perceptual map, the original image is subjected to nonlinear resampling processing to generate the corresponding target image; Based on the visual language model, the target image is subjected to inference processing to generate corresponding text response data; The text response data is then processed for output.
2. The image processing method based on artificial intelligence according to claim 1, characterized in that, The step of constructing the corresponding initial perceptual map based on a preset self-attention mechanism specifically includes: Perform hierarchical analysis on the visual language model to determine a specified intermediate layer from the visual language model; Obtain the attention weights corresponding to the specified intermediate layer; Obtain the preset pooling aggregation strategy; The attention weights are aggregated based on the pooling aggregation strategy to obtain the corresponding basic heatmap. The basic heatmap is used as the initial sensing map.
3. The image processing method based on artificial intelligence according to claim 1, characterized in that, The step of iteratively refining the initial perceptual map to obtain the corresponding first perceptual map specifically includes: Obtain a pre-defined iterative attention refinement strategy; Based on the initial perceptual map, the original image is processed by heatmap generation using the iterative attention refinement strategy to obtain multiple corresponding heatmaps. Determine whether the current iteration meets the preset termination condition; If so, sum all the heatmaps to obtain the corresponding first heatmap; The first heatmap is normalized to obtain the corresponding second heatmap. The second heat map is used as the first sensing map.
4. The image processing method based on artificial intelligence according to claim 1, characterized in that, The step of optimizing the first perceptual map based on a preset post-processing strategy to obtain the corresponding second perceptual map specifically includes: The first perceptual map is normalized to obtain the corresponding first generated perceptual map; Based on a preset objective function, the first generated perception map is subjected to variance amplification processing to obtain the corresponding second generated perception map. The second generated perceptual map is smoothed based on a preset mean filter to obtain the corresponding third generated perceptual map. The third generated perceptual map is upsampled to obtain a fourth generated perceptual map with a resolution that matches that of the original image; The fourth generated perceptual map is used as the second perceptual map.
5. The image processing method based on artificial intelligence according to claim 1, characterized in that, The step of performing nonlinear resampling processing on the original image based on the second perceptual map to generate the corresponding target image specifically includes: Treating the second perceptual map as a probability mass function, the cumulative distribution function is calculated along the horizontal and vertical directions to obtain the corresponding generation function; Based on the generating function, the original image is subjected to inverse transform sampling processing to obtain the corresponding target coordinate information; The target coordinate information is processed by bilinear interpolation algorithm to generate the corresponding image data. The image data is used as the target image.
6. The image processing method based on artificial intelligence according to claim 5, characterized in that, The step of performing inverse transform sampling processing on the original image based on the generation function to obtain the corresponding target coordinate information specifically includes: Obtain the inverse function corresponding to the generating function; Get uniformly distributed random numbers generated within a specified interval; Based on the inverse function, the random number is mapped to the coordinates of the original image to obtain the corresponding coordinate information; The coordinate information is used as the target coordinate information.
7. The image processing method based on artificial intelligence according to claim 1, characterized in that, The step of performing inference processing on the target image based on the visual language model to generate corresponding text response data specifically includes: Get the preset text context; Based on the visual language model, the target image and the text context are analyzed and predicted to generate corresponding probability distribution information. Based on a preset target sampling strategy, the probability distribution information and the text context are processed to generate text until a preset stopping condition is met and the corresponding text output is obtained. The text output is used as the text response data.
8. An image processing device based on artificial intelligence, characterized in that, include: The receiving module is used to receive the input raw image; The construction module is used to construct the corresponding initial perceptual map based on a preset self-attention mechanism during the process of decoding the original image using a preset visual language model. The refinement module is used to perform iterative refinement processing based on the initial perception map to obtain the corresponding first perception map; An optimization module is used to optimize the first perceptual map based on a preset post-processing strategy to obtain a corresponding second perceptual map. The sampling module is used to perform non-linear resampling processing on the original image based on the second perceptual map to generate a corresponding target image; The generation module is used to perform inference processing on the target image based on the visual language model to generate corresponding text response data. The output module is used to process the text response data.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the artificial intelligence-based image processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based image processing method as described in any one of claims 1 to 7.