Data processing method and device based on artificial intelligence, computer equipment and medium
By adjusting the attention matrix of LVLMs through modal bias analysis and bidirectional attention intervention mechanism, and combining it with contrastive decoding strategy, the object illusion problem of LVLMs is solved, improving the accuracy and fluency of image text generation. It is applicable to image analysis in fields such as finance and medicine.
Patent Information
- Application Number
- CN202511299299.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-12-30
AI Technical Summary
Large visual language models (LVLMs) suffer from object illusion when generating image-text descriptions, resulting in low accuracy and limiting their application in real-world scenarios where accuracy is extremely important.
By adjusting the attention matrix of a large visual language model based on modal bias analysis and a bidirectional attention intervention mechanism, and combining it with a contrastive decoding strategy, the generation of illusions is reduced, thereby improving the accuracy and fluency of the generated text.
It effectively reduces the generation of hallucinations, improves the accuracy and fluency of generated text, and is suitable for image analysis scenarios in fields such as finance and healthcare.
Smart Images

Figure CN121233745A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as fintech and digital healthcare, particularly to data processing methods, devices, computer equipment, and storage media based on artificial intelligence. Background Technology
[0002] In the field of large visual language models (LVLMs), LVLMs have demonstrated powerful capabilities in multimodal understanding, reasoning, and human-computer interaction in recent years. These models typically consist of a visual encoder and a large language model (LLM), capable of handling text-image interleaved inputs, and show great application potential in numerous fields such as finance and healthcare.
[0003] However, despite significant progress in LVLMs, a serious object illusion problem persists. This manifests as the model generating text descriptions that include objects that do not actually exist in the image, or incorrectly determining the presence of a specific object in the image. This illusion problem severely limits the application and deployment of LVLMs in real-world scenarios where accuracy is paramount, resulting in low accuracy of text data generated from images.
[0004] For example, in the financial sector, LVLMs (Visual Literacy Models) might be used to identify financial data in images and generate text descriptions. However, due to the object illusion problem, the model might include financial items that are not present in the image (such as non-existent expense categories) in the text description, or incorrectly determine the existence of a key financial indicator (such as profit growth rate). Such erroneous descriptions can lead to inaccurate financial analysis results, affecting financial institutions' decision-making and reducing the efficiency and accuracy of business processing.
[0005] Therefore, there is an urgent need to provide a method or system that can effectively solve the object illusion problem of LVLMs, so as to improve the accuracy of image text generation and meet the practical application needs of multiple fields. Summary of the Invention
[0006] The purpose of this application is to propose a data processing method, apparatus, computer device, and storage medium based on artificial intelligence to solve the problem of object illusion in existing large-scale visual language models, which leads to low accuracy in image text generation.
[0007] Firstly, an artificial intelligence-based data processing method is provided, including:
[0008] Receive the input target image and the corresponding descriptive text;
[0009] The target image and the descriptive text are processed based on a pre-set large-scale visual language model to generate corresponding lexical units;
[0010] The word units are quantified based on a preset modal bias analysis strategy to obtain the modal bias quantification results of the word units;
[0011] If the modal bias quantification result meets the preset attention intervention conditions, then the original attention matrix of the word is calculated, and the weight adjustment processing of the original attention matrix is performed based on the modal bias quantification result to obtain the corresponding specified attention matrix;
[0012] The original attention matrix is processed to obtain the corresponding first conditional probability distribution, and the specified attention matrix is processed to obtain the corresponding second conditional probability distribution;
[0013] The first conditional probability distribution and the second conditional probability distribution are fused to obtain the corresponding target conditional probability distribution.
[0014] The target conditional probability distribution is decoded to generate corresponding target words, and target text corresponding to the target image is generated based on the target words.
[0015] The target text is then processed for output.
[0016] Secondly, an artificial intelligence-based data processing device is provided, comprising:
[0017] A receiving module is used to receive an input target image and descriptive text corresponding to the target image;
[0018] The first processing module is used to process the target image and the descriptive text based on a preset large-scale visual language model to generate corresponding word units;
[0019] The quantization module is used to quantize the word units based on a preset modal bias analysis strategy to obtain the modal bias quantization result of the word units;
[0020] The second processing module is used to calculate the original attention matrix of the word if the modality bias quantification result meets the preset attention intervention conditions, and to perform weight adjustment processing on the original attention matrix based on the modality bias quantification result to obtain the corresponding specified attention matrix.
[0021] The calculation module is used to calculate and process the original attention matrix to obtain the corresponding first conditional probability distribution, and to calculate and process the specified attention matrix to obtain the corresponding second conditional probability distribution;
[0022] The fusion module is used to fuse the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding target conditional probability distribution.
[0023] The third processing module is used to decode the target conditional probability distribution to generate corresponding target words, and generate target text corresponding to the target image based on the target words;
[0024] The output module is used to process the target text for output.
[0025] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based data processing method.
[0026] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned artificial intelligence-based data processing method.
[0027] In the above-mentioned scheme implemented by the data processing method, device, computer equipment, and storage medium based on artificial intelligence, the following steps are taken: First, an input target image and a descriptive text corresponding to the target image are received; then, the target image and the descriptive text are processed based on a preset large-scale visual language model to generate corresponding lexical units; next, the lexical units are quantized based on a preset modal bias analysis strategy to obtain the modal bias quantization result of the lexical units; if the modal bias quantization result meets the preset attention intervention conditions, the original attention matrix of the lexical units is calculated, and the weight of the original attention matrix is adjusted based on the modal bias quantization result to obtain the corresponding specified attention matrix; subsequently, the original attention matrix is calculated to obtain the corresponding first conditional probability distribution, and the specified attention matrix is calculated to obtain the corresponding second conditional probability distribution; the first conditional probability distribution and the second conditional probability distribution are fused to obtain the corresponding target conditional probability distribution; further, the target conditional probability distribution is decoded to generate the corresponding target lexical units, and target text corresponding to the target image is generated based on the target lexical units; finally, the target text is output. Based on the above automated processing flow, this application processes the target image and descriptive text using a large-scale visual language model to generate corresponding lexical units. Then, it quantizes the lexical units using a modal bias analysis strategy. If the modal bias quantization result meets preset attention intervention conditions, the original attention matrix of the lexical units is calculated. Based on the modal bias quantization result, the original attention matrix is weighted to obtain a specified attention matrix. Subsequently, the original attention matrix is processed to obtain a first conditional probability distribution, and the specified attention matrix is processed to obtain a second conditional probability distribution. The first and second conditional probability distributions are then fused to obtain a target conditional probability distribution. This target conditional probability distribution is then decoded to generate target lexical units, and target text corresponding to the target image is generated based on these target lexical units. Finally, the target text is output. Thus, by combining a large-scale visual language model, a modal bias analysis strategy, and an attention intervention mechanism, this application can effectively reduce the generation of illusions, thereby improving the accuracy and fluency of the generated target text. Attached Figure Description
[0028] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1This is an exemplary system architecture diagram to which this application can be applied;
[0030] Figure 2 This is a flowchart of an embodiment of the artificial intelligence-based data processing method according to this application;
[0031] Figure 3 This is a schematic diagram of a structure of an embodiment of the artificial intelligence-based data processing apparatus according to this application;
[0032] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0034] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0036] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0037] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0038] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0039] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0040] It should be noted that the artificial intelligence-based data processing method provided in the embodiments of this application is generally executed by a server / terminal device, and correspondingly, the artificial intelligence-based data processing device is generally set in the server / terminal device.
[0041] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0042] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the AI-based data processing method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The AI-based data processing method provided in this application can be applied to any scenario requiring image analysis, and thus can be applied to products in these scenarios, such as image analysis products in the financial insurance field. The AI-based data processing method includes the following steps:
[0043] Step S201: Receive the input target image and the descriptive text corresponding to the target image.
[0044] In this embodiment, the data processing method based on artificial intelligence runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire the input target image and corresponding descriptive text via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wide wireless) connections, and other currently known or future known wireless connection methods. The executing entity of this application is specifically a data processing system, which may be simply referred to as the system.
[0045] The target image mentioned above is the image to be processed (i.e., the image content needs to be described), and the description text mentioned above is the user instruction input by the relevant user. This application can be applied to image analysis scenarios in the fields of fintech and healthcare. For example, in the car insurance scenario in the financial insurance field, the content of the input target image can be: a picture showing a car involved in a collision, with the front of the car severely deformed, some car parts scattered around, a small tree that has been knocked over and leaning, and obvious skid marks and oil stains on the ground. The corresponding description text includes: "Please describe in detail the objects related to the vehicle accident in this picture and the surrounding environment caused by the accident, so as to provide accurate information for car insurance claims." Or, in the home property insurance scenario in the financial insurance field, the content of the input target image can be: an indoor scene picture, showing a room that has been damaged by fire. The walls are blackened, some of the paint has peeled off, furniture such as sofas and wardrobes are burned beyond recognition, the floor is full of charred debris and water, the window glass is broken, and the curtains have also been burned. The corresponding descriptive text includes: "Describe the damage to household property and belongings caused by the fire, as well as the overall damage to the room, for the purpose of assessing losses under home property insurance, based on this image."
[0046] In addition, in the disease analysis assistance scenario within the digital healthcare field, the input target image could be: a medical X-ray showing a patient's lung area. A distinct shadow area can be seen in the lungs, with an irregular shape, blurred edges, and somewhat disordered surrounding lung texture, showing a clear difference from a normal lung X-ray image. The corresponding descriptive text includes: "Please describe the location, shape, size, and other characteristics of the abnormal shadow in this lung X-ray to provide imaging evidence for disease diagnosis." Alternatively, in the rehabilitation monitoring scenario within the digital healthcare field, the input target image could be: a photograph of a patient performing rehabilitation training, using rehabilitation equipment to perform leg flexion and extension exercises. The photograph clearly shows the patient's leg posture, the type of rehabilitation equipment (such as elastic bands, rehabilitation bicycles, etc.), and the patient's facial expression due to exertion. The corresponding descriptive text includes: "Describe the equipment used by the patient during rehabilitation training in this image, their body movements and posture, and their facial expression to monitor the rehabilitation treatment effect and adjust the rehabilitation plan."
[0047] Specifically, this application proposes a novel, training-free intervention method for the reasoning stage, the core innovation of which lies in:
[0048] This application, based on in-depth empirical analysis, is the first to explicitly identify two distinct modes of illusion in LVLMs—"generative illusion" and "discriminative illusion"—and discovers that their root cause lies in a previously overlooked phenomenon of "modal bias." Specifically, in generating illusions, the model's attention is overly focused on visual information; while in generating discriminative illusions, the model relies excessively on textual information. This discovery provides a completely new perspective and theoretical foundation for solving the illusion problem.
[0049] A bidirectional attention intervention mechanism is proposed: Unlike existing technologies that only unidirectionally enhance vision or suppress language priors, this invention proposes a bidirectional, dynamic "text and visual attention intervention" mechanism. This mechanism can intelligently adjust the attention weights assigned to text and visual tokens during inference, forcing the model to process information from both modalities in a balanced manner, thereby eliminating unimodal bias and fundamentally reducing the generation of illusions.
[0050] Collaborative Contrastive Decoding Strategy: This application combines attention intervention with a contrastive decoding strategy. By comparing the output probability distribution after attention intervention with the original probability distribution, the influence of user commands (visual and textual) is further enhanced, while reducing the model's over-reliance on its own inherent, potentially erroneous, parameterized knowledge, thus synergistically improving the illusion suppression effect.
[0051] Achieving zero-cost, highly versatile deployment: This application presents a completely training-free method that only intervenes in the calculation of the internal attention matrix during the model inference stage, without involving any modification to model parameters. Therefore, there are no additional training costs or data annotation requirements. Furthermore, this method is highly versatile, serving as a plug-and-play module applicable to all LVLMs based on the Transformer architecture, and compatible with various decoding strategies.
[0052] Step S202: Process the target image and the descriptive text based on a preset large-scale visual language model to generate corresponding lexical units.
[0053] In this embodiment, the large-scale visual language model mentioned above is a model that includes a visual encoder, a projector, and a language decoder. The process of processing the target image and descriptive text based on the large-scale visual language model to generate lexical units includes: (1) Encoding stage: Text encoding: The user's descriptive text is encoded into a contextual representation of text lexical units by a text encoder (such as the Transformer layer of BERT or GPT). Visual encoding: The image is encoded into a feature representation of visual lexical units (such as [v_1, v_2, ..., v_N], where each v_i corresponds to a region of the image) by a visual encoder (such as ViT or ResNet+Transformer). (2) Decoding stage (generating lexical units). Initial state: The decoder receives the encoded text and visual features, as well as a special starting lexical unit ( <bos>Word by word generation: Query token: the token to be generated currently (e.g., when the t-th token is generated, the query is q_t). Attention calculation: text self-attention: q_t and the generated text history (h_1, h_2,..., h_{t-1}) calculate attention, capturing the language context. Visual cross-modal attention: q_t and visual tokens (v_1, v_2,..., v_N) calculate attention, capturing image details. Weight fusion: the text and visual attention weights are weighted and summed to obtain the context representation c_t of the current token. Token prediction: predict the probability distribution of the next token (e.g., P(w_t|w_{<t}, I, T), where I is the image and T is the user instruction) through c_t.
[0054] In step S203, the token is quantitatively processed based on a preset modal bias analysis strategy to obtain a modal bias quantitative result of the token.
[0055] In the embodiment, the specific implementation process of quantitatively processing the token based on the preset modal bias analysis strategy to obtain the modal bias quantitative result of the token will be further described in detail in subsequent specific embodiments, and will not be described in detail here.
[0056] In step S204, if the modal bias quantitative result meets a preset attention intervention condition, an original attention matrix of the token is calculated, and the original attention matrix is weight-adjusted based on the modal bias quantitative result to obtain a corresponding specified attention matrix.
[0057] In the embodiment, the modal bias quantitative result meeting the preset attention intervention condition means that the modal bias quantitative result is a result corresponding to detecting a generative hallucination tendency (e.g., VAR is too high, such as VAR>0.7) or a result corresponding to detecting a discriminative hallucination tendency (e.g., TAR is too high, such as TAR>0.6). The token can be calculated and processed based on the large visual language model to obtain the corresponding original attention matrix by using a general attention calculation method.
[0058] The specific implementation process of weight-adjusting the original attention matrix based on the modal bias quantitative result to obtain the corresponding specified attention matrix will be further described in detail in subsequent specific embodiments, and will not be described in detail here.
[0059] In step S205, the original attention matrix is calculated and processed to obtain a corresponding first conditional probability distribution, and the specified attention matrix is calculated and processed to obtain a corresponding second conditional probability distribution.
[0060] In this embodiment, in order to further amplify the effect of attention intervention and reduce the dependence of the model on internal error knowledge, a contrast decoding is introduced. Parallel computing probability distribution: when generating each Token, the system parallelly computes two paths: path one (after intervention): using the modified attention matrix A' l,h , complete the subsequent Transformer calculation to obtain a modified Token conditional probability distribution, i.e., the second conditional probability distribution p'(y k |y <k ), path two (original): using the original attention matrix A l,h , calculate the original Token conditional probability distribution, i.e., the second conditional probability distribution p(y k |y <k ).
[0061] The specific implementation process of calculating and processing the specified attention matrix to obtain the corresponding second conditional probability distribution will be described in further detail in subsequent specific embodiments, and will not be described in detail here.
[0062] Step S206, the first conditional probability distribution and the second conditional probability distribution are fused to obtain a corresponding target conditional probability distribution.
[0063] In this embodiment, the specific implementation process of fusing the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding target conditional probability distribution will be described in further detail in subsequent specific embodiments, and will not be described in detail here.
[0064] Step S207, the target conditional probability distribution is decoded to generate a corresponding target Token, and a target text corresponding to the target image is generated based on the target Token.
[0065] In this embodiment, the specific implementation process of decoding the target conditional probability distribution to generate the corresponding target Token will be described in further detail in subsequent specific embodiments, and will not be described in detail here. By repeatedly performing the step of generating the corresponding target Token for each Token, a complete sequence is generated, which is the target text.
[0066] Step S208, the target text is output.
[0067] In this embodiment, the generated target text can be sent to a related user to complete the output processing of the target text.
[0068] The application first receives an input target image and a description text corresponding to the target image; then processes the target image and the description text based on a preset large visual language model to generate corresponding word units; then quantitatively processes the word units based on a preset modal bias analysis strategy to obtain a modal bias quantitative result of the word units; if the modal bias quantitative result meets a preset attention intervention condition, an original attention matrix of the word units is calculated, and the original attention matrix is weight-adjusted based on the modal bias quantitative result to obtain a corresponding specified attention matrix; subsequently, a first conditional probability distribution corresponding to the original attention matrix is calculated, and a second conditional probability distribution corresponding to the specified attention matrix is calculated; and the first conditional probability distribution and the second conditional probability distribution are fused to obtain a target conditional probability distribution; the target conditional probability distribution is further decoded to generate a target word unit corresponding to the target image, and a target text corresponding to the target image is generated based on the target word unit; and finally, the target text is output. Based on the above automatic processing procedure, the target image and the description text are processed based on the use of the large visual language model to generate the corresponding word units, and then the word units are quantitatively processed based on the use of the modal bias analysis strategy. If the modal bias quantitative result meets the preset attention intervention condition, the original attention matrix of the word units is calculated, and the original attention matrix is weight-adjusted based on the modal bias quantitative result to obtain the specified attention matrix. Subsequently, the first conditional probability distribution is calculated based on the original attention matrix, and the second conditional probability distribution is calculated based on the specified attention matrix. Then, the first conditional probability distribution and the second conditional probability distribution are fused to obtain the target conditional probability distribution. The target word unit is generated by decoding the target conditional probability distribution, and the target text corresponding to the target image is generated based on the target word unit. Finally, the target text is output. In this way, by using the large visual language model, the modal bias analysis strategy and the attention intervention mechanism, the generation of hallucinations can be effectively reduced, and the accuracy and fluency of the generated target text can be improved.
[0069] In some optional implementations, step S203 includes the following steps:
[0070] In generating the specified word unit, a text attention ratio of the specified word unit is calculated based on a preset first calculation strategy; wherein the specified word unit is any one of all the word units.
[0071] In this embodiment, in order to accurately intervene, it is necessary to quantitatively analyze the attention degree of the model to different modalities when generating a specific Token. Two key indicators are defined:
[0072] Textual Attention Ratio (TAR): the sum of the attention weights of all input text Tokens X T when generating the kth Token.
[0073]
[0074] where A l,h (k,i) denotes the attention score of the kth Token being generated to the ith text Token in the lth layer and the hth attention head.
[0075] Visual Attention Ratio (VAR): similarly, the sum of the attention weights of all input image Tokens X V when generating the kth Token.
[0076]
[0077] where A l,h (k,j) denotes the attention score of the kth Token being generated to the jth visual Token in the lth layer and the hth attention head
[0078] Specifically, the strategy content of the first calculation strategy includes: when generating the kth Token, traversing all Transformer layers (l) and attention heads (h), accumulating the attention score of the Token to all text Tokens X T .
[0079] Based on the preset second calculation strategy, the visual attention ratio of the specified word piece is calculated.
[0080] In this embodiment, the strategy content of the second calculation strategy includes: when generating the kth Token, traversing all Transformer layers (l) and attention heads (h), accumulating the attention score of the Token to all image Tokens X V .
[0081] A preset bias threshold is obtained.
[0082] In this embodiment, the existence of modal bias is confirmed by analyzing a large number of hallucination samples in advance: generative hallucination Tokens usually exhibit high VAR and low TAR, while discriminative hallucination Tokens exhibit high TAR and low VAR.
[0083] Among them, the definition process of the bias threshold includes: generative hallucination tendency: condition: VAR>0.7 and TAR<0.3. Explanation: The model over-relied on visual information and generated objects that do not exist in the image (such as "red car").
[0084] Discriminative hallucination tendency: condition: TAR>0.6 and VAR<0.4. Explanation: The model over-relied on text information and ignored visual details (such as describing "still cat" as "flying cat").
[0085] Based on the bias threshold, the text attention ratio and the visual attention ratio are numerically analyzed to generate the modal bias quantification result of the specified token.
[0086] In this embodiment, by using the above-mentioned bias threshold to numerically compare the text attention ratio and the visual attention ratio, if VAR>0.7 and TAR<0.3 for a certain Token, it is marked as generative hallucination tendency (such as fictitious "red car"); if TAR>0.6 and VAR<0.4, it is marked as discriminative hallucination tendency (such as false judgment of object existence), thereby completing the generation process of the modal bias quantification result of the token.
[0087] The present application calculates the text attention ratio of the specified token based on the preset first calculation strategy when generating the specified token; wherein the specified token is any one of all the tokens; then calculates the visual attention ratio of the specified token based on the preset second calculation strategy; then obtains the preset bias threshold; and subsequently, based on the bias threshold, the text attention ratio and the visual attention ratio are numerically analyzed to generate the modal bias quantification result of the specified token. Based on the above processing procedure, when generating the specified token, the present application calculates the text attention ratio of the specified token based on the use of the first calculation strategy, and calculates the visual attention ratio of the specified token based on the use of the second calculation strategy, and then numerically analyzes the text attention ratio and the visual attention ratio based on the use of the bias threshold, so as to efficiently and accurately complete the quantification processing of the specified token and generate the modal bias quantification result of the specified token, thereby ensuring the accuracy of the obtained modal bias quantification result.
[0088] In some optional implementations of the present embodiment, step S204 includes the following steps:
[0089] The large visual language model is analyzed to determine the intervention level in the large visual language model.
[0090] In this embodiment, the goal of intervention level selection is to determine the Transformer layers that need to be intervened to avoid damaging the effective feature fusion of shallow layers while repairing the modal bias of deep layers. The specific determination process includes: (1) analyzing the attention entropy value. Method: Calculate the distribution variance (or entropy value) of the attention weight of each layer to quantify the dispersion degree of attention. For each layer l and head h, the variance of the attention score of all Token pairs is calculated. High variance indicates that the attention is dispersed (such as the multi-modal feature fusion of shallow layers), and low variance indicates that the attention is concentrated (such as the "settlement" phenomenon of deep layers). Observation: Shallow layers (first 6 layers): attention dispersion, involving preliminary interaction of multi-modal information. Deep layers (last 3 layers): multiple heads appear "attention settlement" (i.e., repeatedly focusing on a few Tokens of the same modality). (2) Determine the intervention layer. Selection basis: The attention settlement of deep layers (last 3 layers) directly leads to the solidification of modal bias, which needs to be intervened. The dispersed attention of shallow layers (first 6 layers) helps feature fusion, and intervention may damage effective information transmission. Conclusion: Only the last 3 layers are implemented for attention weight adjustment, and the natural attention distribution of shallow layers is preserved.
[0091] Based on the modal bias quantification result, the corresponding target attention intervention formula is called.
[0092] In this embodiment, the above target attention intervention formula can adopt a text attention intervention formula or a visual attention intervention formula. Among them, if the above modal bias quantification result is a result corresponding to detecting a generative hallucination tendency (such as VAR being too high, such as VAR>0.7), the text attention intervention formula is adopted. If the above modal bias quantification result is a result corresponding to detecting a discriminative hallucination tendency (such as TAR being too high, such as TAR>0.6), the visual attention intervention formula is adopted.
[0093] Based on the target attention intervention formula, the original attention matrix is processed for attention weight adjustment in the intervention level to obtain an adjusted attention matrix.
[0094] In this embodiment, the attention scores of text and visual Tokens are enhanced according to the target attention intervention formula:
[0095] (1) The text attention intervention formula includes:
[0096]
[0097] Wherein, α is a hyperparameter for controlling the strength of text attention enhancement. This operation aims to solve the problem of ignoring text information in generative hallucination.
[0098] (2) The visual attention intervention formula includes:
[0099]
[0100] where β is a hyper-parameter to control the strength of the visual attention boost. This operation aims to address the neglect of visual information in the discriminative hallucination.
[0101] By boosting the weights along their original attention directions, the model can be guided to focus on the user instructions more reliably, instead of introducing noise. The values of a and β can be preset and adjusted according to specific models and tasks.
[0102] Specifically, based on the use of the above target attention intervention formula, the original attention matrix is processed for attention weight adjustment in the above intervention level, and an adjusted attention matrix is obtained as the corresponding specified attention matrix.
[0103] The adjusted attention matrix is taken as the specified attention matrix.
[0104] The present application determines the intervention level in the large visual language model by analyzing the large visual language model; then calls the corresponding target attention intervention formula based on the modal bias quantification result; then processes the original attention matrix for attention weight adjustment in the intervention level based on the target attention intervention formula, and obtains an adjusted attention matrix; subsequently, the adjusted attention matrix is taken as the specified attention matrix. Based on the above processing procedure, the present application determines the intervention level in the large visual language model by analyzing the large visual language model, calls the corresponding target attention intervention formula based on the modal bias quantification result, and then processes the original attention matrix for attention weight adjustment in the intervention level based on the use of the target attention intervention formula, and obtains an adjusted attention matrix as the specified attention matrix, so that the weight adjustment processing can be automatically and accurately completed for the original attention matrix, and the accuracy of the generated specified attention matrix is improved.
[0105] In some optional implementations, the large visual language model includes a language decoder, and the language decoder includes a preset calculation layer and a correction layer; step S205 includes the following steps:
[0106] The specified attention matrix is input into the language decoder.
[0107] In the present embodiment, the above language decoder refers to the language decoder in the large visual language model. The language decoder includes a calculation layer (i.e. a Transformer layer) and a correction layer (i.e. a Softmax layer).
[0108] The specified attention matrix is calculated by the calculation layer to obtain a corresponding reference probability distribution.
[0109] In the embodiment, the calculation of the Transformer layer (such as self-attention, cross-attention) is completed by using the specified attention matrix, and the corresponding reference probability distribution is generated.
[0110] The reference probability distribution is modified by the modification layer to obtain a corresponding modified probability distribution.
[0111] In the embodiment, the reference probability distribution is generated by the Softmax layer to obtain a modified Token probability distribution, i.e., a modified probability distribution, and the modified probability distribution is used as the required second conditional probability distribution.
[0112] The modified probability distribution is used as the second conditional probability distribution.
[0113] The application inputs the specified attention matrix into the language decoder, then calculates the specified attention matrix by the calculation layer to obtain a corresponding reference probability distribution, then modifies the reference probability distribution based on the modification layer to obtain a corresponding modified probability distribution, and then uses the modified probability distribution as the second conditional probability distribution. Based on the above processing procedure, the application inputs the specified attention matrix into the language decoder, then calculates the specified attention matrix by the calculation layer to obtain a reference probability distribution, then modifies the reference probability distribution based on the modification layer, and uses the obtained modified probability distribution as the required second conditional probability distribution, so that the calculation of the specified attention matrix can be efficiently and accurately completed, and the accuracy of the obtained second conditional probability distribution is ensured.
[0114] In some optional implementations, step S206 includes the following steps:
[0115] A preset weighted fusion formula is obtained.
[0116] In the embodiment, the preset weighted fusion formula specifically includes: p final = γ·p'(y k |y <k ) + (1-γ)·p(y k |y <k ), where γ is a hyperparameter greater than 1, used to control the strength of contrast decoding. By setting γ>1, the influence of the probability distribution intervened by attention (i.e., more following user instructions) is amplified.
[0117] The first conditional probability distribution and the second conditional probability distribution are weighted and calculated based on the weighted fusion formula to obtain a corresponding calculation result.
[0118] In the embodiment, the calculation result is taken as the target conditional probability distribution.
[0119] The calculation result is taken as the target conditional probability distribution.
[0120] In the embodiment, the contrast decoding strengthens the constraint of the user instruction through "double path verification". The intervention path corrects the illusion tendency, the reference path retains the original generation ability of the model, and the fusion reduces errors and avoids excessive correction. γ>1 ensures that the corrected distribution dominates, while retaining part of the original information to prevent semantic degradation.
[0121] The application obtains a preset weighted fusion formula; then based on the weighted fusion formula, the first conditional probability distribution and the second conditional probability distribution are weighted and calculated to obtain a corresponding calculation result; and subsequently, the calculation result is taken as the target conditional probability distribution. Based on the above processing procedure, the application performs weighted calculation on the first conditional probability distribution and the second conditional probability distribution based on the use of the weighted fusion formula, and takes the generated calculation result as the required target conditional probability distribution, so as to automatically and accurately complete the fusion processing of the first conditional probability distribution and the second conditional probability distribution, and improve the accuracy and intelligence of the generated target conditional probability distribution.
[0122] In some optional implementation manners of the embodiment, step S207 includes the following steps:
[0123] A preset greedy search strategy is obtained.
[0124] In the embodiment, the strategy flow content of the greedy search strategy includes: in each step of generation, the Token with the highest probability is selected from the target conditional probability distribution as the current output (target Token). In addition, the selected Token is added to the generated sequence and used as the input context of the next step.
[0125] The target conditional probability distribution is processed based on the greedy search strategy to generate a corresponding first Token.
[0126] In the embodiment, the target conditional probability distribution is processed based on the strategy flow content of the greedy search strategy, and the Token with the highest selection probability is taken as the generated first Token, i.e., the target Token.
[0127] The first Token is taken as the target Token.
[0128] In the embodiment, by selecting a decoding strategy (greedy search strategy) and dynamic context updating, the previous quantitative adjustment can be converted into actual results that conform to instructions and have no illusion. The greedy search guarantees efficiency, and the dynamic context ensures global coherence, finally achieving the balance between "correcting errors" and "keeping natural".
[0129] The application obtains a preset greedy search strategy, then processes the target conditional probability distribution based on the greedy search strategy to generate a corresponding first word token, and subsequently takes the first word token as the target word token. Based on the above processing procedure, the application can automatically and efficiently generate a target word token that matches the target conditional probability distribution by processing the target conditional probability distribution based on the use of a greedy search strategy, thereby effectively improving the generation efficiency of the target word token.
[0130] In some optional implementation manners of the embodiment, step S207 includes the following steps:
[0131] A preset beam search strategy is obtained.
[0132] In the embodiment, the strategy procedure content of the beam search strategy includes: setting a beam width K (such as K=3), and retaining K candidate sequences with the highest probabilities at each step. The target conditional probability distribution of each candidate sequence is expanded, and the joint probability (product or logarithmic sum) of the candidate sequence is calculated. Subsequently, the candidate sequence with the highest joint probability is selected as the final output.
[0133] The target conditional probability distribution is processed based on the beam search strategy to generate a corresponding second word token.
[0134] In the embodiment, the target conditional probability distribution is processed based on the strategy procedure content of the beam search strategy, and the generated second word token is taken as the required target word token.
[0135] The second word token is taken as the target word token.
[0136] In the embodiment, by selecting a decoding strategy (beam search strategy) and dynamic context updating, the previous quantitative adjustment can be converted into actual results that conform to instructions and have no illusion. The balance between quality and efficiency, and the dynamic context ensure global coherence, finally achieving the balance between "correcting errors" and "keeping natural".
[0137] The application obtains a preset beam search strategy, then processes the target conditional probability distribution based on the beam search strategy to generate a corresponding second word token, and subsequently takes the second word token as the target word token. Based on the above processing procedure, the application can automatically and efficiently generate a target word token matching the target conditional probability distribution by processing the target conditional probability distribution based on the use of the beam search strategy, effectively ensures the generation efficiency of the target word token, and improves the accuracy of the generated target word token.
[0138] In some optional implementations, the obtained user information seeks user consent and complies with relevant laws and relevant policies.
[0139] In addition, the non-company software tools or components appearing in the embodiments of the application are only examples and do not represent actual use.
[0140] In addition, based on the technical solutions of the application, the following significant advantages and positive effects are brought about:
[0141] 1. Significantly reduce hallucination rate: Experiments show that the application can significantly reduce the object hallucination rate at the instance level and the sentence level on multiple mainstream LVLMs (such as LLaVA-1.5, MiniGPT-4, Shikra), as well as on authoritative hallucination evaluation benchmarks such as CHAIR and POPE, and the performance is better than that of the prior art.
[0142] 2. Zero cost and high efficiency: As an inference period method that does not require training, the application avoids high data labeling and model training costs, and can be directly applied to existing pre-trained models, greatly reducing the technical landing threshold.
[0143] 3. Maintain and improve model general ability: Unlike some hallucination suppression methods that may cause model ability degradation, the application can effectively suppress hallucination while maintaining or even slightly improving the performance of the model on general multi-modal benchmarks.
[0144] 4. Enhance model fine-grained perception ability: By forcing the model to balance attention to visual and textual details, the application can improve the fine-grained perception ability of the model. For example, the model can notice and describe more subtle objects or attributes in the image (such as signs on clothing, states of specific objects, etc.), thereby generating more accurate and richer descriptions.
[0145] High universality and scalability: The application does not depend on a specific model architecture and can be widely applied to various transformer-based LVLMs. Its plug-and-play characteristics make it easy to integrate into existing systems, making it highly practical.
[0146] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0147] It should be emphasized that, to further ensure the privacy and security of the aforementioned target text, the target text can also be stored in a node of a blockchain.
[0148] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0149] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0150] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0151] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0152] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0153] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a data processing device based on artificial intelligence, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0154] like Figure 3 As shown, the artificial intelligence-based data processing device 300 described in this embodiment includes: a receiving module 301, a first processing module 302, a quantization module 303, a second processing module 304, a calculation module 305, a fusion module 306, a third processing module 307, and an output module 308. Wherein:
[0155] The receiving module 301 is used to receive the input target image and the descriptive text corresponding to the target image;
[0156] The first processing module 302 is used to process the target image and the descriptive text based on a preset large-scale visual language model to generate corresponding word units;
[0157] The quantization module 303 is used to quantize the word based on a preset modal bias analysis strategy to obtain the modal bias quantization result of the word.
[0158] The second processing module 304 is used to calculate the original attention matrix of the word if the modality bias quantification result meets the preset attention intervention conditions, and to perform weight adjustment processing on the original attention matrix based on the modality bias quantification result to obtain the corresponding specified attention matrix.
[0159] The calculation module 305 is used to calculate and process the original attention matrix to obtain the corresponding first conditional probability distribution, and to calculate and process the specified attention matrix to obtain the corresponding second conditional probability distribution.
[0160] The fusion module 306 is used to fuse the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding target conditional probability distribution;
[0161] The third processing module 307 is used to decode the target conditional probability distribution to generate corresponding target words, and generate target text corresponding to the target image based on the target words;
[0162] The output module 308 is used to output the target text.
[0163] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here.
[0164] In some optional implementations of this embodiment, the quantization module 303 includes:
[0165] The first calculation submodule is used to calculate the text attention ratio of the specified word element based on a preset first calculation strategy when generating the specified word element; wherein the specified word element is any one of all the word elements;
[0166] The second calculation submodule is used to calculate the visual attention ratio of the specified word based on a preset second calculation strategy;
[0167] The first acquisition submodule is used to acquire the preset bias threshold;
[0168] The first analysis submodule is used to perform numerical analysis on the text attention ratio and the visual attention ratio based on the bias threshold, and generate the modal bias quantification result of the specified word.
[0169] In some optional implementations of this embodiment, the second processing module 304 includes:
[0170] The second analysis submodule is used to analyze the large visual language model and determine the intervention level in the large visual language model.
[0171] The submodule is invoked to call the corresponding target attention intervention formula based on the modality bias quantification result.
[0172] The adjustment submodule is used to adjust the attention weights of the original attention matrix at the intervention level based on the target attention intervention formula, so as to obtain the adjusted attention matrix.
[0173] The first determining submodule is used to use the adjusted attention matrix as the specified attention matrix.
[0174] In some optional implementations of this embodiment, the large visual language model includes a language decoder, which includes a preset computation layer and a correction layer; the computation module 305 includes:
[0175] An input submodule is used to input the specified attention matrix into the language decoder;
[0176] The third calculation submodule is used to calculate and process the specified attention matrix through the calculation layer to obtain the corresponding baseline probability distribution;
[0177] The correction submodule is used to correct the baseline probability distribution based on the correction layer to obtain the corresponding corrected probability distribution.
[0178] The second determining submodule is used to use the modified probability distribution as the second conditional probability distribution.
[0179] In some optional implementations of this embodiment, the fusion module 306 includes:
[0180] The second acquisition submodule is used to acquire the preset weighted fusion formula;
[0181] The fourth calculation submodule is used to perform weighted calculation on the first conditional probability distribution and the second conditional probability distribution based on the weighted fusion formula to obtain the corresponding calculation result;
[0182] The third determining submodule is used to use the calculation result as the target conditional probability distribution.
[0183] In some optional implementations of this embodiment, the third processing module 307 includes:
[0184] The third acquisition submodule is used to acquire the preset greedy search strategy;
[0185] The first processing submodule is used to process the target conditional probability distribution based on the greedy search strategy to generate the corresponding first word element;
[0186] The fourth determining submodule is used to use the first word element as the target word element.
[0187] In some optional implementations of this embodiment, the third processing module 307:
[0188] The fourth acquisition submodule is used to acquire the preset beam search strategy;
[0189] The second processing submodule is used to process the target conditional probability distribution based on the beam search strategy to generate the corresponding second word element;
[0190] The fifth determining submodule is used to use the second lexical as the target lexical.
[0191] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0192] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0193] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0194] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for data processing methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0195] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the artificial intelligence-based data processing method.
[0196] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0197] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based data processing method described above.
[0198] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0199] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.< / bos>
Claims
1. An artificial intelligence-based data processing method, characterized by, The method comprises the following steps: receiving an input target image and a description text corresponding to the target image; processing the target image and the description text based on a preset large visual language model to generate corresponding word units; quantitatively processing the word units based on a preset modal bias analysis strategy to obtain a modal bias quantitative result of the word units; if the modal bias quantitative result meets a preset attention intervention condition, calculating an original attention matrix of the word units and performing weight adjustment processing on the original attention matrix based on the modal bias quantitative result to obtain a corresponding specified attention matrix; calculating and processing the original attention matrix to obtain a corresponding first conditional probability distribution and calculating and processing the specified attention matrix to obtain a corresponding second conditional probability distribution; fusing the first conditional probability distribution and the second conditional probability distribution to obtain a target conditional probability distribution; decoding the target conditional probability distribution to generate a target word unit corresponding to the target image, and generating a target text corresponding to the target image based on the target word unit; outputting the target text. 2.The artificial intelligence-based data processing method of claim 1, wherein, The step of quantitatively processing the word units based on a preset modal bias analysis strategy to obtain a modal bias quantitative result of the word units comprises: when generating a specified word unit, calculating a text attention ratio of the specified word unit based on a preset first calculation strategy; wherein the specified word unit is any one of all the word units; calculating a visual attention ratio of the specified word unit based on a preset second calculation strategy; obtaining a preset bias threshold value; based on the bias threshold value, numerically analyzing the text attention ratio and the visual attention ratio to generate a modal bias quantitative result of the specified word unit. 3.The artificial intelligence-based data processing method of claim 1, wherein, The step of performing weight adjustment processing on the original attention matrix based on the modal bias quantitative result to obtain a corresponding specified attention matrix comprises: analyzing the large visual language model to determine an intervention level in the large visual language model; calling a target attention intervention formula corresponding to the modal bias quantitative result; based on the target attention intervention formula, performing attention weight adjustment processing on the original attention matrix in the intervention level to obtain an adjusted attention matrix; taking the adjusted attention matrix as the specified attention matrix. 4.The artificial intelligence-based data processing method of claim 1, wherein, The large visual language model comprises a language decoder comprising a preset calculation layer and a correction layer; the step of calculating and processing the specified attention matrix to obtain a corresponding second conditional probability distribution comprises: inputting the specified attention matrix into the language decoder; calculating and processing the specified attention matrix through the calculation layer to obtain a corresponding baseline probability distribution; based on the correction layer, correcting the baseline probability distribution to obtain a corresponding corrected probability distribution; taking the corrected probability distribution as the second conditional probability distribution. 5.The artificial intelligence-based data processing method of claim 1, wherein, The step of fusing the first conditional probability distribution and the second conditional probability distribution to obtain a corresponding target conditional probability distribution specifically includes: obtaining a preset weighted fusion formula; performing weighted calculation processing on the first conditional probability distribution and the second conditional probability distribution based on the weighted fusion formula to obtain a corresponding calculation result; taking the calculation result as the target conditional probability distribution. 6.The artificial intelligence-based data processing method of claim 1, wherein, The step of decoding the target conditional probability distribution to generate a corresponding target word element specifically includes: obtaining a preset greedy search strategy; processing the target conditional probability distribution based on the greedy search strategy to generate a corresponding first word element; taking the first word element as the target word element. 7.The artificial intelligence-based data processing method of claim 1, wherein, The step of decoding the target conditional probability distribution to generate a corresponding target word element specifically includes: obtaining a preset beam search strategy; processing the target conditional probability distribution based on the beam search strategy to generate a corresponding second word element; taking the second word element as the target word element.
8. An artificial intelligence-based data processing apparatus, characterized by comprising: It includes: a receiving module configured to receive an input target image and a description text corresponding to the target image; a first processing module configured to process the target image and the description text based on a preset large visual language model to generate a word element; a quantization module configured to perform quantization processing on the word element based on a preset modality bias analysis strategy to obtain a modality bias quantization result of the word element; a second processing module configured to calculate an original attention matrix of the word element if the modality bias quantization result meets a preset attention intervention condition, and perform weight adjustment processing on the original attention matrix based on the modality bias quantization result to obtain a corresponding specified attention matrix; a calculation module configured to perform calculation processing on the original attention matrix to obtain a first conditional probability distribution, and perform calculation processing on the specified attention matrix to obtain a second conditional probability distribution; a fusion module configured to fuse the first conditional probability distribution and the second conditional probability distribution to obtain a corresponding target conditional probability distribution; a third processing module configured to decode the target conditional probability distribution to generate a corresponding target word element, and generate a target text corresponding to the target image based on the target word element; an output module configured to perform output processing on the target text.
9. A computer device, comprising: It includes a memory and a processor, the memory stores computer readable instructions, and the processor executes the computer readable instructions to realize the steps of the artificial intelligence-based data processing method in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by the processor to realize the steps of the artificial intelligence-based data processing method in any one of claims 1 to 7.
Citation Information
Patent Citations
Visual localization and anaphora segmentation method, system and device based on mask anaphora modeling and storage medium
CN118734091A
Elastic converter service system self-adaptive through lexical elements
CN119848173A
Retrieval method, device and equipment based on attention guidance and medium
CN120179878A
Methods and systems for fast inference from machine learning models
WO2024118603A1