Data processing methods, devices, computer equipment, and media based on artificial intelligence
Patent Information
- Application Number
- CN202511299299.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-09-10
AI Technical Summary
[0006]本申请实施例的目的在于提出一种基于人工智能的数据处理方法、装置、计算机设备及存储介质,以解决现有的大型视觉语言模型的对象幻觉问题,导致图像文本生成的准确性较低的技术问题
[0027]上述基于人工智能的数据处理方法、装置、计算机设备及存储介质所实现的方案中,首先接收输入的目标图像,以及与所述目标图像对应的描述文本;然后基于预设的大型视觉语言模型对所述目标图像与所述描述文本进行处理,生成对应的词元;之后基于预设的模态偏见分析策略对所述词元进行量化处理,得到所述词元的模态偏见量化结果;若所述模态偏见量化结果符合预设的注意力干预条件,则计算所述词元的原始注意力矩阵,并基于所述模态偏见量化结果对所述原始注意力矩阵进行权重调整处理,得到对应的指定注意力矩阵;后续对所述原始注意力矩阵进行计算处理得到对应的第一条件概率分布,以及对所述指定注意力矩阵进行计算处理得到对应的第二条件概率分布;并对所述第一条件概率分布与所述第二条件概率分布进行融合处理,得到对应的目标条件概率分布;进一步对所述目标条件概率分布进行解码处理以生成对应的目标词元,并基于所述目标词元生成与所述目标图像对应的目标文本;最后对所述目标文本进行输出处理。基于以上的自动化处理流程,本申请通过基于大型视觉语言模型的使用对目标图像与描述文本进行处理,生成对应的词元,然后基于模态偏见分析策略的使用对词元进行量化处理,若模态偏见量化结果符合预设的注意力干预条件,则计算词元的原始注意力矩阵,并基于模态偏见量化结果对原始注意力矩阵进行权重调整处理得到指定注意力矩阵,之后对原始注意力矩阵进行计算处理得到第一条件概率分布,以及对指定注意力矩阵进行计算处理得到第二条件概率分布,进而对第一条件概率分布与第二条件概率分布进行融合处理得到目标条件概率分布,后续对目标条件概率分布进行解码处理以生成目标词元,并基于目标词元生成与目标图像对应的目标文本,最后对目标文本进行输出处理。如此,本申请通过基于大型视觉语言模型、模态偏见分析策略、注意力干预机制的结合使用,可以有效减少幻觉的产生,进而提升生成的目标文本的准确性和流畅性。
Smart Images

Figure CN121233745B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as fintech and digital healthcare, particularly to data processing methods, devices, computer equipment, and storage media based on artificial intelligence. Background Technology
[0002] In the field of large visual language models (LVLMs), LVLMs have demonstrated powerful capabilities in multimodal understanding, reasoning, and human-computer interaction in recent years. These models typically consist of a visual encoder and a large language model (LLM), capable of handling text-image interleaved inputs, and show great application potential in numerous fields such as finance and healthcare.
[0003] However, despite significant progress in LVLMs, a serious object illusion problem persists. This manifests as the model generating text descriptions that include objects that do not actually exist in the image, or incorrectly determining the presence of a specific object in the image. This illusion problem severely limits the application and deployment of LVLMs in real-world scenarios where accuracy is paramount, resulting in low accuracy of text data generated from images.
[0004] For example, in the financial sector, LVLMs (Visual Literacy Models) might be used to identify financial data in images and generate text descriptions. However, due to the object illusion problem, the model might include financial items that are not present in the image (such as non-existent expense categories) in the text description, or incorrectly determine the existence of a key financial indicator (such as profit growth rate). Such erroneous descriptions can lead to inaccurate financial analysis results, affecting financial institutions' decision-making and reducing the efficiency and accuracy of business processing.
[0005] Therefore, there is an urgent need to provide a method or system that can effectively solve the object illusion problem of LVLMs, so as to improve the accuracy of image text generation and meet the practical application needs of multiple fields. Summary of the Invention
[0006] The purpose of this application is to propose a data processing method, apparatus, computer device, and storage medium based on artificial intelligence to solve the problem of object illusion in existing large-scale visual language models, which leads to low accuracy in image text generation.
[0007] Firstly, an artificial intelligence-based data processing method is provided, including:
[0008] Receive the input target image and the corresponding descriptive text;
[0009] The target image and the descriptive text are processed based on a pre-set large-scale visual language model to generate corresponding lexical units;
[0010] The word units are quantified based on a preset modal bias analysis strategy to obtain the modal bias quantification results of the word units;
[0011] If the modal bias quantification result meets the preset attention intervention conditions, then the original attention matrix of the word is calculated, and the weight adjustment processing of the original attention matrix is performed based on the modal bias quantification result to obtain the corresponding specified attention matrix;
[0012] The original attention matrix is processed to obtain the corresponding first conditional probability distribution, and the specified attention matrix is processed to obtain the corresponding second conditional probability distribution;
[0013] The first conditional probability distribution and the second conditional probability distribution are fused to obtain the corresponding target conditional probability distribution.
[0014] The target conditional probability distribution is decoded to generate corresponding target words, and target text corresponding to the target image is generated based on the target words.
[0015] The target text is then processed for output.
[0016] Secondly, an artificial intelligence-based data processing device is provided, comprising:
[0017] A receiving module is used to receive an input target image and descriptive text corresponding to the target image;
[0018] The first processing module is used to process the target image and the descriptive text based on a preset large-scale visual language model to generate corresponding word units;
[0019] The quantization module is used to quantize the word units based on a preset modal bias analysis strategy to obtain the modal bias quantization result of the word units;
[0020] The second processing module is used to calculate the original attention matrix of the word if the modality bias quantification result meets the preset attention intervention conditions, and to perform weight adjustment processing on the original attention matrix based on the modality bias quantification result to obtain the corresponding specified attention matrix.
[0021] The calculation module is used to calculate and process the original attention matrix to obtain the corresponding first conditional probability distribution, and to calculate and process the specified attention matrix to obtain the corresponding second conditional probability distribution;
[0022] The fusion module is used to fuse the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding target conditional probability distribution.
[0023] The third processing module is used to decode the target conditional probability distribution to generate corresponding target words, and generate target text corresponding to the target image based on the target words;
[0024] The output module is used to process the target text for output.
[0025] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based data processing method.
[0026] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned artificial intelligence-based data processing method.
[0027] In the above-mentioned scheme implemented by the data processing method, device, computer equipment, and storage medium based on artificial intelligence, the following steps are taken: First, an input target image and a descriptive text corresponding to the target image are received; then, the target image and the descriptive text are processed based on a preset large-scale visual language model to generate corresponding lexical units; next, the lexical units are quantized based on a preset modal bias analysis strategy to obtain the modal bias quantization result of the lexical units; if the modal bias quantization result meets the preset attention intervention conditions, the original attention matrix of the lexical units is calculated, and the weight of the original attention matrix is adjusted based on the modal bias quantization result to obtain the corresponding specified attention matrix; subsequently, the original attention matrix is calculated to obtain the corresponding first conditional probability distribution, and the specified attention matrix is calculated to obtain the corresponding second conditional probability distribution; the first conditional probability distribution and the second conditional probability distribution are fused to obtain the corresponding target conditional probability distribution; further, the target conditional probability distribution is decoded to generate the corresponding target lexical units, and target text corresponding to the target image is generated based on the target lexical units; finally, the target text is output. Based on the above automated processing flow, this application processes the target image and descriptive text using a large-scale visual language model to generate corresponding lexical units. Then, it quantizes the lexical units using a modal bias analysis strategy. If the modal bias quantization result meets preset attention intervention conditions, the original attention matrix of the lexical units is calculated. Based on the modal bias quantization result, the original attention matrix is weighted to obtain a specified attention matrix. Subsequently, the original attention matrix is processed to obtain a first conditional probability distribution, and the specified attention matrix is processed to obtain a second conditional probability distribution. The first and second conditional probability distributions are then fused to obtain a target conditional probability distribution. This target conditional probability distribution is then decoded to generate target lexical units, and target text corresponding to the target image is generated based on these target lexical units. Finally, the target text is output. Thus, by combining a large-scale visual language model, a modal bias analysis strategy, and an attention intervention mechanism, this application can effectively reduce the generation of illusions, thereby improving the accuracy and fluency of the generated target text. Attached Figure Description
[0028] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1This is an exemplary system architecture diagram to which this application can be applied;
[0030] Figure 2 This is a flowchart of an embodiment of the artificial intelligence-based data processing method according to this application;
[0031] Figure 3 This is a schematic diagram of a structure of an embodiment of the artificial intelligence-based data processing apparatus according to this application;
[0032] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0034] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0036] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0037] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0038] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0039] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0040] It should be noted that the artificial intelligence-based data processing method provided in the embodiments of this application is generally executed by a server / terminal device, and correspondingly, the artificial intelligence-based data processing device is generally set in the server / terminal device.
[0041] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0042] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the AI-based data processing method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The AI-based data processing method provided in this application can be applied to any scenario requiring image analysis, and thus can be applied to products in these scenarios, such as image analysis products in the financial insurance field. The AI-based data processing method includes the following steps:
[0043] Step S201: Receive the input target image and the descriptive text corresponding to the target image.
[0044] In this embodiment, the data processing method based on artificial intelligence runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire the input target image and corresponding descriptive text via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wide wireless) connections, and other currently known or future known wireless connection methods. The executing entity of this application is specifically a data processing system, which may be simply referred to as the system.
[0045] The target image mentioned above is the image to be processed (i.e., the image content needs to be described), and the description text mentioned above is the user instruction input by the relevant user. This application can be applied to image analysis scenarios in the fields of fintech and healthcare. For example, in the car insurance scenario in the financial insurance field, the content of the input target image can be: a picture showing a car involved in a collision, with the front of the car severely deformed, some car parts scattered around, a small tree that has been knocked over and leaning, and obvious skid marks and oil stains on the ground. The corresponding description text includes: "Please describe in detail the objects related to the vehicle accident in this picture and the surrounding environment caused by the accident, so as to provide accurate information for car insurance claims." Or, in the home property insurance scenario in the financial insurance field, the content of the input target image can be: an indoor scene picture, showing a room that has been damaged by fire. The walls are blackened, some of the paint has peeled off, furniture such as sofas and wardrobes are burned beyond recognition, the floor is full of charred debris and water, the window glass is broken, and the curtains have also been burned. The corresponding descriptive text includes: "Describe the damage to household property and belongings caused by the fire, as well as the overall damage to the room, for the purpose of assessing losses under home property insurance, based on this image."
[0046] In addition, in the disease analysis assistance scenario within the digital healthcare field, the input target image could be: a medical X-ray showing a patient's lung area. A distinct shadow area can be seen in the lungs, with an irregular shape, blurred edges, and somewhat disordered surrounding lung texture, showing a clear difference from a normal lung X-ray image. The corresponding descriptive text includes: "Please describe the location, shape, size, and other characteristics of the abnormal shadow in this lung X-ray to provide imaging evidence for disease diagnosis." Alternatively, in the rehabilitation monitoring scenario within the digital healthcare field, the input target image could be: a photograph of a patient performing rehabilitation training, using rehabilitation equipment to perform leg flexion and extension exercises. The photograph clearly shows the patient's leg posture, the type of rehabilitation equipment (such as elastic bands, rehabilitation bicycles, etc.), and the patient's facial expression due to exertion. The corresponding descriptive text includes: "Describe the equipment used by the patient during rehabilitation training in this image, their body movements and posture, and their facial expression to monitor the rehabilitation treatment effect and adjust the rehabilitation plan."
[0047] Specifically, this application proposes a novel, training-free intervention method for the reasoning stage, the core innovation of which lies in:
[0048] This application, based on in-depth empirical analysis, is the first to explicitly identify two distinct modes of illusion in LVLMs—"generative illusion" and "discriminative illusion"—and discovers that their root cause lies in a previously overlooked phenomenon of "modal bias." Specifically, in generating illusions, the model's attention is overly focused on visual information; while in generating discriminative illusions, the model relies excessively on textual information. This discovery provides a completely new perspective and theoretical foundation for solving the illusion problem.
[0049] A bidirectional attention intervention mechanism is proposed: Unlike existing technologies that only unidirectionally enhance vision or suppress language priors, this invention proposes a bidirectional, dynamic "text and visual attention intervention" mechanism. This mechanism can intelligently adjust the attention weights assigned to text and visual tokens during inference, forcing the model to process information from both modalities in a balanced manner, thereby eliminating unimodal bias and fundamentally reducing the generation of illusions.
[0050] Collaborative Contrastive Decoding Strategy: This application combines attention intervention with a contrastive decoding strategy. By comparing the output probability distribution after attention intervention with the original probability distribution, the influence of user commands (visual and textual) is further enhanced, while reducing the model's over-reliance on its own inherent, potentially erroneous, parameterized knowledge, thus synergistically improving the illusion suppression effect.
[0051] Achieving zero-cost, highly versatile deployment: This application presents a completely training-free method that only intervenes in the calculation of the internal attention matrix during the model inference stage, without involving any modification to model parameters. Therefore, there are no additional training costs or data annotation requirements. Furthermore, this method is highly versatile, serving as a plug-and-play module applicable to all LVLMs based on the Transformer architecture, and compatible with various decoding strategies.
[0052] Step S202: Process the target image and the descriptive text based on a preset large-scale visual language model to generate corresponding lexical units.
[0053] In this embodiment, the large-scale visual language model mentioned above is a model that includes a visual encoder, a projector, and a language decoder. The process of processing the target image and descriptive text based on the large-scale visual language model to generate lexical units includes: (1) Encoding stage: Text encoding: The user's descriptive text is encoded into a contextual representation of text lexical units by a text encoder (such as the Transformer layer of BERT or GPT). Visual encoding: The image is encoded into a feature representation of visual lexical units (such as [v_1, v_2, ..., v_N], where each v_i corresponds to a region of the image) by a visual encoder (such as ViT or ResNet+Transformer). (2) Decoding stage (generating lexical units). Initial state: The decoder receives the encoded text and visual features, as well as a special starting lexical unit ( <bos>). Generate token by token: Query token: the current token to be generated (for example, when generating the t-th token, the query is q_t). Attention calculation: Text self-attention: q_t calculates attention with the generated text history (h_1,h_2,...,h_{t-1}) to capture language context. Visual cross-modal attention: q_t calculates attention with visual tokens (v_1,v_2,...,v_N) to capture image details. Weight fusion: perform weighted summation on the text and visual attention weights to obtain the context representation c_t of the current token. Token prediction: predict the probability distribution of the next token through c_t (e.g., P(w_t|w_{<t},I,T), where I is an image and T is a user instruction).
[0054] Step S203, perform quantization processing on the token based on a preset modal bias analysis strategy to obtain a modal bias quantization result of the token.
[0055] In this embodiment, the specific implementation process of performing quantization processing on the token based on the preset modal bias analysis strategy to obtain the modal bias quantization result of the token will be further described in detail in the subsequent specific embodiments of the present application, and will not be elaborated herein.
[0056] Step S204, if the modal bias quantization result meets a preset attention intervention condition, calculate an original attention matrix of the token, and perform weight adjustment processing on the original attention matrix based on the modal bias quantization result to obtain a corresponding specified attention matrix.
[0057] In this embodiment, that the modal bias quantization result meets the preset attention intervention condition means that the modal bias quantization result is a result corresponding to detecting a generative hallucination tendency (e.g., VAR is too high, such as VAR > 0.7), or a result corresponding to detecting a discriminative hallucination tendency (e.g., TAR is too high, such as TAR > 0.6). Wherein, based on the large vision-language model, a general attention calculation method can be used to perform calculation processing on the token to obtain the corresponding original attention matrix.
[0058] Wherein, the specific implementation process of performing weight adjustment processing on the original attention matrix based on the modal bias quantization result to obtain the corresponding specified attention matrix will be further described in detail in the subsequent specific embodiments of the present application, and will not be elaborated herein.
[0059] Step S205, perform calculation processing on the original attention matrix to obtain a corresponding first conditional probability distribution, and perform calculation processing on the specified attention matrix to obtain a corresponding second conditional probability distribution.
[0060] In this embodiment, to further amplify the effect of attention intervention and reduce the model's reliance on internal erroneous knowledge, contrastive decoding is introduced. Parallel computation of probability distribution: When generating each token, the system computes two paths in parallel: Path 1 (after intervention): using the modified attention matrix A′ l,h After completing the subsequent Transformer calculations, a modified Token conditional probability distribution is obtained, namely the second conditional probability distribution p′(y). k |y <k ), Path 2 (Original): Using the unmodified original attention matrix A l,h Calculate the original token conditional probability distribution, i.e., the second conditional probability distribution p(y). k |y <k ).
[0061] The specific implementation process of calculating and processing the specified attention matrix to obtain the corresponding second conditional probability distribution will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0062] Step S206: The first conditional probability distribution and the second conditional probability distribution are fused to obtain the corresponding target conditional probability distribution.
[0063] In this embodiment, the specific implementation process of fusing the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding target conditional probability distribution will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0064] Step S207: Decode the target conditional probability distribution to generate corresponding target words, and generate target text corresponding to the target image based on the target words.
[0065] In this embodiment, the specific implementation process of decoding the target conditional probability distribution to generate the corresponding target lexical will be further described in detail in subsequent specific embodiments, and will not be elaborated on here. Specifically, the step of generating the corresponding target lexical is repeated for each lexical until a complete sequence is generated, which serves as the target text.
[0066] Step S208: Output the target text.
[0067] In this embodiment, the output processing of the target text can be completed by sending the generated target text to the relevant user.
[0068] This application first receives an input target image and a corresponding descriptive text. Then, it processes the target image and the descriptive text based on a preset large-scale visual language model to generate corresponding lexical units. Next, it quantizes the lexical units based on a preset modal bias analysis strategy to obtain a modal bias quantization result. If the modal bias quantization result meets preset attention intervention conditions, it calculates the original attention matrix of the lexical units and performs weight adjustment processing on the original attention matrix based on the modal bias quantization result to obtain a corresponding specified attention matrix. Subsequently, it calculates and processes the original attention matrix to obtain a corresponding first conditional probability distribution, and calculates and processes the specified attention matrix to obtain a corresponding second conditional probability distribution. It then fuses the first and second conditional probability distributions to obtain a corresponding target conditional probability distribution. Furthermore, it decodes the target conditional probability distribution to generate corresponding target lexical units, and generates target text corresponding to the target image based on the target lexical units. Finally, it outputs the target text. Based on the above automated processing flow, this application processes the target image and descriptive text using a large-scale visual language model to generate corresponding lexical units. Then, it quantizes the lexical units using a modal bias analysis strategy. If the modal bias quantization result meets preset attention intervention conditions, the original attention matrix of the lexical units is calculated. Based on the modal bias quantization result, the original attention matrix is weighted to obtain a specified attention matrix. Subsequently, the original attention matrix is processed to obtain a first conditional probability distribution, and the specified attention matrix is processed to obtain a second conditional probability distribution. The first and second conditional probability distributions are then fused to obtain a target conditional probability distribution. This target conditional probability distribution is then decoded to generate target lexical units, and target text corresponding to the target image is generated based on these target lexical units. Finally, the target text is output. Thus, by combining a large-scale visual language model, a modal bias analysis strategy, and an attention intervention mechanism, this application can effectively reduce the generation of illusions, thereby improving the accuracy and fluency of the generated target text.
[0069] In some alternative implementations, step S203 includes the following steps:
[0070] When generating a specified word, the text attention ratio of the specified word is calculated based on a preset first calculation strategy; wherein, the specified word is any one of all the word.
[0071] In this embodiment, for precise intervention, it is necessary to quantify the degree of attention the model pays to different modalities when generating a specific token. This is achieved by defining two key metrics:
[0072] Textual Attention Ratio (TAR): When generating the k-th token, the model performs a certain number of attention operations on all input text tokens X. T The sum of attention weights.
[0073]
[0074] Among them, A l,h (k,i) represents the attention score of the k-th token being generated for the i-th text token in the h-th attention head of the l-th layer.
[0075] Visual Attention Ratio (VAR): Similarly, it represents the ratio of the input image TokenX to the given image. V The sum of attention weights.
[0076]
[0077] Among them, A l,h (k,j) represents the attention score between the k-th token and the j-th visual token being generated in the h-th attention head of the l-th layer.
[0078] Specifically, the first calculation strategy mentioned above includes the following: when generating the k-th Token, traverse all Transformer layers (l) and attention heads (h), and accumulate the Token for all text Tokens X. T Attention score.
[0079] The visual attention ratio of the specified word is calculated based on a preset second calculation strategy.
[0080] In this embodiment, the second calculation strategy includes: when generating the k-th Token, traversing all Transformer layers (l) and attention heads (h), and accumulating the Token for all image Tokens. V Attention score.
[0081] Obtain the preset bias threshold.
[0082] In this embodiment, the existence of modal bias was confirmed in advance through the analysis of a large number of hallucination samples: generative hallucination tokens usually exhibit high VAR and low TAR, while discriminative hallucination tokens exhibit high TAR and low VAR.
[0083] The definition of the bias threshold includes: generative hallucination tendency: condition: VAR > 0.7 and TAR < 0.3. Explanation: The model over-relies on visual information, generating objects that do not exist in the image (such as "red car").
[0084] Discriminative hallucination tendency: Condition: TAR > 0.6 and VAR < 0.4. Explanation: The model relies too heavily on textual information and ignores visual details (e.g., describing a "still cat" as a "flying cat").
[0085] Numerical analysis is performed on the text attention ratio and the visual attention ratio based on the bias threshold to generate the modal bias quantification result for the specified word.
[0086] In this embodiment, the text attention ratio and the visual attention ratio are numerically compared using the aforementioned bias threshold. If a token has a VAR > 0.7 and a TAR < 0.3, it is marked as having a generative illusion tendency (such as fabricating a "red car"); if a TAR > 0.6 and a VAR < 0.4, it is marked as having a discriminative illusion tendency (such as incorrectly judging the existence of an object), thereby completing the generation process of the modal bias quantification result of the token.
[0087] This application calculates the text attention ratio of a specified lexical unit based on a preset first calculation strategy when generating the lexical unit; wherein the specified lexical unit is any one of all lexical units; then, it calculates the visual attention ratio of the specified lexical unit based on a preset second calculation strategy; subsequently, it obtains a preset bias threshold; and then performs numerical analysis on the text attention ratio and the visual attention ratio based on the bias threshold to generate a modal bias quantification result for the specified lexical unit. Based on the above processing flow, this application, when generating a specified lexical unit, calculates the text attention ratio of the specified lexical unit using the first calculation strategy and the visual attention ratio of the specified lexical unit using the second calculation strategy, and then performs numerical analysis on the text attention ratio and visual attention ratio based on the bias threshold. This enables efficient and accurate quantification of the specified lexical unit and generates a modal bias quantification result for the specified lexical unit, ensuring the accuracy of the obtained modal bias quantification result.
[0088] In some optional implementations of this embodiment, step S204 includes the following steps:
[0089] The large visual language model is analyzed to determine the intervention levels within it.
[0090] In this embodiment, the goal of intervention layer selection is to determine the Transformer layer that needs intervention, avoid disrupting the effective feature fusion of shallow layers, and repair modal bias in deep layers. The specific determination process includes: (1) Analyzing attention entropy values. Method: Calculate the distribution variance (or entropy value) of attention weights in each layer to quantify the degree of attention dispersion. For each layer l and head h, calculate the variance of the attention scores of all token pairs. High variance indicates attention dispersion (such as multimodal feature fusion in shallow layers), and low variance indicates attention concentration (such as the "sinking" phenomenon in deep layers). Observation results: Shallow layers (first 6 layers): attention dispersion, involving the initial interaction of multimodal information. Deep layers (last 3 layers): multiple heads exhibit "attention sinking" (i.e., repeatedly focusing on a few tokens of the same modality). (2) Determining the intervention layer. Selection criteria: Attention sinking in deep layers (last 3 layers) directly leads to the solidification of modal bias and requires intervention. Dispersed attention in shallow layers (first 6 layers) helps feature fusion, and intervention may disrupt effective information transmission. Conclusion: Attention weights are adjusted only for the last three layers, preserving the natural attention distribution in the shallow layers.
[0091] The corresponding target attention intervention formula is invoked based on the modality bias quantification results.
[0092] In this embodiment, the target attention intervention formula can be either a text attention intervention formula or a visual attention intervention formula. Specifically, if the modal bias quantification result corresponds to the detection of a generative hallucination tendency (e.g., excessively high VAR, such as VAR > 0.7), then the text attention intervention formula is used. If the modal bias quantification result corresponds to the detection of a discriminative hallucination tendency (e.g., excessively high TAR, such as TAR > 0.6), then the visual attention intervention formula is used.
[0093] Based on the target attention intervention formula, the original attention matrix is adjusted in the intervention level to obtain the adjusted attention matrix.
[0094] In this embodiment, enhancement is achieved by referring to the attention scores of the text and visual tokens according to the target attention intervention formula:
[0095] (1) The text attention intervention formula includes:
[0096]
[0097] Here, α is a hyperparameter used to control the strength of text attention enhancement. This operation aims to address the problem of ignoring textual information in generative illusions.
[0098] (2) The formula for visual attention intervention includes:
[0099]
[0100] Here, β is a hyperparameter used to control the intensity of visual attention enhancement. This operation aims to address the problem of ignoring visual information in discriminative hallucinations.
[0101] By enhancing the weights along their original attention direction, the model can be guided to more reliably follow user commands, rather than introducing noise. The values of α and β can be preset and adjusted according to the specific model and task.
[0102] Specifically, based on the above-mentioned target attention intervention formula, the original attention matrix is adjusted in the above-mentioned intervention level to obtain the adjusted attention matrix as the corresponding designated attention matrix.
[0103] The adjusted attention matrix is used as the designated attention matrix.
[0104] This application analyzes a large visual language model to determine the intervention levels within the model. Then, based on the modality bias quantification results, it calls the corresponding target attention intervention formula. Following this, based on the target attention intervention formula, it adjusts the attention weights of the original attention matrix at each intervention level to obtain an adjusted attention matrix. This adjusted attention matrix is then used as the designated attention matrix. Based on this process, this application analyzes a large visual language model to determine the intervention levels, calls the corresponding target attention intervention formula based on the modality bias quantification results, and then adjusts the attention weights of the original attention matrix at each intervention level using the target attention intervention formula. The resulting adjusted attention matrix is then used as the designated attention matrix. This allows for automatic and accurate weight adjustment of the original attention matrix, improving the accuracy of the generated designated attention matrix.
[0105] In some optional implementations, the large visual language model includes a language decoder, which includes a preset computation layer and a correction layer; step S205 includes the following steps:
[0106] The specified attention matrix is input into the language decoder.
[0107] In this embodiment, the language decoder refers to the language decoder within the aforementioned large-scale visual language model. This language decoder includes a computational layer (i.e., a Transformer layer) and a correction layer (i.e., a Softmax layer).
[0108] The specified attention matrix is processed by the computation layer to obtain the corresponding baseline probability distribution.
[0109] In this embodiment, the calculation of the Transformer layer (such as self-attention and cross-attention) is completed by using the specified attention matrix mentioned above, and the corresponding baseline probability distribution is generated.
[0110] The baseline probability distribution is corrected based on the correction layer to obtain the corresponding corrected probability distribution.
[0111] In this embodiment, the above-mentioned baseline probability distribution is used to generate a modified token probability distribution, i.e., a modified probability distribution, through a Softmax layer, and this modified probability distribution is used as the required second conditional probability distribution.
[0112] The modified probability distribution is used as the second conditional probability distribution.
[0113] This application inputs a specified attention matrix into the language decoder; then, the computation layer processes the specified attention matrix to obtain a corresponding baseline probability distribution; subsequently, the correction layer corrects the baseline probability distribution to obtain a corresponding corrected probability distribution; finally, the corrected probability distribution is used as the second conditional probability distribution. Based on the above processing flow, this application achieves efficient and accurate computation of the specified attention matrix by inputting the specified attention matrix into the language decoder, then processing the specified attention matrix using the computation layer to obtain a baseline probability distribution, further correcting the baseline probability distribution using the correction layer, and using the obtained corrected probability distribution as the required second conditional probability distribution. This ensures the accuracy of the obtained second conditional probability distribution.
[0114] In some alternative implementations, step S206 includes the following steps:
[0115] Obtain the preset weighted fusion formula.
[0116] In this embodiment, the aforementioned preset weighted fusion formula specifically includes: p final =γ·p'(y k |y <k )+(1-γ)·p(y k |y <k ), where γ is a hyperparameter greater than 1 used to control the intensity of contrastive decoding. By setting γ>1, the influence of the probability distribution after attentional intervention (i.e., more compliant with user instructions) is amplified.
[0117] The first conditional probability distribution and the second conditional probability distribution are weighted and calculated based on the weighted fusion formula to obtain the corresponding calculation results.
[0118] In this embodiment, the first conditional probability distribution and the second conditional probability distribution are input into the corresponding positions in the weighted fusion formula for calculation, and the calculation result is used as the final output probability distribution, i.e., the target conditional probability distribution.
[0119] The calculation result is used as the target conditional probability distribution.
[0120] In this embodiment, the contrastive decoding strengthens the binding force of user commands through "dual-path verification." The intervention path corrects the illusionary tendency, while the baseline path retains the model's original generative capabilities. After fusion, errors are reduced while over-correction is avoided. γ>1 ensures that the corrected distribution is dominant, while retaining some original information to prevent semantic degradation.
[0121] This application obtains a preset weighted fusion formula; then, based on the weighted fusion formula, it performs weighted calculations on the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding calculation result; subsequently, the calculation result is used as the target conditional probability distribution. Based on the above processing flow, this application performs weighted calculations on the first conditional probability distribution and the second conditional probability distribution using the weighted fusion formula, and uses the generated calculation result as the required target conditional probability distribution, thereby automatically and accurately completing the fusion processing of the first conditional probability distribution and the second conditional probability distribution, improving the accuracy and intelligence of the generated target conditional probability distribution.
[0122] In some optional implementations of this embodiment, step S207 includes the following steps:
[0123] Obtain the preset greedy search strategy.
[0124] In this embodiment, the strategy flow of the greedy search strategy includes: in each generation step, selecting the token (word) with the highest probability from the target conditional probability distribution as the current output (target word). Additionally, the selected token is added to the generated sequence and used as the input context for the next step.
[0125] The target conditional probability distribution is processed based on the greedy search strategy to generate the corresponding first word element.
[0126] In this embodiment, the target conditional probability distribution is processed according to the strategy flow content based on the above-mentioned greedy search strategy, and the token with the highest probability is selected as the first word element generated, i.e., the target word element.
[0127] The first word element is used as the target word element.
[0128] In this embodiment, by selecting a decoding strategy (greedy search strategy) and dynamic context updates, the preceding quantization adjustments can be transformed into an actual result that conforms to the instructions and is free of illusions. Greedy search ensures efficiency, while dynamic context ensures global consistency, ultimately achieving a balance between "correcting errors" and "maintaining naturalness".
[0129] This application obtains a preset greedy search strategy; then processes the target conditional probability distribution based on the greedy search strategy to generate a corresponding first word element; subsequently, the first word element is used as the target word element. Based on the above processing flow, this application, by processing the target conditional probability distribution using a greedy search strategy, can automatically and efficiently generate target word elements that match the target conditional probability distribution, effectively improving the generation efficiency of target word elements.
[0130] In some optional implementations of this embodiment, step S207 includes the following steps:
[0131] Obtain the preset bundle search strategy.
[0132] In this embodiment, the strategy flow of the above-mentioned beam search strategy includes: setting a beam width K (e.g., K=3), and retaining the K candidate sequences with the highest probabilities at each step. For each candidate sequence, extend the target conditional probability distribution for the next step, and calculate the joint probability (product or logarithmic summation) of the candidate sequences. Subsequently, select the candidate sequence with the highest joint probability as the final output.
[0133] The target conditional probability distribution is processed based on the beam search strategy to generate the corresponding second word element.
[0134] In this embodiment, the target conditional probability distribution is processed according to the strategy flow content based on the above-mentioned beam search strategy, and the generated second word is used as the required target word.
[0135] The second word element is used as the target word element.
[0136] In this embodiment, by selecting a decoding strategy (beam search strategy) and dynamic context updates, the preceding quantization adjustments can be transformed into a consistent and unillusory actual result. This balances quality and efficiency, while the dynamic context ensures global consistency, ultimately achieving a balance between "correcting errors" and "maintaining naturalness."
[0137] This application obtains a preset beam search strategy; then processes the target conditional probability distribution based on the beam search strategy to generate a corresponding second word element; subsequently, the second word element is used as the target word element. Based on the above processing flow, this application, by processing the target conditional probability distribution using a beam search strategy, can automatically and efficiently generate target word elements that match the target conditional probability distribution, effectively ensuring the generation efficiency of target word elements and improving the accuracy of the generated target word elements.
[0138] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0139] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0140] Furthermore, the technical solution based on this application brings the following significant advantages and positive effects:
[0141] 1. Significantly reduced illusion rate: Experiments have shown that this application can significantly reduce the object illusion rate at the instance level and sentence level on multiple mainstream LVLMs (such as LLaVA-1.5, MiniGPT-4, Shikra) and on authoritative illusion evaluation benchmarks such as CHAIR and POPE, outperforming existing technologies.
[0142] 2. Zero cost and high efficiency: As an inference-time method that does not require training, this application avoids the high costs of data annotation and model training, and can be directly applied to existing pre-trained models, greatly reducing the threshold for technology implementation.
[0143] 3. Maintain and improve the model’s general capabilities: Unlike some illusion suppression methods that may lead to the degradation of model capabilities, this application can maintain or even slightly improve the model’s performance on general multimodal benchmarks while effectively suppressing illusions.
[0144] 4. Enhanced Fine-Grained Perception Capability: By forcing the model to pay equal attention to visual and textual details, this application can improve the model's fine-grained perception capability. For example, the model can notice and describe more subtle objects or attributes in the image (such as logos on clothing, the state of a specific object, etc.), thereby generating more accurate and richer descriptions.
[0145] High versatility and scalability: This application does not depend on a specific model architecture and can be widely applied to various LVLMs based on Transformer. Its plug-and-play feature makes it easy to integrate into existing systems, giving it great practical value.
[0146] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0147] It should be emphasized that, to further ensure the privacy and security of the aforementioned target text, the target text can also be stored in a node of a blockchain.
[0148] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0149] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0150] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0151] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0152] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0153] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a data processing device based on artificial intelligence, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0154] like Figure 3 As shown, the artificial intelligence-based data processing device 300 described in this embodiment includes: a receiving module 301, a first processing module 302, a quantization module 303, a second processing module 304, a calculation module 305, a fusion module 306, a third processing module 307, and an output module 308. Wherein:
[0155] The receiving module 301 is used to receive the input target image and the descriptive text corresponding to the target image;
[0156] The first processing module 302 is used to process the target image and the descriptive text based on a preset large-scale visual language model to generate corresponding word units;
[0157] The quantization module 303 is used to quantize the word based on a preset modal bias analysis strategy to obtain the modal bias quantization result of the word.
[0158] The second processing module 304 is used to calculate the original attention matrix of the word if the modality bias quantification result meets the preset attention intervention conditions, and to perform weight adjustment processing on the original attention matrix based on the modality bias quantification result to obtain the corresponding specified attention matrix.
[0159] The calculation module 305 is used to calculate and process the original attention matrix to obtain the corresponding first conditional probability distribution, and to calculate and process the specified attention matrix to obtain the corresponding second conditional probability distribution.
[0160] The fusion module 306 is used to fuse the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding target conditional probability distribution;
[0161] The third processing module 307 is used to decode the target conditional probability distribution to generate corresponding target words, and generate target text corresponding to the target image based on the target words.
[0162] The output module 308 is used to output the target text.
[0163] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based data processing method in the aforementioned embodiments, and will not be repeated here.
[0164] In some optional implementations of this embodiment, the quantization module 303 includes:
[0165] The first calculation submodule is used to calculate the text attention ratio of the specified word element based on a preset first calculation strategy when generating the specified word element; wherein the specified word element is any one of all the word elements;
[0166] The second calculation submodule is used to calculate the visual attention ratio of the specified word based on a preset second calculation strategy;
[0167] The first acquisition submodule is used to acquire the preset bias threshold;
[0168] The first analysis submodule is used to perform numerical analysis on the text attention ratio and the visual attention ratio based on the bias threshold, and generate the modal bias quantification result of the specified word.
[0169] In some optional implementations of this embodiment, the second processing module 304 includes:
[0170] The second analysis submodule is used to analyze the large visual language model and determine the intervention level in the large visual language model.
[0171] The submodule is invoked to call the corresponding target attention intervention formula based on the modality bias quantification result.
[0172] The adjustment submodule is used to adjust the attention weights of the original attention matrix at the intervention level based on the target attention intervention formula, so as to obtain the adjusted attention matrix.
[0173] The first determining submodule is used to use the adjusted attention matrix as the specified attention matrix.
[0174] In some optional implementations of this embodiment, the large visual language model includes a language decoder, which includes a preset computation layer and a correction layer; the computation module 305 includes:
[0175] An input submodule is used to input the specified attention matrix into the language decoder;
[0176] The third calculation submodule is used to calculate and process the specified attention matrix through the calculation layer to obtain the corresponding baseline probability distribution;
[0177] The correction submodule is used to correct the baseline probability distribution based on the correction layer to obtain the corresponding corrected probability distribution.
[0178] The second determining submodule is used to use the modified probability distribution as the second conditional probability distribution.
[0179] In some optional implementations of this embodiment, the fusion module 306 includes:
[0180] The second acquisition submodule is used to acquire the preset weighted fusion formula;
[0181] The fourth calculation submodule is used to perform weighted calculation on the first conditional probability distribution and the second conditional probability distribution based on the weighted fusion formula to obtain the corresponding calculation result;
[0182] The third determining submodule is used to use the calculation result as the target conditional probability distribution.
[0183] In some optional implementations of this embodiment, the third processing module 307 includes:
[0184] The third acquisition submodule is used to acquire the preset greedy search strategy;
[0185] The first processing submodule is used to process the target conditional probability distribution based on the greedy search strategy to generate the corresponding first word element;
[0186] The fourth determining submodule is used to use the first word element as the target word element.
[0187] In some optional implementations of this embodiment, the third processing module 307:
[0188] The fourth acquisition submodule is used to acquire the preset beam search strategy;
[0189] The second processing submodule is used to process the target conditional probability distribution based on the beam search strategy to generate the corresponding second word element;
[0190] The fifth determining submodule is used to use the second word element as the target word element.
[0191] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0192] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0193] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0194] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for data processing methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0195] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the artificial intelligence-based data processing method.
[0196] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0197] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based data processing method described above.
[0198] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0199] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.< / bos>
Claims
1. An artificial intelligence-based data processing method, characterized by, Includes the following steps: Receive the input target image and the corresponding descriptive text; The target image and the descriptive text are processed based on a pre-set large-scale visual language model to generate corresponding lexical units; The word units are quantified based on a preset modal bias analysis strategy to obtain the modal bias quantification results of the word units; If the modal bias quantification result meets the preset attention intervention conditions, then the original attention matrix of the word is calculated, and the weight adjustment processing of the original attention matrix is performed based on the modal bias quantification result to obtain the corresponding specified attention matrix; The original attention matrix is processed to obtain the corresponding first conditional probability distribution, and the specified attention matrix is processed to obtain the corresponding second conditional probability distribution; The first conditional probability distribution and the second conditional probability distribution are fused to obtain the corresponding target conditional probability distribution. The target conditional probability distribution is decoded to generate corresponding target words, and target text corresponding to the target image is generated based on the target words. The target text is then processed for output. The step of quantifying the lexical units based on a preset modality bias analysis strategy to obtain the modality bias quantification result of the lexical units specifically includes: When generating a specified word, the text attention ratio of the specified word is calculated based on a preset first calculation strategy; wherein, the specified word is any one of all the word elements. The visual attention ratio of the specified word element is calculated based on a preset second calculation strategy; Obtain the preset bias threshold; Numerical analysis is performed on the text attention ratio and the visual attention ratio based on the bias threshold to generate the modal bias quantification result for the specified word. 2.The artificial intelligence-based data processing method of claim 1, wherein, The step of adjusting the weights of the original attention matrix based on the modality bias quantification result to obtain the corresponding specified attention matrix specifically includes: The large visual language model is analyzed to determine the intervention levels within it. The corresponding target attention intervention formula is invoked based on the modality bias quantification results; Based on the target attention intervention formula, the original attention matrix is adjusted in the intervention level to obtain the adjusted attention matrix; The adjusted attention matrix is used as the designated attention matrix. 3.The artificial intelligence-based data processing method of claim 1, wherein, The large-scale visual language model includes a language decoder, which includes a pre-defined computation layer and a correction layer; the step of calculating and processing the specified attention matrix to obtain the corresponding second conditional probability distribution specifically includes: The specified attention matrix is input into the language decoder; The specified attention matrix is processed by the computation layer to obtain the corresponding baseline probability distribution; The baseline probability distribution is corrected based on the correction layer to obtain the corresponding corrected probability distribution. The modified probability distribution is used as the second conditional probability distribution. 4.The artificial intelligence-based data processing method of claim 1, wherein, The step of fusing the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding target conditional probability distribution specifically includes: Obtain the preset weighted fusion formula; Based on the weighted fusion formula, the first conditional probability distribution and the second conditional probability distribution are weighted and calculated to obtain the corresponding calculation results; The calculation result is used as the target conditional probability distribution. 5.The artificial intelligence-based data processing method of claim 1, wherein, The step of decoding the target conditional probability distribution to generate the corresponding target lexical units specifically includes: Obtain the preset greedy search strategy; The target conditional probability distribution is processed based on the greedy search strategy to generate the corresponding first word element; The first word element is used as the target word element.
6. The data processing method based on artificial intelligence according to claim 1, characterized in that, The step of decoding the target conditional probability distribution to generate the corresponding target lexical units specifically includes: Obtain the preset beam search strategy; The target conditional probability distribution is processed based on the beam search strategy to generate the corresponding second word element; The second word element is used as the target word element.
7. A data processing device based on artificial intelligence, characterized in that, include: A receiving module is used to receive an input target image and descriptive text corresponding to the target image; The first processing module is used to process the target image and the descriptive text based on a preset large-scale visual language model to generate corresponding word units; The quantization module is used to quantize the word units based on a preset modal bias analysis strategy to obtain the modal bias quantization result of the word units; The second processing module is used to calculate the original attention matrix of the word if the modality bias quantification result meets the preset attention intervention conditions, and to perform weight adjustment processing on the original attention matrix based on the modality bias quantification result to obtain the corresponding specified attention matrix. The calculation module is used to calculate and process the original attention matrix to obtain the corresponding first conditional probability distribution, and to calculate and process the specified attention matrix to obtain the corresponding second conditional probability distribution; The fusion module is used to fuse the first conditional probability distribution and the second conditional probability distribution to obtain the corresponding target conditional probability distribution. The third processing module is used to decode the target conditional probability distribution to generate corresponding target words, and generate target text corresponding to the target image based on the target words; The output module is used to process the target text for output. The quantization module includes: The first calculation submodule is used to calculate the text attention ratio of the specified word element based on a preset first calculation strategy when generating the specified word element; wherein the specified word element is any one of all the word elements; The second calculation submodule is used to calculate the visual attention ratio of the specified word based on a preset second calculation strategy; The first acquisition submodule is used to acquire the preset bias threshold; The first analysis submodule is used to perform numerical analysis on the text attention ratio and the visual attention ratio based on the bias threshold, and generate the modal bias quantification result of the specified word.
8. A computer device, characterized in that, The system includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the artificial intelligence-based data processing method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data processing method based on artificial intelligence as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Elastic converter service system self-adaptive through lexical elements
CN119848173A
Retrieval method, device and equipment based on attention guidance and medium
CN120179878A