Image and text-based reasoning method and device, equipment and medium
By generating inference chains that include text inference data and bounding box coordinate data, the problem of insufficient visual detail association in multimodal models is solved, achieving a more accurate and transparent inference process, which is suitable for intelligent assisted decision-making in the fields of smart healthcare and finance.
Patent Information
- Application Number
- CN202511187467.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-11
AI Technical Summary
Existing multimodal large-scale language models lack visual detail associations when processing mixed text and image inputs, resulting in insufficient clarity and questionable authenticity in the reasoning process. Furthermore, they struggle to maintain coherent semantic associations across images when dealing with complex problems, and existing methods have failed to achieve end-to-end deep fusion.
By generating an inference chain containing text inference data and bounding box coordinate data through a pre-set inference model, the collaborative thinking between the problem and the image is realized, improving the relevance of visual input. The GRPO algorithm is used for model training, reducing the number of samples required and improving accuracy.
It improves the authenticity and accuracy of the reasoning process, enhances the transparency of logical reasoning, and is suitable for scenarios such as visual question answering, object counting, and spatial relationship judgment, while reducing the dependence on prompting engineering or auxiliary modules.
Smart Images

Figure CN120930804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and is applied to the fields of smart healthcare and finance. In particular, it relates to a reasoning method, apparatus, device, and medium based on images and text. Background Technology
[0002] Stepwise reasoning methods based on inference chains have been widely applied in language and vision-language multimodal tasks to improve the transparency of model decisions through step-by-step explanation. However, current open-source visual reasoning models have significant limitations when handling mixed text and image inputs: the generated inference chains only contain natural language text and fail to establish explicit associations with visual details (such as object attributes and spatial relationships) in the input images, resulting in insufficient clarity and questionable authenticity of the reasoning process, and the model cannot truly achieve "thinking in conjunction with images." This deficiency stems from two technical challenges: on the one hand, mainstream large multimodal language models (MLLMs) are based on language token generation as their core architecture and lack a mechanism for directly generating or referencing visual elements in the inference chain; on the other hand, when dealing with complex problems that require interleaving multiple images (such as temporal analysis or contrastive reasoning), most MLLMs are limited by their context processing capabilities and struggle to maintain coherent semantic associations across images.
[0003] Although solutions exist at present, there are still obvious shortcomings: some methods separate visual authenticity verification from text inference chain generation into independent processes, failing to achieve deep end-to-end integration; other methods, although they use external tools (such as object detectors) or prompting engineering to extract visual features, do not essentially endow the model with the ability to autonomously generate interleaved multimodal inference sequences, which limits the applicability of the technology. Summary of the Invention
[0004] This invention provides a reasoning method, apparatus, computer device, and storage medium based on images and text to solve the technical problem of insufficient association between the reasoning chain and visual input, and limited applicability.
[0005] In a first aspect, embodiments of the present invention provide an image and text-based reasoning method applied to a reasoning platform, comprising: receiving question data containing images and text input by a client; inputting the question data into a preset reasoning model; the reasoning model performing reasoning based on the question data to generate a reasoning chain, wherein the reasoning chain includes text reasoning data and bounding box coordinate data; and the reasoning model generating a target answer based on the text reasoning data.
[0006] Secondly, embodiments of the present invention also provide an image and text-based reasoning apparatus applied to a reasoning platform, which includes a unit for performing the above-described method.
[0007] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0008] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the above-described method.
[0009] This application provides a reasoning method, apparatus, device, and medium based on images and text. It generates a reasoning chain containing textual reasoning data and bounding box coordinate data through a preset reasoning model, enabling collaborative thinking between questions and images. This improves the correlation between the reasoning chain and visual input, allowing users to more accurately understand the reasoning process and answers in scenarios requiring logical reasoning based on image content, such as visual question answering, object counting, and spatial relationship judgment. This enhances the reliability and accuracy of the reasoning. Furthermore, the reasoning model does not rely on prompting or auxiliary modules during the reasoning process, thus improving its applicability. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A schematic flowchart illustrating the image and text-based reasoning method provided in an embodiment of the present invention;
[0012] Figure 2 A schematic diagram of a sub-process of the image and text-based reasoning method provided in an embodiment of the present invention;
[0013] Figure 3 A schematic diagram of a sub-process of the image and text-based reasoning method provided in an embodiment of the present invention;
[0014] Figure 4 A schematic diagram of a sub-process of the image and text-based reasoning method provided in an embodiment of the present invention;
[0015] Figure 5 A schematic diagram of a sub-process of the image and text-based reasoning method provided in an embodiment of the present invention;
[0016] Figure 6 A schematic diagram of a sub-process of the image and text-based reasoning method provided in an embodiment of the present invention;
[0017] Figure 7 A schematic diagram of a sub-process of the image and text-based reasoning method provided in an embodiment of the present invention;
[0018] Figure 8 A schematic block diagram of an image and text-based reasoning device provided in an embodiment of the present invention;
[0019] Figure 9 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0022] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0023] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0024] Please see Figure 1This is a schematic flowchart illustrating the image and text-based reasoning method provided in this invention. In this application, the image and text-based reasoning method is applied in a reasoning platform, particularly in business areas such as smart healthcare and fintech. Through collaborative reasoning using text and bounding boxes, the method achieves dual protection of "visual evidence anchoring + logical chain tracing" within the reasoning chain. This improves the accuracy of the reasoning results and enhances the interpretability of the decision-making process, providing efficient and reliable intelligent assistance for business scenarios. For example, in the field of smart healthcare, it can be applied to medical image diagnosis assistance scenarios: when a doctor inputs a patient's CT image and the text question "identify tumor location and analyze benignity / malignancy," the reasoning chain generates text reasoning data (e.g., "The image shows an irregular nodule in the lower lobe of the left lung with obvious spiculated edges") and bounding box coordinate data (precisely defining the nodule region). The text reasoning data, combined with medical knowledge, interprets the image features, while the bounding box coordinate data provides intuitive visual positioning. The two corroborate each other, helping doctors quickly locate lesions and understand the analysis logic. Alternatively, in the fintech field, this method is applicable to market dynamics analysis scenarios: After an analyst inputs a stock chart and a text question asking "determine the support level and the reason for the fluctuation," the inference chain generates text inference data (such as "three rebounds occurred at the 30 yuan level, forming strong support, which is related to positive quarterly financial reports") and bounding box coordinate data (marking the price range corresponding to the support level). The text inference data correlates with market information to interpret chart patterns, while the bounding box coordinate data clarifies key points, helping analysts quickly grasp market trends.
[0025] Therefore, by generating the inference chain containing the text inference data and the bounding box coordinate data through the inference model, the collaborative thinking of the problem and the image is realized, the correlation between the inference chain and the visual input is improved, and the basis for inference is traceable, thereby improving the accuracy and credibility of the inference.
[0026] This application provides a reasoning method, apparatus, computer device, and storage medium based on images and text. The reasoning method based on images and text includes: receiving question data containing images and text input by a client; inputting the question data into a preset reasoning model; and generating a reasoning chain and a target answer through the reasoning model, wherein the reasoning chain includes text reasoning data and bounding box coordinate data.
[0027] This application generates a reasoning chain containing the text reasoning data and the bounding box coordinate data through a preset reasoning model, enabling collaborative thinking between questions and images, improving the correlation between the reasoning chain and visual input, and thus enabling users to more accurately understand the reasoning process and the reasoning answer in scenarios that require logical reasoning based on image content, such as visual question answering, object counting, and spatial relationship judgment, thereby improving the authenticity, reliability, and accuracy of the reasoning; moreover, the reasoning model does not rely on prompting engineering or auxiliary modules during the reasoning process, thereby improving the applicability of the reasoning model.
[0028] Figure 1 This is a flowchart illustrating the image and text-based reasoning method provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S10-S30.
[0029] S10. Receive question data containing images and text input from the client, and input the question data into a preset inference model;
[0030] Specifically, as more scenarios require logical reasoning based on image content, higher demands are placed on the technology of reasoning models. For example, in scenarios such as visual question answering, object counting, and spatial relationship judgment, the model can reason based on image content and questions to help users obtain answers.
[0031] In this embodiment, the question data refers to data containing images and text. The question data has detailed images and a clear question, so that the model can obtain known conditions and questions from the images and text, and then perform reasoning to obtain an accurate answer. The reasoning model refers to a model pre-set in the computer, which is an algorithmic framework for implementing reasoning functions. The reasoning model is used to perform reasoning based on the question data to obtain the final accurate answer.
[0032] When a user needs to use the reasoning model to calculate an answer, the user can input the question data on the client. Specifically, the user can input the question data into the computer by directly taking a picture of the image and the question, or by directly taking a picture of the image and manually inputting the question, or by extracting the question data containing the image and text from another document. The specific method is not limited.
[0033] When the computer receives the question data input by the client, it inputs the question data into the reasoning model, so that the reasoning model can reason based on the images and text in the question data, and finally obtain the target answer. The target answer and the complete reasoning process (i.e., the reasoning chain) are then returned to the client to improve the accuracy and reliability of the reasoning. At the same time, directly inputting the question data into the reasoning model to obtain the answer is efficient and the data input process is less prone to errors, thus improving the accuracy of the reasoning.
[0034] In one embodiment, such as Figure 2 As shown, steps S11-S12 precede step S10.
[0035] S11. Receive a training set, which includes multiple types of training data;
[0036] S12. The inference model is trained based on the GRPO algorithm and each of the training data.
[0037] Specifically, the training set is a sample set used for model learning, providing the model with a learning basis. The training set includes various types of training data, covering VSR (Spatial Relationships), TallyQA (Counting), and GQA (Gross Quality Assignment), enabling the inference model to reason about various types of question data. The training data consists of image-question-answer triplets, and only 20 samples of each type are needed for effective training. The GRPO algorithm is a reinforcement learning algorithm specifically designed for multimodal inference chain generation. Its core addresses the key issues of unstable reward signals and model deviation from prior knowledge in complex tasks through group reward distribution calibration and policy update constraints.
[0038] In this embodiment, model training is the core step in achieving accurate inference. Before using the inference model, it needs to be trained. In this embodiment, the inference model is trained based on the GRPO algorithm, which eliminates the need for dense annotation and requires fewer samples, thus reducing costs. Furthermore, the training process does not require inference chain annotation or bounding box annotation. Through reinforcement learning, the model masters the real inference paradigm, resulting in high accuracy and reliability of the inference model.
[0039] In training the inference model, the training set is first received. This training set contains various types of training data, covering image and text combinations in different scenarios. For example, in the field of smart healthcare, it may include patients' medical images (such as CT scan images) and corresponding diagnostic description text; in the financial field, it may include stock charts and related market analysis reports. This diverse training data provides the model with rich learning materials, enabling it to adapt to the inference needs of different fields.
[0040] The training of the inference model is based on the GRPO algorithm. The learning process is completed through interaction with each training data, thereby enabling it to master the real reasoning paradigm through reinforcement learning, thus improving the inference accuracy and reliability of the inference model.
[0041] For example, in a smart healthcare scenario, the inference model receives multiple training data sets containing medical images and diagnostic questions, such as "determine the possible cause of illness based on a patient's lung CT image and description of cough symptoms." The inference model first attempts to generate an inference result, and then receives feedback based on a preset reward mechanism—if the inference result conforms to medical logic (e.g., accurately identifying nodules in the image and associating them with pneumonia symptoms), a higher reward is given; if a misjudgment occurs (e.g., classifying a benign nodule as malignant), a lower reward is given. This feedback guides the inference model to adjust its parameters and gradually learn the correlation between medical images and symptoms.
[0042] Alternatively, in training within the financial field, the training data processed by the inference model might be "combining a stock's candlestick chart with macroeconomic news to predict its trend for the coming week." The GRPO algorithm encourages the inference model to balance accuracy and stability during the learning process: on the one hand, it incentivizes the model to correctly analyze the volatility characteristics in the candlestick chart and key information in the news (such as policy adjustments and company financial reports) through a reward mechanism, generating reasonable predictions; on the other hand, it limits drastic changes in the inference model's strategy, avoiding imbalances in prediction logic due to individual extreme data, and ensuring robust inference capabilities when facing complex market dynamics.
[0043] Therefore, in this embodiment, through continuous learning of various types of training data, the inference model gradually grasps the deep correlation between images and text. Whether it is assisting doctors in interpreting images and medical records in smart healthcare, or supporting analysts in analyzing charts and market information in the financial field, it can output inference results that conform to professional logic, providing reliable decision support for practical application scenarios.
[0044] In one embodiment, such as Figure 3 As shown, step S12 includes steps S121-S122.
[0045] S121. Perform inference on the training data and calculate rewards to obtain a composite reward value, wherein the reward calculation includes real inference format reward calculation, optional real target count reward calculation and answer accuracy reward calculation;
[0046] S122. The inference model is optimized by calculating the group normalization advantage and the GRPO objective function for the composite reward value.
[0047] Specifically, the composite reward value refers to the comprehensive reward value obtained after evaluating multiple rewards. Generally, it is the sum of the scores calculated from multiple rewards, used to quantify the quality of the inference model output. The reward calculation includes three rewards: the calculation of the true reasoning format reward, the calculation of the optional true target count reward, and the calculation of the answer accuracy reward.
[0048] The true inference format reward calculation is used to encourage the model to generate inference output that conforms to a specific format, including the correct use of special inference format markers (e.g., <think> 、< / think> , <rethink> 、< / rethink> The use of special inference format markers (such as those for bounding box coordinates) and other formatting inference chains makes the model's inference process more structured and standardized, enhancing the logic and interpretability of the inference. For example, when checking the inference chains generated by the inference model, each correctly placed marker pair increases the reward by 0.5 if the special inference format markers are used correctly; a reward of 0.5 is also given if at least one grammatically correct bounding box exists in the inference chain (detected through regular expression matching). This guides the inference model to learn to switch flexibly and systematically between text and coordinates, generating formatted inference chains.
[0049] The real-object counting reward calculation, referring to training examples on visual counting-related datasets, is an optional reward calculation used to encourage the inference model to accurately generate bounding boxes that match the number of target objects during inference, thereby improving the model's accuracy and completeness in counting tasks. For example, when training the model on tasks involving object counts, the number of bounding boxes generated in the model's inference chain is compared with the number of target objects in the true answer. If they match, the reward is set to 0.5; otherwise, no reward is given. This ensures that the model considers completeness when labeling targets, avoiding multiple boxes, missing boxes, or arbitrary box drawing, making the inference results generated by the model more consistent with reality in terms of counting.
[0050] The accuracy reward calculation combines the judgment of an external Vision-Language Model (such as GPT-4o) and BLEU similarity to more comprehensively and accurately evaluate the correctness of the model's final answer, improve the accuracy of the model's responses, prevent the model from exploiting loopholes, and ensure that the answers generated by the model are semantically closer to the true answers. For example, GPT-4o is used to evaluate the triple consisting of the question, the predicted answer, and the true answer to obtain a binary correctness score (0 or 1); simultaneously, the sentence-level BLEU-1 similarity between the predicted answer and the true answer is calculated. Finally, the accuracy of the model's reasoned answer is comprehensively evaluated through the correctness score and similarity score, and a corresponding reward is given to guide the model to generate more accurate final answers.
[0051] The group normalization advantage calculation standardizes the composite reward value of all training data to eliminate scale differences and stabilize the training process. The GRPO objective function generates a policy update signal based on the group normalized composite reward value through probability ratio constraints and KL penalties, ultimately outputting model parameters with high inference quality. The GRPO objective function aims to balance maximizing expected rewards with maintaining closeness to the reference policy to promote stable learning. Therefore, during the optimization process, the inference model adjusts its policy parameters (i.e., the GRPO objective function) based on the calculated group normalization advantage (i.e., the group normalized composite reward value), making the inference sequence generated by the policy more consistent with expectations, continuously improving the model's performance in inference tasks, and helping to enhance the model's inference ability and training efficiency.
[0052] During the training process of the inference model, after receiving the training data, the inference model performs inference on each piece of training data to generate an inference result containing an inference chain and an inference answer. Then, a reward calculation is performed on the inference result to obtain the composite reward value. Specifically, rewards are calculated for three categories: true inference format reward, optional true target count reward, and answer accuracy reward, to quantify the quality of the inference result output by the inference model. Then, based on the generated composite reward value, a group normalization advantage calculation is used to compare the reward value of this sample with that of other samples in the same batch. The rewards are standardized to eliminate the impact of differences in reward scales across different scenarios. For example, the reward for "correct lesion identification" in medical data and the reward for "accurate trend prediction" in financial data will be normalized to the same evaluation scale, enabling the model to fairly compare the inference performance of samples from different domains. Subsequently, the GRPO objective function optimization will adjust the model parameters based on these normalized advantage values. For example, in the medical scenario, it ensures that the model does not deviate from the basic anatomical knowledge it has mastered when learning to identify new lesion features; in the financial scenario, it avoids the model from over-adjusting the long-term trend analysis logic due to short-term market fluctuations.
[0053] In this way, the model can continuously improve its accuracy in reasoning tasks across different domains while maintaining the stability of its strategy, ultimately achieving efficient fusion reasoning of image and text information. Furthermore, based on the training method of the aforementioned reasoning model, the GRPO reinforcement learning algorithm is employed. This algorithm guides the model's learning through rewards based on the true reasoning format, optional rewards based on the true target count, and rewards based on answer accuracy. This eliminates the need for dense annotations, reducing costs and improving efficiency and applicability. Moreover, the method provides a structured output format, which reduces the learning difficulty. Additionally, the training data is carefully selected, focusing on core task scenarios, thus reducing the number of samples required.
[0054] S20. The inference model performs inference based on the question data to generate an inference chain, wherein the inference chain includes text inference data and bounding box coordinate data.
[0055] Specifically, the reasoning chain is a complete record of the logical deduction process of the reasoning model, including intermediate steps from problem analysis to conclusion formation, and reflecting the reasoning logic and basis. The reasoning chain includes textual reasoning data and bounding box coordinate data. The textual reasoning data is the analysis process presented in textual form within the reasoning chain, including the interpretation, association, and deductive logic of information. The bounding box coordinate data is the positional coordinate information of key areas marked in the image, accurately locating target objects in the image through coordinate values, providing visual evidence for textual reasoning.
[0056] In the image- and text-based reasoning method, after the question data is input into the reasoning model, the model generates a reasoning chain containing textual reasoning data and bounding box coordinate data. The textual reasoning data demonstrates the model's logical process in analyzing the question data, and the bounding box coordinate data is interwoven into the textual reasoning data to aid in its interpretation. The coordinate values accurately locate the target region of the image in the question data, providing visual evidence for the textual reasoning data. This enables collaborative thinking between the question and the image in the reasoning chain, improving the correlation between the reasoning chain and the visual input. This makes the reasoning process of the reasoning model clearer and more understandable to the user, enhancing the reliability of the reasoning model's inference.
[0057] For example, in the field of smart healthcare, when the client inputs a patient's brain MRI image and the text question "Determine if there is a cerebral infarction and its specific location," the inference model first identifies abnormal regions in the image, and then analyzes and generates the inference chain in conjunction with the text information. In the generated inference chain, the text inference data records the analysis process in detail, such as "The image shows an abnormal signal in the left basal ganglia region; the T1-weighted image shows a low signal, and the T2-weighted image shows a high signal, consistent with the imaging characteristics of cerebral infarction." The bounding box coordinate data precisely marks the location of the left basal ganglia region, defining the abnormal area through specific coordinate values, giving the text inference a clear visual direction. Based on this inference chain, the final target answer generated by the inference model clearly indicates "a cerebral infarction exists, located in the left basal ganglia region," providing a reference for the doctor's diagnosis.
[0058] Therefore, this reasoning chain, which includes the text reasoning data and the bounding box coordinate data, enables the reasoning model to possess both the clarity of textual logic and the accuracy of image positioning. The two mutually corroborate each other, enhancing the credibility and interpretability of the target answer. In practical applications across different fields, it can better meet users' needs for transparency in the reasoning process. At the same time, the reasoning model does not rely on hint engineering or auxiliary modules during the reasoning process, thereby improving the applicability and efficiency of the reasoning model.
[0059] S30. The reasoning model generates the target answer based on the text reasoning data.
[0060] Specifically, the target answer is the final conclusion drawn by the reasoning model from the input question data, and it must be based on the logical deduction of the reasoning chain, possessing clarity and accuracy.
[0061] The process by which the reasoning model generates the target answer essentially involves deep integration and logical refinement of the textual reasoning data. Through structured analysis, it transforms scattered reasoning content into precise conclusions. In practice, the model first comprehensively analyzes the hierarchical structure of the textual reasoning data, identifies core arguments, supporting evidence, and potential logical connections, and then performs targeted processing based on domain knowledge.
[0062] For example, in the field of smart healthcare, if the text inference data is "The patient's chest CT scan shows patchy shadows in the lower lobes of both lungs (preliminary inference). Combining clinical symptoms and laboratory tests, the shadow density is consistent with the characteristics of bacterial pneumonia, and the pleura has not been involved (secondary inference)," the model will first extract key information: imaging features (shadows in the lower lobes of both lungs), clinical correlation (bacterial pneumonia), and lesion extent (not involving the pleura). Subsequently, this information is integrated according to medical diagnostic standards, redundant expressions are eliminated, and the target answer is finally generated: "The patient is diagnosed with bacterial pneumonia, the lesion is located in the lower lobes of both lungs, and has not yet spread to the pleura." Throughout the process, the model ensures that the target answer contains both the core diagnostic results and retains key location information, providing a clear basis for treatment plan formulation.
[0063] In the financial field, when the textual reasoning data includes statements such as "A certain stock's revenue has increased by 15% in the past three months, but its net profit has decreased by 5% (preliminary reasoning), with a 20% increase in raw material prices being the main reason, and industry data showing that raw material prices are unlikely to fall in the short term (secondary reasoning)," the model focuses on the core contradiction: the inverse relationship between revenue and net profit and its causes. By associating with market patterns, the scattered data analysis is condensed into the target answer: "The stock's revenue growth is dragged down by rising costs, putting pressure on net profit, and short-term profit expectations are cautiously influenced by raw material prices." This target answer summarizes key financial characteristics and identifies influencing factors, providing clear guidance for investment decisions.
[0064] When using the image- and text-based reasoning method, the reasoning model generates the target answer based on the text reasoning data following the logic of "extracting core information—verifying logical consistency—aligning with problem requirements": first, it filters content directly related to the problem data from the text reasoning data; then, it checks whether there are contradictions in the reasoning chain (such as the matching between image features and diagnoses in medicine); finally, it outputs the conclusion in concise and clear language. This process retains the key evidence for reasoning while avoiding redundant information interference, making the target answer both accurate and practical, effectively supporting actual business decisions.
[0065] More specifically, the reasoning model first generates the reasoning chain based on the question data, and then derives the target answer based on the text reasoning data in the reasoning chain. Both the reasoning chain and the target answer are displayed on the client, so that users can have a clear understanding of the reasoning process of the reasoning model and improve the authenticity, reliability and traceability of the target answer.
[0066] In one embodiment, such as Figure 4 As shown, the text reasoning data includes preliminary text reasoning data and secondary text reasoning data, and step S20 may include steps S21-S23.
[0067] S21. Perform preliminary reasoning based on the problem data to obtain the preliminary text reasoning data and the bounding box coordinate data;
[0068] S22. Perform secondary reasoning based on the question data, the preliminary text reasoning data, and the bounding box coordinate data to obtain the secondary text reasoning data;
[0069] S23. Obtain the target answer based on the secondary text reasoning data.
[0070] Specifically, the textual reasoning data refers to the logical deductions presented in text form within the reasoning chain, including the interpretation, association, and analysis of information. The textual reasoning data includes preliminary textual reasoning data and secondary textual reasoning data. The preliminary textual reasoning data is the reasoning content generated by the reasoning model after its initial analysis of the problem data, providing a foundation for subsequent reasoning. The secondary textual reasoning data is the reasoning content obtained through further in-depth analysis based on the preliminary reasoning results, which is closer to the final conclusion. The bounding box coordinate data is the location information of key areas marked in the image, accurately locating the target through coordinates and providing visual support for textual reasoning.
[0071] In this embodiment, during the reasoning process of the reasoning model, the model first performs preliminary reasoning based on the question data (i.e., the preliminary reasoning stage). After receiving the image and text question, the model first understands and analyzes the question. Combining existing knowledge and preliminary perception of the image, it determines the image region that needs attention, describes the thought process in natural language, and inserts bounding box coordinates as appropriate (e.g., "There are several animals in the picture, one of which is a cow, and its coordinates are (42, 73, 433, 296)"). After completing this series of operations, the thinking output of this stage ends.
[0072] Therefore, the preliminary reasoning stage obtains the preliminary text reasoning data and the bounding box coordinate data from the question data. The preliminary text reasoning data is derived from the initial analysis of the question data, and the bounding box coordinate data is the location coordinates of key areas identified through image recognition of the question data. This preliminary reasoning stage enables the reasoning model to have a structured reasoning process, making the model's thought process clearer and more organized, enhancing the logic and interpretability of the reasoning; it helps the model integrate visual information (the bounding box coordinate data) and linguistic information (the preliminary text reasoning data), improving reasoning accuracy; and it facilitates the subsequent secondary reasoning stage in reviewing and optimizing the preceding thought processes, further refining the reasoning process and improving the reliability of the final answer.
[0073] Then, the inference model performs secondary inference (i.e., the secondary inference stage) based on the question data, the preliminary text inference data, and the bounding box coordinate data. The inference model first reviews the image regions corresponding to the bounding box coordinate data mentioned in the preliminary inference stage to confirm again whether the understanding of these regions (i.e., the preliminary text inference data) is accurate. For example, when determining whether "the truck is below the cat," the positional relationship determined by the bounding box coordinates of the truck and the cat is re-examined. Then, based on the reconfirmed image information and the text question, the reasoning logic and conclusions of the initial reasoning stage are reviewed to consider any loopholes or irrationalities. If problems are found, the reasoning model adjusts its reasoning approach and reorganizes the logic. The results of this rethinking are integrated to form more complete reasoning content (i.e., the secondary text reasoning data). The optimized reasoning content is then output, which more accurately reflects the relationship between the image information and the question's answer. For instance, in the question "How many zebras are in the picture?", the secondary reasoning stage might confirm or correct the number of zebras and coordinates given in the initial reasoning stage, ultimately outputting something like, "The previously given coordinates accurately cover all visible zebras; after reconfirmation, there are a total of 7 zebras in the picture." This provides a more reliable basis for arriving at the final answer.
[0074] Therefore, the secondary reasoning stage verifies the reasoning content of the preliminary text reasoning data, determines whether the preliminary text reasoning data is correct, and obtains more detailed reasoning results. That is, in the secondary reasoning stage, the focus is on solving the consistency verification of text and image and the completion of the evidence chain. It mainly corrects or supplements the reasoning results of the preliminary reasoning to obtain the secondary text reasoning data.
[0075] Finally, the secondary text reasoning data is integrated and refined to obtain the target answer, thereby improving the accuracy and reliability of the target answer. The initial text reasoning data, the bounding box coordinate data, and the secondary text reasoning data constitute the reasoning chain. The reasoning chain and the target answer are displayed together on the client side so that users can understand the reasoning process, thus improving the reliability of the reasoning model.
[0076] For example, in the field of smart healthcare, when a chest CT image and the text "Please identify high-risk pulmonary nodules" are input into the inference model, the inference model first performs preliminary inference on the image and text, and generates the preliminary text inference data: 3 solid nodules were found, with a diameter of 6-8 mm; the preliminary inference results of the bounding box coordinate data: (120,88,165,130), (210,152,248,190), (305,75,345,110). At this time, the bounding box accurately marks the location of the nodules, but the text only provides a preliminary description and does not define the target answer (i.e., identify 3 nodules) or judge the risk of malignancy.
[0077] Secondly, the inference model performs secondary inference based on the question data, the preliminary text inference data, and the bounding box coordinate data to obtain the secondary text inference data: the previously given coordinates accurately cover all visible nodules. After reconfirmation, there are 3 solid nodules in the image. Nodule 1 (120, 88, 165, 130): lobulated edge, high probability of malignancy; Nodule 2 (210, 152, 248, 190): calcification features, possibly benign; Nodule 3 (305, 75, 345, 110): vascular clustering sign, biopsy recommended as a secondary inference result. At this time, the inference model associates the image details with the bounding box coordinate data, upgrades the "3 nodules" in the preliminary inference result to a targeted diagnosis, and exposes the contradictions that were not found in the preliminary stage (the calcification features of nodule 2 overturn the malignancy hypothesis).
[0078] Finally, based on the results of the secondary reasoning, the target answer was extracted: 2 high-risk nodules (nodules 1 and 3), and puncture biopsy is recommended.
[0079] Therefore, in this embodiment, the image and text-based reasoning method adopts a phased reasoning approach, so that the initial text reasoning data and the bounding box coordinate data form a basic support, and the secondary text reasoning data deepens the logical connection on this basis. The final generated target answer has both the accuracy of visual evidence and the rigor of logical deduction, making it more reliable in decision support in professional fields.
[0080] In one embodiment, such as Figure 5 As shown, the problem data includes image data and text data, and step S21 may include steps S211-S212.
[0081] S211. Locate the target region based on the image data and the text data;
[0082] S212. Generate the bounding box coordinate data based on the target area and the preset reference system.
[0083] Specifically, the problem data is the information to be reasoned upon input by the client, including image data and text data; wherein, the image data is a graphic or image containing visual information, and the text data is textual content describing the problem. The target region is a key area in the image data that is related to the problem in the text data, and is the core object of reasoning and analysis. The preset reference system is a coordinate system built into the model, used to unify the measurement standard of the target position in the image and ensure the consistency of the bounding box coordinates. The bounding box coordinate data is the numerical information of the target region position marked in the image data, and the coordinates determine the specific range of the core object in the image data.
[0084] In this embodiment, the bounding box coordinate data is generated based on the image data and text data. The specific steps are as follows: First, the target region is located by combining the image data and text data in the question data. Specifically, the inference model first parses the semantic focus of the text data instruction, relies on pre-trained visual grounding capabilities to locate the target region in the image related to the question, and locks the associated target region in the image data. The target region can be one or more. For example, in a smart healthcare scenario, the inference model is input with the image data of a brain MRI image and the text data "Locate the infarct lesion in the left basal ganglia region". First, the inference model parses the text keyword "left basal ganglia region" in the text data and determines that the region is located in the deep left hemisphere of the brain in the medical knowledge base. Then, the target region that conforms to anatomical features (such as a high signal area in T2 sequence) is identified in the image data of the MRI image.
[0085] Then, the inference model generates the bounding box coordinates based on the target region and the preset reference frame. That is, the inference model directly generates bounding box coordinates that conform to the syntax rules (such as quadruples in the form of (x1,y1,x2,y2)) without the need for additional pixel input or external tools.
[0086] More specifically, the generated bounding box coordinate data must meet the format requirements of the reward mechanism in the training process of the inference model. Regular matching is used to ensure that the bounding box coordinate data is integer and has a valid structure in order to obtain corresponding rewards and promote the model to learn and standardize the output.
[0087] The bounding box coordinates are interwoven with the initial textual reasoning data in the reasoning chain, clearly pointing to the visual evidence upon which the reasoning relies, such as "the zebra's position in the image is (200, 168, 248, 202)," providing a basis for verification in the subsequent secondary reasoning stage. This process is optimized through reinforcement learning, enabling the model to gradually learn to generate accurate bounding box coordinate data at appropriate reasoning steps, achieving an organic integration of visual information and textual reasoning, thereby improving the reasoning accuracy and reliability of the reasoning model.
[0088] Therefore, by first locating the target area by combining the image data and the text data, and then generating the bounding box coordinate data based on the preset reference system, the preliminary reasoning results not only clearly define the location of the analysis object, but also provide accurate visual anchors for subsequent secondary reasoning, making the entire reasoning process clear in its direction and traceability from the very beginning.
[0089] In one embodiment, such as Figure 6 As shown, step S22 may include steps S221-S224.
[0090] S221. Determine whether the bounding box coordinate data and the target region match;
[0091] S222. If a match is found, determine whether the preliminary text inference data is correct based on the bounding box coordinate data and the text data.
[0092] S223. If correct, then generate the secondary text reasoning data based on the preliminary text reasoning data;
[0093] S224. If incorrect, then the secondary text reasoning data is obtained by re-inferring based on the bounding box coordinate data and the text data.
[0094] Specifically, during the generation of the secondary text inference data by the inference model, the matching between the bounding box coordinate data and the target region needs to be verified first. Specifically, the system compares the image range marked by the bounding box coordinate data with the actual range of the target region, checking whether the bounding box coordinate data completely covers the target region and whether there is any positional offset or excessive / inappropriate range. If the coordinate range of the bounding box coordinate data basically matches the actual contour of the target region, and all key features are included, it is determined to be a match; if the bounding box coordinate data significantly deviates from the target region or does not completely cover the key features, it is determined to be a mismatch.
[0095] When the bounding box coordinate data matches the target region, the inference model further combines the bounding box coordinate data and the text data to verify the correctness of the preliminary text inference data. At this point, the system checks whether the description of the target region in the preliminary text inference data (such as features, attributes, state, quantity, positional relationships, etc.) is consistent with the actual situation of the region marked by the bounding box coordinate data, and whether it conforms to the question orientation of the text data. For example, if the text data requires analysis of "the color of a circular target in an image," while the preliminary text inference data describes it as "a square target is red," even if the bounding box coordinate data accurately marks the circular target, it will still be judged as incorrect because the description does not match the target features; if the preliminary text inference data accurately describes the color of the circular target, it will be judged as correct.
[0096] If the initial textual reasoning data is correct, the model will generate secondary textual reasoning data based on it. By reconfirming the content of the initial textual reasoning data or introducing more in-depth analytical logic (such as combining domain knowledge and associating with other information), the initial conclusions will be expanded and deepened, making the reasoning content more complete and logical.
[0097] If the initial text inference data is incorrect, the model will discard the erroneous conclusion and re-infer based on the target area marked by the bounding box coordinate data, combined with the requirements of the text data, to correct the erroneous description and supplement the necessary analysis, and finally generate the accurate secondary text inference data.
[0098] In this embodiment, during the process of regenerating the secondary text reasoning data, the accuracy of visual positioning is first verified, followed by the logicality of text reasoning. This ensures that the secondary text reasoning data relies on accurate image information and aligns with the core requirements of the text question, thus providing a reliable reasoning basis for subsequently generating the target answer. The entire process requires no manual intervention; the model autonomously completes verification and correction, improving the automation level of reasoning and the credibility of the results.
[0099] In one embodiment, such as Figure 7 As shown, steps S213-S214 precede step S211.
[0100] S213. Determine whether the bounding box coordinate data needs to be generated based on the image data and the text data;
[0101] S214. If generation is required, start the step of generating the bounding box coordinate data.
[0102] In this embodiment, before locating the target region based on the image data and text data, the inference model needs to determine whether it is necessary to generate the bounding box coordinate data. This process is an important prerequisite for improving inference efficiency and accuracy. Specifically, the model simultaneously performs joint analysis on the input image data and text data, determining whether the bounding box coordinate data needs to be generated by parsing their correlation. When a specific visual region needs to be referenced to support inference (such as counting, spatial relationship judgment, etc.), the generation of the bounding box coordinate data is initiated.
[0103] Specifically, firstly, the model performs semantic understanding on the text data, extracting the core requirements and inference objectives. If the text data contains instructions that require locating, marking, counting, spatial relationship judgment, or identifying specific objects in the image (such as "point out abnormal areas in the image" or "mark the location of an object"), it is preliminarily determined that the bounding box coordinate data may need to be generated. If the text data only involves queries about the overall features of the image (such as "describe the overall style of the image" or "determine whether the image is clear"), and does not involve the location information of specific objects, then the bounding box coordinate data may not need to be generated.
[0104] Subsequently, the model will verify this initial judgment by combining the image data. If the image data contains specific objects or regions that are clearly definite and consistent with the description in the text data, and the location information of these objects or regions is crucial for subsequent reasoning (e.g., the location of lesions in medical images, the coordinates of specific data in charts), then it is further confirmed that the bounding box coordinate data needs to be generated; if the image data does not contain a clear target object, or the text data requires reasoning that does not depend on a specific location (e.g., determining whether the brightness of the image is appropriate), then it is confirmed that the bounding box coordinate data does not need to be generated.
[0105] When the model determines that the bounding box coordinate data needs to be generated, it will immediately start the step of generating the bounding box coordinate data, namely steps S211-S212. Based on the target features in the image data and the positioning requirements of the text data, the target region is obtained, and the precise coordinate range is calculated and output according to the target region and the preset reference system. If it is determined that no generation is required, it will directly enter the subsequent preliminary inference stage to avoid unnecessary coordinate calculations occupying system resources.
[0106] Therefore, this embodiment ensures that the generation of bounding box coordinate data is only performed when necessary by closely combining the visual features of the image data with the semantic requirements of the text data. This provides accurate visual anchors for inference tasks that require localization and avoids the impact of invalid calculations on inference efficiency, making the entire inference process more efficient and in line with actual needs.
[0107] This application generates a reasoning chain containing the text reasoning data and the bounding box coordinate data through a preset reasoning model, enabling collaborative thinking between questions and images, improving the correlation between the reasoning chain and visual input, and thus enabling users to more accurately understand the reasoning process and the reasoning answer in scenarios that require logical reasoning based on image content, such as visual question answering, object counting, and spatial relationship judgment, thereby improving the authenticity, reliability, and accuracy of the reasoning; moreover, the reasoning model does not rely on prompting engineering or auxiliary modules during the reasoning process, thereby improving the applicability of the reasoning model.
[0108] Figure 8 This is a schematic block diagram of an image and text-based reasoning device 300 provided in an embodiment of the present invention. Figure 8 As shown, corresponding to the above-described image- and text-based reasoning method, the present invention also provides an image- and text-based reasoning apparatus 300. This image- and text-based reasoning apparatus 300 includes a unit for performing the above-described image- and text-based reasoning method, and the apparatus can be configured in a computer device. Specifically, please refer to... Figure 8 The image and text-based reasoning device 300 includes a receiving unit 301 and a generating unit 302.
[0109] The receiving unit 301 receives question data containing images and text input from the client, inputs the question data into a preset inference model, and receives a training set, which includes various types of training data.
[0110] The generation unit 302, wherein the reasoning model performs reasoning based on the question data to generate a reasoning chain, wherein the reasoning chain includes text reasoning data and bounding box coordinate data; the reasoning model generates a target answer based on the text reasoning data; performs preliminary reasoning based on the question data to obtain preliminary text reasoning data and bounding box coordinate data; performs secondary reasoning based on the question data, the preliminary text reasoning data, and the bounding box coordinate data to obtain secondary text reasoning data; obtains the target answer based on the secondary text reasoning data; locates a target region based on the image data and the text data; generates the bounding box coordinate data based on the target region and a preset reference frame; if correct, generates the secondary text reasoning data based on the preliminary text reasoning data; if incorrect, re-reasons based on the bounding box coordinate data and the text data to obtain the secondary text reasoning data; if generation is required, the step of generating the bounding box coordinate data is initiated.
[0111] In one embodiment, the receiving unit 301 includes a training unit.
[0112] The training unit trains the inference model based on the GRPO algorithm and each training data point; it performs inference on the training data and calculates rewards to obtain a composite reward value, wherein the reward calculation includes a true inference format reward calculation, an optional true target count reward calculation, and an answer accuracy reward calculation; the composite reward value is then used to optimize the inference model using population normalization advantage calculation and the GRPO objective function.
[0113] In one embodiment, the generation unit 302 includes a judgment unit.
[0114] The judgment unit determines whether the bounding box coordinate data and the target region match; if they match, it determines whether the preliminary text inference data is correct based on the bounding box coordinate data and the text data; and it determines whether the bounding box coordinate data needs to be generated based on the image data and the text data.
[0115] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned image and text-based reasoning device and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0116] The aforementioned image- and text-based reasoning device 300 can be implemented as a computer program, which can, for example... Figure 9 It runs on the computer device shown.
[0117] Please see Figure 9 , Figure 9 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.
[0118] See Figure 8 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0119] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform an image- and text-based reasoning method.
[0120] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0121] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an image and text-based reasoning method.
[0122] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0123] The processor 502 is used to run the computer program 5032 stored in the memory to implement the steps of the above-described image and text-based reasoning method.
[0124] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0125] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0126] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the steps of the above-described image- and text-based reasoning method.
[0127] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0128] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0129] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0130] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0131] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0132] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A reasoning method based on images and text, characterized in that, Applied to an inference platform, the method includes: Receive question data containing images and text from the client, and input the question data into a preset inference model; The inference model performs inference based on the question data to generate an inference chain, wherein the inference chain includes text inference data and bounding box coordinate data; The reasoning model generates the target answer based on the text reasoning data.
2. The method according to claim 1, characterized in that, The text reasoning data includes preliminary text reasoning data and secondary text reasoning data. The reasoning model performs reasoning based on the question data to generate a reasoning chain. The step of the reasoning chain including text reasoning data and bounding box coordinate data includes: Preliminary reasoning is performed based on the question data to obtain the preliminary text reasoning data and the bounding box coordinate data; The secondary text inference data is obtained by performing secondary inference based on the question data, the preliminary text inference data, and the bounding box coordinate data. The target answer is obtained based on the secondary text reasoning data.
3. The method according to claim 2, characterized in that, The problem data includes image data and text data. The step of performing preliminary inference based on the problem data to obtain the preliminary text inference data and the bounding box coordinate data includes: The target region is located based on the image data and the text data; The bounding box coordinate data is generated based on the target region and the preset reference system.
4. The method according to claim 3, characterized in that, The step of performing secondary inference based on the question data, the preliminary text inference data, and the bounding box coordinate data to obtain the secondary text inference data includes: Determine whether the bounding box coordinate data matches the target region; If a match is found, the correctness of the preliminary text inference data is determined based on the bounding box coordinate data and the text data. If correct, then generate the secondary text reasoning data based on the preliminary text reasoning data; If incorrect, the secondary text inference data is obtained by re-inferring based on the bounding box coordinate data and the text data.
5. The method according to claim 3, characterized in that, Prior to the step of locating the target region based on the image data and the text data, the following is included: Determine whether the bounding box coordinate data needs to be generated based on the image data and the text data; If generation is required, then initiate the step of generating the bounding box coordinate data.
6. The method according to claim 1, characterized in that, Before the step of receiving question data containing images and text input from the client and inputting the question data into a preset inference model, the following steps are included: Receive a training set, which includes various types of training data; The inference model is trained based on the GRPO algorithm and each of the training data.
7. The method according to claim 6, characterized in that, The steps for learning and training the inference model based on the GRPO algorithm and the training set include: The training data is used to perform inference and reward calculation to obtain a composite reward value, wherein the reward calculation includes a true inference format reward calculation, an optional true target count reward calculation, and an answer accuracy reward calculation; The inference model is optimized by calculating the group normalized advantage and using the GRPO objective function to calculate the composite reward value.
8. An image- and text-based reasoning device, applied to a reasoning platform, characterized in that, Includes a unit for performing the method as described in any one of claims 1-7.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, can implement the method as described in any one of claims 1-7.
Citation Information
Cited By
Image content auditing method and system based on multilayer scene graph structure
CN121861291A