Cross-modal knowledge reasoning method based on multi-modal large model

Through the improved multimodal cognitive framework and training strategies, the cross-modal understanding and reasoning capabilities of multimodal large models are improved, and the cross-modal data alignment and fusion problems are solved, achieving more accurate and interpretable inference results, which are suitable for multiple high-precision application fields.

CN120409639APending Publication Date: 2025-08-01SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510488858.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing multimodal large models have limited cross-modal understanding capabilities and insufficient inference accuracy, especially in cross-modal data alignment and fusion, long video analysis, dynamic scene understanding and human-computer interaction capabilities.

Method used

Using a multimodal cognitive framework based on Transformer, multimodal data preprocessing, improved rotational position embedding and visual encoder window attention mechanism, combined with rejection sampling and direct preference optimization strategies, multimodal neural network is trained to achieve cross-modal knowledge inference.

Benefits of technology

The model's information fusion and generalization capabilities in multimodal tasks are improved, making the inference results more accurate and interpretable, and are suitable for application areas with high precision requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409639A_ABST
    Figure CN120409639A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-modal knowledge reasoning method based on a multi-modal large model. In a cross-modal knowledge reasoning process, an existing model is usually limited by single-modal information extraction and shallow feature fusion, so that deep semantic association among data such as texts, images and videos is difficult to fully capture. In order to solve the problem, the invention provides a model for fusing multi-modal information such as texts, images, videos, documents and the like, and processing of multi-modal data is converted into unified feature extraction, interaction and deep reasoning tasks by fully utilizing a supervision fine tuning strategy, a self-adaptive attention mechanism and a cross-language processing technology. The model adopts a modular design, integrates multi-source data complementary analysis, spatial-temporal feature modeling and emotional semantic analysis, and realizes multi-modal collaborative interaction, dynamic scene understanding, long video key event analysis and man-machine co-emotional response. Through sufficient training, the multi-modal large model shows excellent logical reasoning ability and emotion understanding ability in a complex cognitive task, and a brand new solution is provided for efficient extraction, deep semantic analysis and intelligent response of cross-modal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal intelligent reasoning, and more specifically, to a cross-modal knowledge reasoning method based on a multimodal large model. The present invention utilizes computer deep learning, natural language processing, computer vision, and multimodal information fusion technologies to achieve the joint understanding, semantic parsing, and deep reasoning of multimodal data such as text, images, videos, and documents, thereby enhancing the intelligent level of complex cognitive tasks. Background Art

[0002] Multimodal intelligent technology is in a rapid development stage, bringing profound changes to the field of artificial intelligence. This trend is reflected in multiple aspects, including cross-modal data fusion, dynamic scene understanding, long video analysis, intelligent human-computer interaction, and the improvement of deep reasoning capabilities. Currently, the development of multimodal large models not only promotes the progress of general artificial intelligence but also demonstrates extensive application value in multiple fields such as healthcare, finance, education, law, and industry.

[0003] With the breakthrough of multimodal technology, computers can not only understand text information but also perform joint reasoning by combining various data such as images, videos, audio, and documents, thereby achieving more intelligent cognitive capabilities. For example, in the field of autonomous driving, intelligent systems need to simultaneously process camera images, lidar data, and voice commands to ensure the safe driving of vehicles; in medical image analysis, doctors not only need to view the patient's image data but also need to combine medical record text and gene sequencing data for comprehensive diagnosis. With the development of technology, multimodal cognitive models are gradually applied in multiple industries, improving the intelligent level and application value of artificial intelligence. Against this background, the research on multimodal large models continues to break through to enhance the capabilities of cross-modal data fusion, spatio-temporal feature modeling, and complex cognitive reasoning. Currently, mainstream multimodal methods mainly focus on directions such as cross-modal feature alignment, dynamic scene understanding, multimodal knowledge graph construction, and intelligent interaction optimization. For example, in the field of intelligent human-computer interaction, traditional natural language dialogue systems often rely only on text input, while multimodal interaction systems can combine information such as voice, expression, gesture, and image, making the dialogue more natural and realistic. In addition, intelligent video analysis is becoming one of the important fields of multimodal research. Long video content often contains rich information such as people, scenes, and time clues. How to efficiently analyze long videos and extract key information is a key issue in current research. Through deep learning and cross-modal fusion technologies, automatic summarization, event recognition, and structured storage of video content can be achieved, providing support for applications such as automatic news generation, intelligent monitoring, and content recommendation.

[0004] However, current multimodal large models still face many challenges. On the one hand, the alignment and fusion of cross-modal data are difficult, and the information expression methods of different modalities are different, making unified information modeling a technical difficulty. On the other hand, tasks such as long video analysis and dynamic scene understanding require in-depth modeling of spatio-temporal information, and the generalization ability of existing methods in complex scenarios is still limited. In addition, there is still much room for improvement in the empathy ability, logical reasoning ability, and cross-language adaptation ability of human-computer interaction. How to make multimodal large models more robust, interpretable, and capable of real-time response is one of the core issues in current research.

[0005] In recent years, with the development of large-scale pre-trained models, multimodal large models that combine technologies such as deep learning, knowledge reasoning, and natural language processing have gradually become the mainstream. By constructing large-scale labeled datasets, optimizing cross-modal feature alignment algorithms, and introducing domain knowledge to enhance reasoning ability, researchers have continuously promoted the progress of multimodal technologies. Especially in the fields of intelligent document parsing, long video understanding, cross-language knowledge transfer, etc., methods based on large models have shown significant advantages, laying a solid foundation for the development of multimodal intelligence. Summary of the Invention

[0006] The purpose of the present invention is to provide an efficient cross-modal knowledge reasoning method for the problems of limited cross-modal understanding ability and insufficient reasoning accuracy of current multimodal large models. This method realizes in-depth fusion of information, cross-modal knowledge transfer, and accurate reasoning by integrating multimodal data such as text, images, and videos, so as to improve the generalization ability and reasoning ability of the model in multimodal tasks.

[0007] The technical solution adopted by the present invention to achieve the above purpose is:

[0008] A cross-modal knowledge reasoning method based on a multimodal large model, comprising the following steps:

[0009] Collect multimodal data sources for systematic preprocessing, align and label different modality data according to predefined rules to form a joint dataset of knowledge reasoning question-answer pairs containing text, image, video, and document image information; and divide the dataset according to a ratio;

[0010] Establish a neural network model based on the general multimodal cognitive framework of Transformer, and improve the multimodal rotary position embedding and visual encoder window attention mechanism; use the joint optimization paradigm of rejection sampling and direct preference optimization to iteratively train the model, and fine-tune the parameters during the training process to enhance its knowledge acquisition ability and context relevance, and obtain an ideal model;

[0011] According to the query question input by the user, the ideal model is used to perform question-answer retrieval on the content to be identified and output the knowledge reasoning response corresponding to the user query.

[0012] The systematic pretreatment includes:

[0013] (1) For plain text data, a step-by-step cleaning and structured extraction method is used to obtain the semantically representative “premise-inference rule-conclusion” question-answer pairs;

[0014] (2) Perform image refinement on text-image data; through entity extraction → coreference calibration → relationship modeling → reasoning chain construction → natural language mapping, the image-text data is converted into question-answer pairs containing "cross-modal associations between images and text";

[0015] (3) For text-video data, through keyframe anchoring → text fine extraction → cross-modal association annotation → question semantic mapping, the text and visual information in the video are converted into question-answer pairs containing "video-text with cross-modal association". Each question-answer pair contains three elements: time anchor, text content, and visual association, which are used to support factual queries and reasoning questions.

[0016] (4) Unify the format and annotate structured document data, including but not limited to invoices, forms, and contracts → extract key information and convert it into "structured attribute rule-conclusion" question and answer pairs.

[0017] The steps to obtain question-answer pairs for plain text data are as follows:

[0018] Step-by-step cleaning: First, perform preliminary cleaning on the original text, including removing special symbols and redundant spaces and blank lines; second, keep the capitalization of proper nouns; finally, use the word segmentation tool to perform lexical segmentation;

[0019] Construct triple reasoning chains: extract triple reasoning chains through entity sharing or relationship transfer;

[0020] In terms of entity extraction, we first use the named entity recognition (NER) method to identify entity information in the text, and then combine manually predefined rules with the GPT-4o model to assist learning, explore the implicit logical relationships in the text and complete the annotation: "premise-inference rule-conclusion" question and answer pairs.

[0021] The steps to obtain question-answer pairs for text-image data are as follows:

[0022] Image refinement: image denoising and text region enhancement; using bounding box annotation to locate entity coordinates and distinguish between homonymous objects, while also recording the spatial relationships between entities;

[0023] Question-answer pair construction includes:

[0024] Text semantic structured extraction: Extract basic entity recognition, logical keywords, and quantifiers; logically split long text into clauses and annotate the causal or conditional reasoning type corresponding to each clause;

[0025] Entity-level coreference calibration: Establish a cross-modal entity mapping table to associate unique IDs with the same entity in images and text; annotate the multimodal attributes of entities;

[0026] Relationship and event association annotation: Extract explicit relationships from images and texts, and the intermediate conditions required for implicit reasoning; annotate the event sequence chain and environmental factors based on the scenario description task;

[0027] Construct reasoning chains in layers: break down complex reasoning into sub-steps, each of which includes premises, reasoning rules based on common sense or domain knowledge, and conclusions;

[0028] Natural language mapping: annotate entity appearance, action posture and spatial distance; define scene categories and functional attributes;

[0029] Question-answer pair construction: Through entity extraction → coreference calibration → relationship modeling → reasoning chain construction → natural language mapping, the image and text data are converted into question-answer pairs containing cross-modal associations, which are used to preserve visual image details and incorporate logical reasoning.

[0030] The steps to obtain question-answer pairs for text-video data are as follows:

[0031] Keyframe extraction: uniformly samples frames from long videos and removes consecutive similar frames based on content; records frame attributes: timestamp, video ID, time position in the original video, and frame type; marks frames containing text; video frame types include but are not limited to dialogue frames, action frames, and text display frames;

[0032] Text extraction and processing: First, use an OCR tool to extract visible text within the frame, annotating the text content and confidence level. Then, use bounding box coordinates to annotate the specific area of the text within the frame, distinguishing between multiple lines of text or scattered text blocks. Finally, annotate the text type by purpose. Text types include but are not limited to subtitles, slogans, graphic data, and object labels, which are used to associate text with video content.

[0033] Text-visual cross-modal association annotation: Perform position alignment to indicate whether the text is located within a certain object / area, or its spatial relationship with visual elements; if there is a logical association between text and visual elements, annotate the relationship type as "explanation", "indication", or "supplementation"; and annotate the semantic association between text and visual content; annotate the priority and record the text integrity for multi-language mixed text, dynamic subtitles, and occluded text; if the keyframe involves continuous action or the text subtitles change line by line, annotate the temporal relationship between frames; and desensitize text containing personal information or sensitive identifiers according to standard labeling.

[0034] Q&A Pair Construction: Through key frame anchoring → fine-grained text extraction → cross-modal correlation annotation → question semantic mapping, convert the text and visual information in the video into structured data that can be used for Q&A; each Q&A pair contains three elements: time anchor, text content, and visual association, which are used to support factual queries and reasoning questions.

[0035] Converting structured documents into Q&A pairs includes the following steps:

[0036] Format Unification: Use OCR tools to convert structured documents into editable text, preserving the original layout; for electronic documents in Word and PDF formats, use parsing libraries to extract text coordinates, font styles, and table structures; for scanned or photographed document images, use image processing tools to denoise, enhance text contrast, remove image skew, and unify image resolution and scale to a standard size.

[0037] Entity and Relationship Annotation: Annotate entities in forms, invoices, and contracts; at the same time, extract relationships between entities: the corresponding values in the rows and columns of the table; identify table headers and annotate the corresponding relationships between cell contents and rows / columns; record the row and column continuation relationships of continued tables when processing multi-page tables.

[0038] Q&A Pair Generation Rules: Direct answers need to be directly extracted from entity coordinates or table cells; or for indirect answers, calculation logic needs to be annotated and the logical calculation results of multiple entities or tables need to be integrated; or convert the titles and data of charts in the document into text descriptions to generate questions.

[0039] The Transformer model architecture designs multi-modal rotational position embeddings and a visual encoder window attention mechanism, including:

[0040] Window Attention Mechanism: Introduce a window attention mechanism in the visual encoder. By dividing the input image into local windows, the attention calculation is only performed within each window, rather than globally.

[0041] MRoPE Position Encoding with Absolute Time Alignment: Attach an absolute time vector to each video frame and combine it with relative position encoding to form a "absolute + relative" dual time representation.

[0042] The iterative training adopts the following two-stage optimization paradigm, including:

[0043] Rejection Sampling: Use an intermediate evaluation model to generate responses for the annotated dataset and compare the responses generated by the model with the correct annotated answers to filter out unsatisfactory outputs.

[0044] Direct preference optimization: Using image-text and plain text data to align the model with human preferences, for achieving deep fusion and semantic interaction of different modal information.

[0045] It also includes adopting strategies of deep semantic understanding, cross-modal contrastive learning, video content condensation and sentiment semantic analysis to iteratively train the model and fine-tune parameters during the training process to enhance its knowledge acquisition ability and context relevance.

[0046] The knowledge inference response corresponding to the user query includes structured knowledge, relational inference chains or text generation results to support multiple downstream tasks;

[0047] Parse the inference chain of the query statement input by the user, and extract the core elements in the inference process: premise conditions, logical rules, intermediate conclusions;

[0048] Combining the core elements in the inference process, use a symbolic inference engine to process mathematical symbols and scientific formulas, and automatically solve problems to obtain the structured knowledge conclusion of the image-text.

[0049] Through the structured relational inference chain, transform the structured recognition conclusion of the image-text into a human-readable natural language text generation result and output it to the user.

[0050] The present invention has the following beneficial effects and advantages:

[0051] 1. The present invention adopts an adaptive attention mechanism and a multi-head self-attention module to improve the information fusion ability of the model in multi-modal tasks and achieve cross-modal reasoning.

[0052] 2. The present invention enhances the generalization ability of the model through high-quality preprocessing and annotation of multi-modal data, making its performance more stable in different modal tasks.

[0053] 3. The present invention combines the rejection sampling and direct preference optimization strategies to make the inference results more accurate and conform to human preferences, improving the controllability and reliability of the model in inference tasks.

[0054] 4. The present invention adopts a structured output method, making the inference results more interpretable and applicable to multiple application fields with high-precision requirements such as medicine, finance, and law. Brief Description of the Drawings

[0055] Figure 1 It is the overall flowchart of the present invention;

[0056] Figure 2 It is the model structure diagram of the present invention. Detailed Embodiments

[0057] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific implementation methods of the present invention in conjunction with the accompanying drawings. A lot of specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the invention. Therefore, the present invention is not limited by the specific implementations disclosed below.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.

[0059] The present invention provides a cross-modal knowledge reasoning method based on a multi-modal large model, which realizes the deep fusion and logical reasoning of multi-modal data such as text, images, and videos through standardized data preprocessing rules, refined model training strategies, and interpretable reasoning output mechanisms. The method includes the following steps:

[0060] I. Construct a multi-modal knowledge reasoning data set

[0061] This step aims to collect and obtain relevant data from multi-modal data sources (including text, images, videos, documents, etc.) and perform systematic preprocessing on it to lay a foundation for subsequent knowledge reasoning tasks. The preprocessing work mainly includes the following aspects:

[0062] (1) Pure text data processing: Adopt a step-by-step cleaning and structured extraction strategy. First, perform preliminary cleaning on the original text, including removing special symbols (such as HTML tags, garbled characters) and redundant spaces and blank lines. In specific scenarios, retain the original capitalized form for content such as proper nouns. Subsequently, use a word segmentation tool for lexical segmentation. For example, jieba can be used for Chinese text, and nltk and other tools can be used for English text. An example of segmentation is: [Nature, magazine, published, an, article, about, AI, by, John, Smith, whose, team, is, from, MIT].

[0063] For the text processing of the triple inference chain, the logical structure needs to be focused on. The inference chain usually consists of multiple triples and forms a complete logical path through shared entities or transitive relationships. The triple form includes: (entity A, relationship, entity B) or (entity, attribute, attribute value), which is used to represent factual knowledge in the objective world. For example:

[0064] (Apple, belongs to, fruit)

[0065] (Earth, radius, 6371 km)

[0066] Reasoning chains are connected by entity sharing or relationship transitiveness. For example:

[0067] Premise Triad a: (Socrates, belongs to, human)

[0068] Premise triple b: (human, has attributes, mortal)

[0069] The inferred conclusion triple c: (Socrates, has attributes, is mortal)

[0070] Here, the shared entity "human" is used to connect the two premise triples, leading to the conclusion, forming a "2 premises + 1 conclusion" chain of reasoning. The formation of a chain of reasoning relies on the transitive nature of relationships, for example, the transitive law of the "belongs to" relationship: if A belongs to B, and B belongs to C, then it follows that A belongs to C.

[0071] For entity extraction, named entity recognition (NER) is first used to identify entities in the text. Then, manually written rules (such as "graduated from a certain place → permanently resides there") are combined with GPT-4o model-assisted learning to uncover implicit logical relationships in the text and complete the annotation process. For text containing complex logical structures (such as conditional sentences and multiple causal relationships), more detailed annotation is required.

[0072] For example:

[0073] Text: "If the weather is sunny and the temperature is suitable, outdoor activities will proceed as normal", marked with:

[0074] [The weather is sunny (condition 1) and the temperature is suitable (condition 2)] → [Outdoor activities (events) proceed normally, which is (result)]

[0075] The text "A is B's friend, B is C's colleague, so A and C have an indirect connection" is annotated as: [A(person), friend, B(person)] → [B(person), colleague, C(person)] → [A(person), indirect connection, C(person)]

[0076] Finally, GPT-4o is used to assist in identifying relationship transfer patterns, completing implicit premises (such as the universality of "humans have attributes"), and optimizing question formulation (such as converting "derive conclusion" into a natural language question). Ultimately, a question-answer pair of "premise-inference rule-conclusion" is formed, ensuring a clear and traceable logical chain. Quality control of the annotation results is performed through manual spot checks or cross-validation to ensure the accuracy and consistency of the annotation information.

[0077] (2) Text-Image Data Processing: The annotation of long inference data and scenario description tasks for text-image pairs needs to take into account cross-modal logical associations and natural language generation requirements, which can be specifically divided into the following core steps:

[0078] Image Refinement Processing: For inference requirements, denoise the image (e.g., use median filtering to remove salt-and-pepper noise) and enhance it (e.g., use adaptive threshold binarization to highlight the text area) to ensure that key entities (such as objects, text) are clearly distinguishable; when annotating bounding boxes, accurately locate the entity coordinates and distinguish objects with the same name but different meanings (e.g., the "apple" fruit and the "Apple" company logo are labeled separately), and at the same time record the spatial relationships between entities (e.g., "the cat is located above the sofa").

[0079] Semantic Structured Extraction of Text: In addition to basic entity recognition (person names, organization names), focus on extracting logical keywords (such as "because", "causes") and quantifiers (such as "more than", "all") to provide semantic support for the inference chain; split long texts into clauses logically (e.g., "Premise 1 → Premise 2 → Conclusion"), and label the corresponding inference type for each clause (such as causal relationship, conditional judgment).

[0080] Co-reference Calibration at the Entity Level: Establish a cross-modal entity mapping table to ensure that the same entity in the image and text (such as "boy") is associated with a unique ID, eliminating ambiguity (e.g., judging whether the text "apple" refers to the fruit or the company according to the image content); label the multimodal attributes of the entity (such as the consistency between the color and size of the "car" in the image and the text description).

[0081] Relationship and Event Association Annotation: Directly extract explicit relationships (such as "kick", "be located" co-displayed in text and image), and record the intermediate conditions required for inference for implicit logic (such as "too high temperature may trigger an alarm"); for scenario description tasks, additionally label the event chain (such as "the boy kicks the ball → the ball rolls → the puppy chases") and environmental elements (such as time, weather) to lay the foundation for generating natural language scenarios.

[0082] Hierarchical Construction of Inference Chain: Decompose complex inferences into atomic steps, each step including premises (text-image input), inference rules (common sense or domain knowledge), and conclusions (intermediate or final judgments). For example: Premise 1 "The temperature sensor in the image shows 35°C" + Premise 2 "Text rule 'When temperature > 30°C, the cooling system needs to be started'" → Conclusion "Start the cooling system", and label the modal input (image / text) and logical basis for each step.

[0083] Scenario description element refinement (natural language mapping): Annotate entity appearances (such as "blue shirt", "red football"), action postures (such as "kick", "run"), and spatial distances (such as "1 meter ahead", "0.5 meter to the right") to ensure the generation of detailed scene descriptions; Define scene categories (such as "outdoor sports", "medical emergency") and functional attributes (such as "entertainment", "alarm trigger") to provide context semantic guidance for the model.

[0084] Q&A pair construction: Through entity extraction → coreference calibration → relationship modeling → reasoning chain construction → natural language mapping, convert image-text data into Q&A pairs containing cross-modal associations, which not only retain visual details (such as colors, positions) but also incorporate logical reasoning (such as causality, conditions), suitable for training multi-modal models to understand complex scenarios (such as "Analyze the reasons and results of an event based on image and text descriptions"). The questions are manually written.

[0085] Example:

[0086] Fact-based: Directly map entity attributes (such as "The boy is wearing a blue shirt" corresponding to the color annotation in the image);

[0087] Inference-based: Integrate image-text relationships (such as "Why does the ball roll?" Answer: "Because the boy kicks the ball, and the force makes the ball move", combining image actions and physical common sense).

[0088] (3) Text-video data processing: The core steps for key frame extraction and text-visual information alignment annotation in text-video data processing are as follows:

[0089] Key frame extraction: Uniformly sample long videos at fixed time intervals (such as 1 frame every 5 seconds), and combine content screening to avoid redundancy (such as removing consecutive similar frames). For complex scenarios (such as fast movement, multi-object interaction), annotators manually supplement or adjust the key frames extracted by the machine to ensure coverage of core events, object states, or frames where text appears. Then record the frame timestamp, the video ID it belongs to, and its time position in the original video (such as "Frame 2 in the scene from 00:12:34 to 00:12:36"), and at the same time annotate the frame type (such as dialogue frame, action frame, text display frame), with a focus on marking frames containing recognizable text (such as subtitles, logos, handwritten content).

[0090] Text extraction and processing: First, use an OCR tool (such as Tesseract, Baidu AI Open Platform) to extract all visible text within the frame, mark the text content and confidence level (filter out text with low confidence level, and manually correct blurred or misrecognized content). Then, use the bounding box coordinates (x1, y1, x2, y2) to mark the specific area of the text in the frame (the upper left and lower right coordinates), and distinguish multi-line text or scattered text blocks. Finally, mark the text type according to its use (such as subtitles, slogans, chart data, object labels), and clarify the association between the text and the video content (such as "subtitles of the dialogue of the characters in the picture", "text on the background billboard").

[0091] Text-visual association annotation: Implement position alignment to explain whether the text is located within a specific object / area (such as "the text 'No Parking' is located on the red warning sign on the right side of the picture"), or the spatial relationship with visual elements (such as "the subtitles are located at the bottom of the picture, corresponding to the speaking actions of the characters"). If there is a logical association between the text and the visual elements (such as the text indicates an object, explains an action), mark the type of relationship ("explanation", "indication", "supplement"), for example, "the paper held by the character in the picture reads 'Meeting Agenda', corresponding to the current meeting scene". At the same time, mark the semantic association between the text and the visual content (such as "the data text in the chart corresponds to the values of the bar chart in the picture", "the slogan text describes the main event in the picture"). For mixed-language, dynamic text (such as scrolling subtitles), and occluded text, it is necessary to clearly mark the priority (such as giving priority to marking the clearly visible part), and record the text integrity (such as "the two characters 'Safety' blocked by the character"). If the key frame involves continuous actions or text changes (such as subtitles appearing line by line), supplement the annotation of the inter-frame timing relationship (such as "the text in the 3rd frame is the subsequent content of the text in the 2nd frame"). For text containing personal information and sensitive identifiers, desensitize it according to the specification (such as blurring or replacing it with a general label).

[0092] Refinement of scenario description elements (problem semantic mapping): Mark the entity appearance (such as "blue shirt", "red football"), action postures (such as "kick", "run"), and spatial distances (such as "1 meter ahead", "0.5 meter to the right") to ensure the generation of a detailed scenario description; define the scenario category (such as "outdoor sports", "medical emergency") and functional attributes (such as "entertainment", "alarm trigger") to provide context semantic guidance for the model.

[0093] Construction of question-answer pairs: Through key frame anchoring → text fine extraction → cross-modal association annotation → problem semantic mapping, convert the text and visual information in the video into structured data that can be questioned and answered. Each question-answer pair contains three elements: time anchor, text content, and visual association, which supports both fact-based queries (such as "the subtitle content at a certain moment") and reasoning-based questions (such as "how does the text explain the action in the picture"), providing multi-dimensional training data for the video understanding model. The construction of questions mainly uses manual and machine assistance.

[0094] Example:

[0095] Factual questions: directly map the OCR text content and its location (e.g., "What is the subtitle content?" → extract the subtitle text within the frame).

[0096] Inferential questions: integrate the text-visual relationship (e.g., "Which object in the picture corresponds to the 'Emergency Exit' sign?" → associate with the green arrow icon on the left side of the picture).

[0097] Case question-answer pairs:

[0098] Question: "What text is on the paper held by the person in the 00:05:10 frame?"

[0099] Answer: "The paper reads 'Meeting Agenda', located slightly to the left of the center of the picture, corresponding to the scene where the person is explaining the meeting process (based on text position annotation and visual action association).

[0100] (4) Document parsing and preprocessing: The following is a detailed description of the data annotation work for structured documents such as invoices, forms, and contracts, covering the annotation process, key information extraction methods, and quality control points:

[0101] Format standardization: Convert to editable text through OCR tools (such as ABBYY FineReader, Google Vision), retaining the original layout (such as table borders, paragraph indents). For electronic documents in Word or PDF formats, use parsing libraries (PyPDF2, Docx2Text in Python) to extract text coordinates, font styles (distinguish headings / text), and table structures (recognize merged cells, multi-page tables). For scanned or photographed document images (such as invoice scans, contract photos), use image processing tools (such as Photoshop, online tool Fotor) to denoise (remove scanning stripes, shooting shadows), enhance contrast (make the text clearer). If the image is tilted (such as a contract photographed with a mobile phone not horizontally), rotate and correct it through an image editing tool (such as the 'Auto-correct horizontal line' function). Unify the image resolution (such as 300 dpi), and scale to the standard size (such as 2480×3508 pixels for A4 document images) to avoid field positioning deviations caused by size differences.

[0102] Entity and Relationship Annotation: For forms, invoices, etc., identify the entities therein, such as invoice numbers and amounts in invoices, dates and parties in contracts, etc. At the same time, extract the relationships between these entities, such as the corresponding values of rows and columns in a table. Identify the table headers (such as the "Commodity Name", "Quantity", and "Unit Price" columns of an invoice), and annotate the corresponding relationships between the cell content and the rows and columns (such as "the cell in the 2nd row and 3rd column corresponds to the 'Unit Price' column, and the value is '100 yuan'"). When processing multi-page tables, record the row and column continuation relationships of the continued table (such as "the 5th row of Contract Table 3-1 continues the 4th row of Table 3-1 on Page 4").

[0103] Q&A Pair Generation Rules: Direct answers need to be directly extracted from entity coordinates or table cells (such as invoice numbers corresponding to OCR text content, and contract terms corresponding to the original text of paragraphs). Example: For the question "What is the invoice number?", the answer is "NO.20231234" extracted by OCR, and the source is marked as "the coordinate area in the upper right corner of the invoice (400,50,500,80)". Indirect answers need to integrate multiple entities or perform table calculations (such as "Total Amount = Quantity × Unit Price", "Liquidated Damages of the Contract = Total Contract Amount × 0.1%"), Example: For the question "What is the total contract amount?", the answer needs to sum up the "Amount" column in all rows of the table, and the calculation logic is marked as "SUM(all values in the 3rd column of the table)". If the document contains charts (such as the line chart of amounts in an invoice), convert the chart title and data labels into text descriptions and then generate questions (such as "What is the sales amount in Q3 of 2023 in the line chart?", and the answer extracts the corresponding data points of the chart). For multi-hop reasoning Q&A pairs, it means that the answer needs to be derived through the logical integration of 2 or more independent document information units, rather than directly extracting a single sentence or field. By manually combining with GPT4-o for the input document fragments, potential questions are generated, and finally, it is manually verified to check whether multiple premises come from the same document and are logically related (such as whether "Clause A" and "Clause B" belong to the same contract).

[0104] Example:

[0105]

Question

[0106]

Premise Conditions

[0107] 1. Clause 4.5: "If the delay exceeds 60 days, Party A has the right to terminate the contract" (Page 8);

[0108] 2. Clause 3.2: "Pay liquidated damages at 0.05% of the total contract amount per day for the delay" (Page 7).

[0109]

Reasoning Process

[0110] 1. The 65-day delay exceeds the 60-day threshold of Clause 4.5, meeting the condition for terminating the contract;

[0111] 2. The overdue fact simultaneously triggers the liquidated damages rule in Clause 3.2;

[0112] 3. There is no conflict between the two, and both can be claimed simultaneously.

[0113]

Answer

[0114] (5) Dataset division: Divide the data into training set, validation set and test set according to the ratio of 7:1:2 to ensure the generalization ability of the model.

[0115] II. Model Architecture Design

[0116] Window Attention Mechanism: Introduce the window attention mechanism in the visual encoder. By dividing the input image into local windows (such as sub-regions of 7x7 pixels), the attention calculation is only carried out within each window instead of the global scope. This lightweight design significantly reduces the computational complexity (the complexity is reduced from O(N2) to O(M2×N / M2), where N is the number of global pixels and M is the number of window pixels), while maintaining the accurate capture of local semantics. The feature aggregation of repetitive textures (such as checkerboards, grids) and dense text areas (such as document tables) is more efficient, avoiding the redundant calculation of global attention.

[0117] MRoPE with Absolute Time Alignment: Attach an absolute time vector to each video frame (such as converting "00:12:34" into a time encoding sequence), and combine it with the relative position encoding (such as the adjacent frame interval) to form a "absolute + relative" dual time representation. By reducing the number of video input tokens + absolute time alignment, the processing cost of the model for long videos (such as 2-hour conference videos) is reduced by 60%, while maintaining the event localization accuracy (the second-level error rate < 3%). Through the three-dimensional component design of MRoPE → differential encoding of input modalities → improvement of absolute time alignment, a unified position representation system covering text, images, and videos is achieved.

[0118] III. Model Training, Validation and Testing

[0119] Use a multi-modal large model based on the Transformer architecture, and perform fine-tuning and training on the multi-modal dataset to enable it to perform cross-modal knowledge reasoning. The steps include:

[0120] (1) Loading the pre-trained model: Use an existing multi-modal pre-trained model as the base model and load the weight parameters pre-trained on a large-scale text-image / video corpus to ensure the basic feature extraction capabilities (such as image semantic representation and text context understanding). For specific tasks (such as document parsing and video reasoning), selectively freeze the parameters of the underlying vision / language encoders (the first 6 layers) and only fine-tune the high-level interaction module to balance the transfer efficiency and task-specific optimization. For structured modalities such as tables and formulas, design hybrid position embeddings (such as two-dimensional coordinate encoding + sequence index) to ensure the compatibility of position information for different modal inputs (such as the row and column positions of table cells being associated with the order of text paragraphs).

[0121] (2) Contrastive learning: Construct positive and negative sample pairs (such as matching / mismatching text-image pairs), and use the InfoNCE loss function to enhance the semantic alignment accuracy between modalities (such as distinguishing cross-modal ambiguities between the "apple" fruit and the "Apple" company logo). Generate mixed-modal samples (such as "real image + generated text" combinations) through linear interpolation to expand the modal diversity of the training data. Adopt FP16 mixed-precision optimization to reduce the GPU memory usage while maintaining the accuracy and support larger batch training (such as a 50% increase in batch size on an 8-GPU cluster).

[0122] (3) Rejection sampling strategy to enhance the quality of CoT data: On the labeled dataset, use the current version of the model to generate CoT reasoning steps (such as step-by-step calculations for math problems and logical derivations of contract terms). Introduce a more powerful evaluation model (a reward model trained with the Qwen-VL series) to evaluate the intermediate conclusions and real annotations during model training (such as intermediate numerical values in steps and the correctness of clause references). Only retain samples with all steps correct (such as each step of the math calculation being correct and each step of the VQA accurately referencing visual information). Use a rule engine to eliminate overly long (number of steps > 15) or repetitive (such as circularly referencing the same modal information) CoT results to ensure that the reasoning chain is concise and effective (the average number of steps is controlled within 5 - 10 steps). The labeled dataset is evaluated by a reward model trained with the Qwen-VL series to obtain high-quality data question-answer pairs, which are used for model fine-tuning and cross-modal knowledge reasoning tasks.

[0123] Example of the rejection sampling process:

[0124] The model generates CoT steps (intermediate version model):

[0125] Step 1: Since this is a parking lot (no warning signs are mentioned, relying only on text common sense and not using visual information);

[0126] Step 2: Parking is allowed in the parking lot, which is a contradiction and cannot be answered (wrong).

[0127] Verification and rejection:

[0128] Reason for rejection: In Step 1, the "No Parking" warning sign in the image is not mentioned, and the visual information is not fully utilized;

[0129] Processing: The "modal integration degree" verification fails to trigger, and the sample is rejected.

[0130] Examples of qualified samples (passed the screening):

[0131] Step 1: There is a red "No Parking" warning sign in the upper right corner of the image (locating the visual entity);

[0132] Step 2: According to the warning sign, parking is prohibited here (combining visual information with rules, retained).

[0133] (4) Human preference alignment: Construct a preference dataset through crowdsourcing annotation (such as "clearer reasoning steps", "more concise answer expression"), and use the cross-entropy loss function to make the model output distribution converge to human preferences (such as the interpretability score of the answer is increased by 40%).

[0134] IV. Result Output and Application

[0135] The structured knowledge obtained through reasoning can be organized and output in the following form to support different application scenarios:

[0136] Technical implementation path: Reasoning chain analysis: Extract the core elements in the reasoning process (such as preconditions, logical rules, intermediate conclusions), for example: Mathematical reasoning: "Step 1: According to the triangle interior angle sum theorem, ∠A + ∠B + ∠C = 180°; Step 2: Given ∠A = 60°, ∠B = 70°, so ∠C = 180° - 60° - 70° = 50°." Legal derivation: "Based on Article 563 of the Civil Code (Premise 1), Party B's overdue delivery exceeds 30 days (Premise 2), which meets the conditions for contract termination (Conclusion)." Educational scenario: "Question: Calculate the area of a rectangle. Reasoning process: First measure the length and width, which are 10 cm and 5 cm respectively; then apply the area formula (length × width), that is 10 × 5 = 50 cm²; finally, the area is obtained as 50 square centimeters." Customer service scenario: "According to the order information you provided (tail number 1234), your package has exceeded the expected delivery time by 3 days (Premise), according to Article 4. of the logistics agreement (Rule), we will apply for a compensation for delayed delivery for you (Conclusion)." Image reasoning: "The red warning sign in the picture (coordinates X = 100, Y = 200) shows 'No Parking' (visual entity), according to the traffic regulations (Rule), this area belongs to the no-parking zone (Conclusion)."

[0137] Automatic Problem Solving: Symbolic Reasoning Engine: Supports the formal processing of mathematical symbols (such as integrals, matrices) and scientific formulas (such as physical laws, chemical equations). For example: Math problem: "Solve the equation 2x + 5 = 15" → Automatically deduce "2x = 15 - 5 → x = 10 ÷ 2 → x = 5"; Physics problem: "Given force F = ma, m = 2 kg, a = 3 m / s², find F" → Automatically substitute into the formula to calculate "F = 2 × 3 = 6 N". Multi-step Logical Verification: Decompose complex problems into steps and verify intermediate results. For example: Geometry proof problem: "Prove triangle congruence" → Automatically match congruence theorems (SSS / SAS / ASA) and verify whether each step condition is satisfied; Programming problem: "Write a bubble sort algorithm" → Generate a code framework and verify the logical correctness (such as whether the inner loop correctly compares adjacent elements and whether the outer loop controls the number of sorting times).

[0138] Natural Language Generation: Enhanced Transparency and Interpretability of the Reasoning Process: By converting structured reasoning chains (such as triple logic, multi-hop steps, modal associations) into human-readable natural language text, the model can explain the origin of the conclusion with a clear logical chain, meeting the interpretability requirements of scenarios such as education, customer service, and decision support.

[0139] The present invention also provides a system architecture:

[0140] An inference system based on a multi-modal large model, including a processor and a memory. The following program modules are stored in the memory and perform inference tasks when the processor loads the program; The program modules include:

[0141] Data Processing Module: Preprocess multi-modal data and construct a multi-modal data set;

[0142] Multi-modal Cognitive Modeling and Optimization Training Module: Establish a multi-modal inference network based on Transformer, and use the labeled data set for iterative training to enhance the model's cross-modal understanding and reasoning ability;

[0143] Inference Result Output Module: Map user queries to structured inference results, including knowledge graphs, inference chains, or text generation, to support various application requirements.

[0144] Based on multimodal information fusion and deep semantic understanding, the present invention proposes a multimodal large model and its method that support cross-modal knowledge reasoning. By adopting a cross-modal Transformer structure, the ability of multimodal feature alignment is enhanced, and the generalization and adaptability of the model are improved. Through unified multimodal data modeling, dynamic scene reasoning, long video intelligent analysis, human-computer empathy interaction, and cross-language adaptability optimization, the present invention aims to solve the deficiencies in the prior art and provide an efficient and accurate solution for multimodal intelligent reasoning. The above are only the preferred embodiments of the present invention and do not impose any limitations on the present invention. Any simple modifications, changes, and equivalent structural changes made to the above embodiments according to the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A cross-modal knowledge reasoning method based on a multi-modal large model, characterized in that, The following steps are involved: Collect multimodal data sources and perform systematic preprocessing. Align and annotate data from different modalities according to predefined rules to form a joint dataset of knowledge reasoning question-answer pairs containing text, images, videos, and document image information. The dataset is then divided proportionally. A neural network model based on a general multimodal cognitive framework of Transformer was established, with improvements to the multimodal rotation position embedding and the visual encoder window attention mechanism. The model was iteratively trained using a joint optimization paradigm of rejection sampling and direct preference optimization, with parameter fine-tuning during training to enhance its knowledge acquisition and contextual relevance, resulting in an ideal model. According to the query question input by the user, the ideal model is used to perform question-answer retrieval on the content to be identified and output the knowledge reasoning response corresponding to the user query.

2. The cross-modal knowledge reasoning method based on a multi-modal large model according to claim 1, wherein, The systematic pretreatment includes: (1) For plain text data, a step-by-step cleaning and structured extraction method is used to obtain the "premise-inference rule-conclusion" question-answer pairs that represent semantics; (2) Perform image refinement on text-image data; through entity extraction → coreference calibration → relationship modeling → reasoning chain construction → natural language mapping, the image-text data is converted into question-answer pairs containing "cross-modal associations between images and text"; (3) For text-video data, through keyframe anchoring → text fine extraction → cross-modal association annotation → question semantic mapping, the text and visual information in the video are converted into question-answer pairs containing "video-text with cross-modal association". Each question-answer pair contains three elements: time anchor, text content, and visual association, which are used to support factual queries and reasoning questions. (4) Unify the format and annotate structured document data, including but not limited to invoices, forms, and contracts → extract key information and convert it into "structured attribute rule-conclusion" question and answer pairs.

3. A cross-modal knowledge reasoning method based on a multi-modal large model according to claim 2, characterized in that, The steps to obtain question-answer pairs for plain text data are as follows: Step-by-step cleaning: First, perform preliminary cleaning on the original text, including removing special symbols and redundant spaces and blank lines; second, keep the capitalization of proper nouns; finally, use the word segmentation tool to perform lexical segmentation; Construct triple reasoning chains: extract triple reasoning chains through entity sharing or relationship transfer; In terms of entity extraction, we first use the named entity recognition (NER) method to identify entity information in the text. Then, we combine manually predefined rules with the GPT-4o model to assist in learning, explore the implicit logical relationships in the text, and complete the annotation of the "premise-inference rule-conclusion" question and answer pairs.

4. A cross-modal knowledge reasoning method based on a multi-modal large model according to claim 2, characterized in that The steps to obtain question-answer pairs for text-image data are as follows: Image refinement: image denoising and text region enhancement; using bounding box annotation to locate entity coordinates and distinguish between homonymous objects, while also recording the spatial relationships between entities; Question-answer pair construction includes: Text semantic structured extraction: Extract basic entity recognition, logical keywords, and quantifiers; logically split long text into clauses and annotate the causal or conditional reasoning type corresponding to each clause; Entity-level coreference calibration: Establish a cross-modal entity mapping table to associate unique IDs with the same entity in images and text; annotate the multimodal attributes of entities; Relationship and event association annotation: Extract explicit relationships from images and texts, and the intermediate conditions required for implicit reasoning; annotate the event sequence chain and environmental factors based on the scenario description task; Construct reasoning chains in layers: break down complex reasoning into sub-steps, each of which includes premises, reasoning rules based on common sense or domain knowledge, and conclusions; Natural language mapping: annotate entity appearance, action posture and spatial distance; define scene categories and functional attributes; Question-answer pair construction: Through entity extraction → coreference calibration → relationship modeling → reasoning chain construction → natural language mapping, the image and text data are converted into question-answer pairs containing cross-modal associations, which are used to preserve visual image details and incorporate logical reasoning.

5. A cross-modal knowledge reasoning method based on a multi-modal large model according to claim 2, characterized in that The steps to obtain question-answer pairs for text-video data are as follows: Keyframe extraction: uniformly samples frames from long videos and removes consecutive similar frames based on content; records frame attributes: timestamp, video ID, time position in the original video, and frame type; marks frames containing text; video frame types include but are not limited to dialogue frames, action frames, and text display frames; Text extraction and processing: First, use an OCR tool to extract visible text within the frame, annotating the text content and confidence level. Then, use bounding box coordinates to annotate the specific area of the text within the frame, distinguishing between multiple lines of text or scattered text blocks. Finally, annotate the text type by purpose. Text types include but are not limited to subtitles, slogans, graphic data, and object labels, which are used to associate text with video content. Text-visual cross-modal association annotation: Position alignment is performed to indicate whether the text is located within a certain object / area, or its spatial relationship with visual elements. If there is a logical association between text and visual elements, the relationship type is annotated as "explanation," "indication," or "supplementation." The semantic association between text and visual content is also annotated. For mixed language text, dynamic subtitles, and obscured text, priority is annotated and text integrity is recorded. If keyframes involve continuous action or text subtitles change line by line, the temporal relationship between frames is additionally annotated. Text containing personal information or sensitive identifiers is desensitized according to standard labels. Question-answer pair construction: Through keyframe anchoring → fine text extraction → cross-modal association annotation → question semantic mapping, the text and visual information in the video are converted into structured data that can be asked and answered. Each question-answer pair contains three elements: time anchor, text content, and visual association, which are used to support factual queries and reasoning questions.

6. A cross-modal knowledge reasoning method based on a multi-modal large model according to claim 2, characterized in that, Converting structured documents into question-answer pairs involves the following steps: Format standardization: Use OCR tools to convert structured documents into editable text, preserving the original layout. For electronic documents in Word and PDF formats, use parsing libraries to extract text coordinates, font styles, and table structures. For scanned or photographed document images, use image processing tools to remove noise and enhance text contrast. Remove image tilt. Standardize image resolution and scale to a standard size. Entity and relationship annotation: Annotate entities in forms, invoices, and contracts; extract relationships between entities: the corresponding values of rows and columns in a table; identify table headers and annotate the correspondence between cell content and rows and columns; record the continuation relationship of rows and columns when processing cross-page tables. Q&A pair generation rule: The direct answer needs to be directly extracted from the entity coordinates or table cells; or for the indirect answer, the calculation logic needs to be marked and the logical calculation results of multiple entities or tables need to be integrated; or the titles and data in the charts in the document are converted into text descriptions to generate questions.

7. A cross-modal knowledge reasoning method based on a multi-modal large model according to claim 1, characterized in that, The Transformer model architecture designs multi-modal rotary position embedding and the visual encoder window attention mechanism, including: Window attention mechanism: Introduce the window attention mechanism in the visual encoder. By dividing the input image into local windows, the attention calculation is only carried out within each window, rather than in the global scope. MRoPE position encoding with absolute time alignment: Attach an absolute time vector to each video frame and combine it with the relative position encoding to form a "absolute + relative" dual-time representation.

8. A cross-modal knowledge reasoning method based on a multi-modal large model according to claim 1, characterized in that The iterative training adopts the following two-stage optimization paradigm, including: Rejection sampling: Use an intermediate evaluation model to generate responses for the labeled dataset, and compare the responses generated by the model with the correct answers in the labels to filter out unsatisfactory outputs. Direct preference optimization: Use image-text and pure text data to align the model with human preferences for achieving in-depth fusion and semantic interaction of different modal information.

9. A cross-modal knowledge reasoning method based on a multi-modal large model according to claim 1, characterized in that It also includes adopting strategies such as deep semantic understanding, cross-modal contrast learning, video content condensation, and emotional semantic analysis to iteratively train the model and fine-tune the parameters during the training process to enhance its knowledge acquisition ability and context relevance.

10. A cross-modal knowledge reasoning method based on a multi-modal large model according to claim 1, characterized in that, The output of the knowledge inference response corresponding to the user query includes structured knowledge, relationship inference chains, or text generation results to support multiple downstream tasks. Parse the inference chain of the query statement input by the user and extract the core elements in the inference process: preconditions, logical rules, and intermediate conclusions. Combine the core elements in the inference process and use a symbolic inference engine to process mathematical symbols and scientific formulas to automatically solve problems and obtain the structured knowledge conclusions of the image-text. Convert the structured recognition conclusions of the image-text into human-readable natural language text generation results through the structured relationship inference chain and output them to the user.

Citation Information

Patent Citations

  • Rich semantic dialogue generation method fusing visual situation

    CN115964467A

  • Tuning generative models using latent variable inference

    CN118468868A

Cited By

  • Expressway scene-oriented interpretable hierarchical reasoning multi-modal method and system

    CN120822624A

  • Heterogeneous data conversion method and system based on multi-modal large model

    CN120973851A

  • Fine-grained visual target recognition expert knowledge agent generation method and device

    CN121009989A

  • A fine-grained visual target recognition expert knowledge intelligent agent generation method and device

    CN121009989B

  • Bidding field information extraction method and device

    CN121074929A