Decision-making method and device based on visual thinking chain, equipment and medium
By constructing a multimodal data processing method in the fields of fintech and healthcare, a visualized thought chain is generated, which solves the problem of lack of consistency in joint reasoning of text semantics and image spatial information, realizes a more transparent and interpretable decision-making process, and improves accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies cannot effectively combine textual semantics and image spatial information for joint reasoning in the fields of fintech and healthcare, resulting in a lack of consistency, transparency, and interpretability between the language reasoning chain and the visual reasoning trajectory.
By receiving multimodal data, performing feature extraction and fusion processing, generating multimodal fusion features, performing semantic and spatial joint encoding, constructing a visual thought chain of language reasoning chain and visual reasoning trajectory, and performing adaptive constraint alignment, finally generating interpretable decision output.
It enables more consistent, transparent and traceable decision-making processes in the fintech and healthcare sectors, and improves the accuracy and interpretability of reasoning results in multimodal scenarios.
Smart Images

Figure CN121834705A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent decision-making technology, and in particular to a decision-making method, apparatus, device, and medium based on a visual thinking chain. Background Technology
[0002] In the fintech sector, current intelligent risk control, underwriting, and claims systems typically rely on textual data or structured data to perform risk assessment, rule matching, and decision generation. While large language modeling technology has been gradually applied to text understanding in questionnaire parsing, medical record review, and claims scenarios in recent years, these systems generally treat text and images as independent information sources, lacking the ability to perform deep fusion and reasoning between the two. In business scenarios involving spatial orientation, object states, or scene changes, relying solely on text models makes it difficult to establish a unified semantic and spatial understanding. For example, in analyzing collision direction in auto insurance claims or identifying crop lodging in agricultural insurance cases, existing technologies often cannot simultaneously combine textual descriptions and image content to establish consistent spatial inference results. This prevents the identification of potential contradictions between cross-modal information, reducing the accuracy of judgments.
[0003] In the healthcare field, existing intelligent diagnostic assistance, clinical review, and health assessment systems largely employ text-based reasoning chains, enhancing judgment through layer-by-layer linguistic reasoning. However, these mechanisms excel at handling structured or semi-structured linguistic logic, often using independent image models to judge spatial relationships and semantic cues in image data, lacking a unified reasoning framework across modal information. When interpreting the correspondence between medical images and disease progression descriptions, existing systems are prone to a disconnect between semantic and spatial understanding, failing to reliably capture spatiotemporal features such as symptom evolution or changes in lesion areas. Furthermore, because the reasoning process unfolds implicitly within the model, clinicians struggle to obtain clear inferences from the system output, impacting the interpretability and traceability of results. Summary of the Invention
[0004] The main objective of this invention is to provide a decision-making method, apparatus, device, and storage medium based on a visual thinking chain, aiming to solve the technical problem that existing technologies cannot achieve joint reasoning of text semantics and image spatial information at the reasoning level, resulting in a lack of consistency, transparency, and interpretable decision output between the language reasoning chain and the visual reasoning trajectory.
[0005] To achieve the above objectives, this invention provides a decision-making method based on a visual thinking chain, comprising: Receives multimodal data, including text and image data; Feature extraction and fusion processing are performed on the text data and image data to obtain multimodal fusion features; The multimodal fusion features are subjected to joint semantic and spatial encoding to obtain a joint feature representation; Based on the joint feature representation, a visual thought chain including a language reasoning chain and a visual reasoning trajectory is generated. Adaptive constraint alignment processing is performed on the language reasoning chain and visual reasoning trajectory of the visualized thinking chain to obtain the aligned visualized thinking chain. Based on the aligned visual thought chain, an interpretable decision output is generated.
[0006] Furthermore, to achieve the above objectives, the present invention provides a decision-making device based on a visual thinking chain, comprising: A multimodal input module is used to receive multimodal data, including text data and image data; The multimodal feature fusion module is used to extract and fuse features from the text data and image data to obtain multimodal fused features; The joint encoding module is used to perform semantic and spatial joint encoding processing on the multimodal fusion features to obtain a joint feature representation; The thought chain generation module is used to generate a visual thought chain, including a language reasoning chain and a visual reasoning trajectory, based on the joint feature representation. The thought chain alignment module is used to perform adaptive constraint alignment processing on the language reasoning chain and visual reasoning trajectory of the visualized thought chain to obtain the aligned visualized thought chain. The decision generation module is used to generate interpretable decision outputs based on the aligned visual thought chain.
[0007] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a decision-making program based on a visual thinking chain stored in the memory and executable on the processor, wherein when the decision-making program based on a visual thinking chain is executed by the processor, it implements the steps of the decision-making method based on a visual thinking chain as described above.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a non-volatile computer-readable storage medium storing a decision-making program based on a visual thinking chain, wherein when the decision-making program based on the visual thinking chain is executed by a processor, it implements the steps of the decision-making method based on the visual thinking chain as described above.
[0009] Beneficial Effects: This invention relates to the field of intelligent decision-making technology and can be applied to business scenarios such as fintech and healthcare. It discloses a decision-making method, apparatus, device, and medium based on a visual thinking chain, comprising: receiving multimodal data of text and image data; performing feature extraction and fusion processing to generate multimodal fusion features; performing semantic and spatial joint encoding processing on the multimodal fusion features to obtain a joint feature representation; generating a visual thinking chain composed of a linguistic reasoning chain and a visual reasoning trajectory based on the joint feature representation; performing adaptive constraint alignment processing on the linguistic reasoning chain and the visual reasoning trajectory to obtain an aligned visual thinking chain; and generating an interpretable decision output based on the aligned visual thinking chain. This invention, by constructing a joint expression system of linguistic reasoning chains and visual reasoning trajectories and performing adaptive constraint alignment processing on both, enables the model to understand semantic and spatial relationships within a unified reasoning structure, thereby achieving a more consistent, transparent, and traceable decision-making process and improving the accuracy and interpretability of reasoning results in multimodal scenarios. Attached Figure Description
[0010] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a decision-making method based on a visual thinking chain according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the decision-making method based on a visual thinking chain according to the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the decision-making device based on a visual thinking chain of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0011] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0012] The decision-making method based on visual thinking chains provided in this invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can receive multimodal data (text and image data) from the client, perform feature extraction and fusion processing to generate multimodal fusion features; perform semantic and spatial joint encoding processing on the multimodal fusion features to obtain joint feature representations; generate a visual thought chain composed of a language inference chain and a visual inference trajectory based on the joint feature representations; perform adaptive constraint alignment processing on the language inference chain and the visual inference trajectory to obtain an aligned visual thought chain; and generate an interpretable decision output based on the aligned visual thought chain. This invention constructs a joint expression system of language inference chain and visual inference trajectory, and performs adaptive constraint alignment processing on both, enabling the model to understand semantic and spatial relationships within a unified inference structure, thereby achieving a more consistent, transparent, and traceable decision-making process and improving the accuracy and interpretability of inference results in multimodal scenarios. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0013] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the decision-making method based on a visual thinking chain provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0014] like Figure 2 As shown, the decision-making method based on visual thinking chain proposed in this invention includes the following steps: S10, receive multimodal data including text data and image data; In this embodiment, receiving multimodal data, including text and image data, involves input management of two types of information with different sources and expressive attributes. Implementing this processing typically requires five aspects: data source determination, data flow entry configuration, data structure parsing, synchronization maintenance, and data quality verification. Text data can originate from character sequences, tokenized language content, structured text units, or natural language descriptions. Its sources may include user input, business recording systems, acquisition devices, or third-party access services. Text data requires character encoding parsing during the input process, converting external formats into internally processable semantic expression units, such as converting Unicode encoding, JSON fragments, or HTML fragments into plain text sequences. Image data refers to pixel matrices generated by optical sensors, mobile device acquisition units, business upload terminals, or external vision systems. This type of data may arrive at the processing end in formats such as JPEG, PNG, WEBP, and binary pixel streams. During the input stage, format parsing, pixel matrix recovery, color space normalization, or resolution normalization must be performed.
[0015] Receiving multimodal data requires building a unified data input interface. This interface can be implemented through network request endpoints, message queue access points, device data channels, or local cache read modules. Its function is to establish a controllable input entry point for text and image data streams. Cross-source input introduces different timestamps, different flow rates, and different encoding methods. Therefore, the input interface needs to include data boundary judgment and flow integrity detection mechanisms to identify abnormal truncation, missing segments, or duplicate packets.
[0016] To ensure the ability of subsequent feature processing to identify associations in multimodal data, it is necessary to introduce the ability to build associations between text content and image content during the input stage. These associations can be established based on business fields, timestamp proximity, user behavior events, scene numbers, or cross-modal reference identifiers. For example, in financial transactions, an image might correspond to a textual description in a claims image record; in healthcare transactions, an image might correspond to body part information or supplementary descriptive text. The logic for establishing these associations must ensure that the text sequence and pixel matrix have consistent referencing identities before entering subsequent processing modules, thereby maintaining the foundation for cross-modal semantic alignment.
[0017] Before entering the fusion processing, the text and image data streams must maintain synchronization. The input management module needs to attach time-series markers or unique identifiers to the received results, enabling data from different modalities to be used as a unified event processing unit in subsequent processes. Synchronization maintenance can be achieved based on reception time, service-triggered events, or sequence numbers from external transmission protocols. For data quality checks, content integrity checks are required at the input end, such as determining whether the text character sequence is abnormally interrupted or whether the image matrix has severe defects, to ensure sufficient data stability before entering the feature extraction environment.
[0018] This embodiment achieves joint reception, format parsing, synchronous management, and association construction of text and image data at the input stage, enabling subsequent semantic processing modules to operate within a consistent data reference space, thus providing a stable input foundation for multimodal fusion processing. This structure improves the reliability of cross-modal associations, avoiding source mismatches or semantic confusion between text and image information during the inference stage, and providing a solid foundation for subsequent inference chain generation and interpretive output.
[0019] S20, perform feature extraction and fusion processing on the text data and image data to obtain multimodal fusion features; In this embodiment, feature extraction and fusion processing of text and image data involves encoding the features of the data content that has already been parsed at the input end for both modalities, and constructing an interactive expression space based on cross-modal relationships, enabling linguistic and visual content to establish corresponding items in a unified semantic domain. The semantic features of the text data originate from the semantic extraction process of character sequences, token sequences, or natural language units. This extraction process typically generates semantic vector representations through embedding mechanisms. The text encoder constructs semantic expressions through word embedding, subword embedding, or character embedding, transforming the text content into a semantic tensor for subsequent modules to perform context inference. The semantic feature interpretation requires explaining that the text encoder captures semantic dependencies within the text through multi-layer attention structures or sequence modeling structures, ensuring the coherence of event descriptions, spatial descriptions, or logical descriptions at the vector level.
[0020] The visual features of image data originate from the process of extracting spatial patterns and texture structures from the image pixel matrix. Image encoders extract the dependencies between shapes, edges, textures, and local regions through convolutional structures, hierarchical feature pyramids, or visual attention structures, thereby forming a stable visual vector representation in visual space. The existence of visual features is used to describe the relationships between objects, the layout of regions, or spatial orientation in an image. These visual elements may form cross-modal constraints with linguistic descriptions in subsequent reasoning.
[0021] Cross-modal fusion aims to introduce semantic and visual features into a shared representation space, enabling the two types of features to establish consistent associations during interaction. The fusion process can be achieved through cross-modal attention modules, mutual attention mechanisms, or bidirectional weighted mechanisms, allowing semantic vectors to focus on corresponding regions in the image, while visual features focus on key semantic words in the text description. For example, when the text contains spatial descriptors, action descriptors, or entity names, the cross-modal fusion module aligns the text vector portion to the corresponding region in the image matrix using attention weights, constructing a bidirectional interactive association tensor.
[0022] To improve the reliability of representation during the fusion process, feature enhancement mechanisms can be introduced. Feature enhancement means denoising, fine-grained gain, spatial reweighting, or semantic unification of the tensor after interaction fusion, so that the interaction representation can fully reflect multimodal semantics. For example, feature enhancement for images can emphasize edge textures or region shapes through local weighting, while feature enhancement for text can apply vector gain to text segments containing key semantics or logical threads. The enhanced tensors constitute multimodal fusion features, enabling textual semantics and visual information to have a consistent structure in a unified space, allowing subsequent inference modules to extract cross-modal association patterns from a single tensor.
[0023] This embodiment constructs semantic and visual features from text and image data respectively, and achieves interactive fusion through a cross-modal mechanism. This enables linguistic and visual expressions to establish a consistent semantic relationship in a unified space, allowing subsequent inference modules to simultaneously utilize linguistic and visual cues for inference. This structure enhances the consistency of responses between cross-modal information, reduces biases caused by independent modal processing, and establishes a consistent cognitive foundation for subsequent inference chain generation and interpretive output.
[0024] S30, perform semantic and spatial joint encoding processing on the multimodal fusion features to obtain joint feature representation; In this embodiment, the process of semantic and spatial joint encoding of multimodal fusion features includes hierarchical semantic extraction and spatial structure reconstruction of the fused cross-modal expression, enabling linguistic semantics, visual layout, and cross-modal dependencies to form a hierarchical expression in a unified vector space. Multimodal fusion features typically contain interactive representations of the original text semantics and the original image visuals, and already contain preliminary cross-modal attention relationships. The goal of the joint encoding process is to further extract word-level, sentence-level, and paragraph-level semantic associations from this fusion tensor, while reconstructing the spatial distribution structure, so that subsequent inference chain generation possesses interpretable semantic depth and spatial consistency.
[0025] Semantic joint encoding includes word-level semantic aggregation, sentence-level semantic aggregation, and paragraph-level semantic aggregation. Word-level semantic aggregation involves extracting local representations corresponding to words, word fragments, or semantic phrases from the fusion tensor. Through local context modeling, fine-grained semantic segments highly correlated with visual regions in the text are extracted as vectors. Word-level aggregation relies on multiple modal interaction weights in the fusion features, enabling language segments to be associated with corresponding visual regions, forming a semantically aligned set of word vectors.
[0026] Sentence-level semantic aggregation integrates multiple word-level vectors into a higher-level representation according to the internal structure of the language. It captures event logic, causal relationships, and semantic span through attention networks or sequence modeling structures, combining the overall meaning of the language description with the spatial relationships in the image into a consistent sentence vector expression. Sentence-level aggregation not only represents the semantic relationships across words in the language but also maintains cross-modal consistency between language and images.
[0027] The role of paragraph-level semantic aggregation is to identify paragraph-level thematic clues, event background, and overall narrative context from multiple sentence-level representations. Through a hierarchical modeling mechanism, it forms paragraph vectors with a relatively long semantic span, enabling the entire input content to form a stable semantic thread. Thematic clues in paragraph-level representations automatically absorb scene background information from visual features, making the text narrative and image structure form a consistent overall interpretive framework.
[0028] Spatial joint coding introduces spatial structure reconstruction capabilities on top of paragraph-level representation, enabling cross-modal fusion features to reflect the correspondence between the spatial order of entities in an image and the spatial directional words, locative words, or action paths appearing in the semantics. Spatial joint coding achieves a consistent mapping relationship between paragraph-level semantic vectors and the layout structure extracted from visual features through spatial weight adjustment, spatial feature mapping, or spatial embedding injection. For example, when locative words such as "ahead," "right," and "near" appear in an accident description, spatial coding maps these semantic cues to specific regions in the image, giving the spatial representation geometric consistency.
[0029] Hierarchical attention compression is used to compress and gradient-weight hierarchical representations after word-level, sentence-level, and paragraph-level extraction, resulting in a compact yet comprehensive joint feature representation. The hierarchical attention structure weights different semantic levels based on the importance of the input content, giving higher attention to key semantic segments, key visual regions, or key spatial relationships, thereby improving the overall expressive power of the joint feature representation. The final vector representation obtained through hierarchical compression maintains both semantic continuity and spatial consistency, providing stable input for subsequent inference chain generation.
[0030] This embodiment uses semantic and spatial joint encoding on multimodal fusion features to establish consistent semantic and spatial links between linguistic and visual content at multiple levels, thereby forming a semantically clear and spatially consistent joint feature representation. This joint representation enhances the accuracy of subsequent inference chain generation, allowing linguistic and visual inference to be logically developed based on the same representational foundation, while reducing errors caused by information fragmentation and ensuring the interpretability and consistency of the inference process.
[0031] S40, Based on the joint feature representation, generate a visual thought chain including a language reasoning chain and a visual reasoning trajectory; In this embodiment, the process of generating linguistic reasoning chains and visual reasoning trajectories based on joint feature representations utilizes multi-level information that has already undergone semantic and spatial integration in the joint representation. Through a generation mechanism, the implicit reasoning logic, event sequence, spatial cues, and cross-modal associations in the representation are gradually made explicit, so that the reasoning process is presented in two forms: linguistic chains and visual trajectories. Joint feature representations typically include textual semantics, visual region information, cross-modal consistency cues, and spatial layout embeddings, thus possessing a sufficient informational foundation to facilitate the generation of reasoning chains.
[0032] The generation of language reasoning chains relies on a language reasoning chain generator. This generator is a type of sequence generation network that selects semantic units from joint representations and reconstructs event logic through sequential generation, such as causal logic, temporal order, state changes, and conditional triggering. The meaning of the language reasoning unit sequence is a linear sequence of sentence fragments, logical phrases, or semantic segments output by the generator. These segments correspond to interpretable reasoning steps, such as "contact occurs," "directional impact," and "positional change." The generator determines the semantic region and spatial information that each language unit should reference through the attention distribution of the joint representations, thus ensuring that the generated language logic chain has factual basis.
[0033] Visual reasoning trajectory generation relies on a visual reasoning trajectory generator. This generator extracts visual region vectors, spatial coordinates, and cross-modal attention maps from the joint representation, organizing this information into a visual sequence. The visual reasoning unit sequence typically includes key location markers, region numbers, and spatial orientation or state change cues in the image, which can be represented by vector trajectories, region sequences, or spatial coordinate sequences. The essence of visual trajectory generation is to reconstruct the spatial logic within visual evidence according to time or reasoning order, providing not only verbal descriptions but also a visual chain of spatial evidence for the reasoning process.
[0034] The role of the cross-modal synchronization controller is to ensure consistency between the language reasoning unit sequence and the visual reasoning unit sequence in terms of reasoning order, event granularity, and spatial logic. The synchronization control mechanism utilizes the cross-modal attention matrix in the joint representation to synchronize the two types of sequences, ensuring that each logical node in the language reasoning chain corresponds to a spatial node in the visual reasoning trajectory. The synchronization process generates temporally coordinated language and visual sequences, essentially aligning two originally independently generated sequences to the same time or logical axis, enabling bimodal reasoning to proceed in parallel.
[0035] Once the sequence coordination is complete, the language reasoning unit sequence forms a language reasoning chain through a sequence combination mechanism. The combination process structurally integrates the language units in the sequence according to the order of events, giving the language chain a clear form of reasoning development. Similarly, the visual reasoning unit sequence forms a visual reasoning trajectory through a combination mechanism, that is, connecting spatial units in a logical order to form a parseable structure with directional, line segment, or regional sequences.
[0036] Ultimately, a combination mechanism is used to unite the linguistic reasoning chain and the visual reasoning trajectory to form a visual thinking chain. The meaning of the visual thinking chain is to represent the semantic logic chain and the visual trajectory structure in a unified structure, so that logical content and spatial content can be presented simultaneously, providing a unified reasoning basis for subsequent alignment and interpretation output.
[0037] This embodiment generates linguistic reasoning chains and visual reasoning trajectories based on joint feature representations. This transforms the reasoning process from simply presenting the final answer to publicly generating logic through linguistic chains and visual trajectories, explicitly recording spatial evidence, semantic reasoning, and cross-modal consistency. This enhances the transparency of the reasoning process, enabling decisions to be tracked, verified, and explained, while simultaneously improving the reliability, completeness, and user understandability of the reasoning.
[0038] S50, perform adaptive constraint alignment processing on the language reasoning chain and visual reasoning trajectory of the visualized thinking chain to obtain the aligned visualized thinking chain; In this embodiment, after generating the linguistic reasoning chain and the visual reasoning trajectory, they need to be fused under unified constraints to form a structurally coordinated and logically consistent aligned visual thought chain. The linguistic reasoning chain typically focuses on semantic logic paths, including inference relationships, semantic nodes, and textual explanation clues; the visual reasoning trajectory focuses on spatial attention regions, including detection regions, feature highlights, and visual jump sequences. Because semantic paths and spatial paths differ in their representation and organizational structure, directly presenting them side-by-side would cause a fragmentation of reasoning and interpretation. To address this difference, adaptive constraint alignment processing needs to be introduced for the linguistic reasoning chain and the visual reasoning trajectory, ensuring that the two types of reasoning structures form a consistent alignment mapping in terms of semantics, space, and attention rhythm.
[0039] The process first extracts semantic and spatial path nodes as alignment objects based on the path nodes contained in the linguistic inference chain and the visual inference trajectory, respectively. The system constructs a constraint set based on factors such as text semantic level, inference relevance, visual saliency, and spatial positional relationship. This constraint set describes semantic and spatial constraints, including semantic consistency requirements, spatial range matching relationships, and alignment priorities based on node relevance.
[0040] Given a constrained set, the semantic path nodes in the linguistic reasoning chain and the spatial path nodes in the visual reasoning trajectory are compared layer by layer. By analyzing the differences in their temporal order, semantic span, and visual attention area, potential biases are identified, such as semantic interpretation lagging behind visual judgment, visual scanning range preceding textual reasoning triggering, or inconsistent path node density distribution. For the identified biases, alignment corrections are made by adjusting semantic weights, refining the visual attention area, and coordinating node order, so that the linguistic reasoning chain and the visual reasoning trajectory achieve unified expression in semantic logic, attention area, and reasoning rhythm.
[0041] This embodiment performs adaptive constraint alignment on the linguistic reasoning chain and the visual reasoning trajectory, establishing a mapping relationship between semantic path nodes and spatial path nodes under unified constraints, thereby reducing structural offset between the two types of reasoning results. Because the linguistic reasoning chain and the visual reasoning trajectory are coordinated in terms of semantic logic, spatial distribution, and reasoning rhythm, the aligned visual thinking chain can simultaneously display the semantic reasoning process and spatial attention changes within the same representation structure. This allows subsequent decision generation stages to form a more stable explanatory chain based on consistent reasoning, achieving clearer reasoning presentation and more traceable decision output.
[0042] S60, Based on the aligned visual thought chain, generate an interpretable decision output.
[0043] In this embodiment, when generating interpretable decision output based on the aligned visual thought chain, the linguistic reasoning chain nodes and visual reasoning trajectory nodes contained in the aligned visual thought chain are first utilized. These nodes have already established semantic-spatial correspondences in the preprocessing, thus serving as a stable basis for interpretive decisions. Each linguistic reasoning chain node records key semantic units in the textual reasoning logic, such as judgment conditions, inference fragments, or intermediate concepts; each visual reasoning trajectory node records key areas of interest in the image region, such as salient areas, spatial relationships, or local change trends. During the generation of interpretable decision output, decision components need to be extracted from these nodes, including semantic judgments, visual evidence, causal paths, and inference order.
[0044] To enable reasoning chains to directly drive interpretable decision outputs, an explanation generation mechanism is needed to combine linguistic and visual reasoning nodes. This mechanism scans the aligned visual thought chain, identifies the logical order and trends in the areas of interest between reasoning nodes, and transforms these relationships into a structured set of decision-making evidence. Specifically, the explanation generation mechanism maps logical paths in the linguistic reasoning chain to explanatory sentences and spatial evidence in the visual reasoning trajectory to graphical cues or location references through text template filling, visual evidence binding, or node relationship combination.
[0045] In the actual generation process, interpretable decision output needs to include two types of content: a conclusion and an explanation. The conclusion originates from the final inference expressed by the terminal nodes in the aligned visual thought chain; the explanation originates from the combination of intermediate nodes in the reasoning chain, including semantic transition relationships, changing trends in visual attention areas, and causal mappings between modalities. The generation process can rely on text generation models, rule engines for combination, or graph structure parsing mechanisms to form visual explanations. The explanation content corresponds one-to-one with the aligned visual thought chain, making the output results traceable.
[0046] In practical applications, the generation method of explainable decision output can be adjusted according to different expression needs and operating environments. For example, a language model-based explanation generation method can be used to directly generate text output with reasoning explanations based on the aligned visual thought chain, making the explanation content more natural language expressive; alternatively, a structured rule combination method can be used to abstract the nodes of the reasoning chain into rule expressions, and form a highly readable explanation string through rule weight combination; furthermore, when adapting to graphic display scenarios, a hybrid output containing text explanations and image overlays can be constructed to simultaneously present the reasoning chain and image evidence in the visual interface.
[0047] For environments requiring high-precision output, a weighting mechanism can be used to prioritize inference chain nodes with higher stability as core explanatory basis and assign them higher explanatory power. For environments requiring real-time processing, the density of image evidence parsing can be reduced, and only the main node can be used for rapid interpretation. For environments requiring business review, more structured generation methods can be introduced to generate standardized explanatory documents that include conclusion descriptions, evidence location, and semantic inference links.
[0048] In scenarios with significant differences in multimodal inputs, the node matching strategy can be adjusted to make the explanation generation mechanism focus more on language reasoning chain fragments with high confidence, or to emphasize key areas in the visual trajectory. When adapting across domains, the output content can conform to the semantic specifications of the target industry by replacing the explanation template or explanation structure.
[0049] This embodiment generates interpretable decision output based on aligned visual thought chains, allowing both linguistic reasoning chains and visual reasoning trajectories to participate in decision construction as output bases simultaneously. This mitigates the inconsistency in interpretation caused by the independent operation of semantic and visual paths in multimodal reasoning. Since the final output is directly generated based on the aligned visual thought chains, the interpretation maintains consistency in logical order, evidence citation, and causal structure, thereby improving the traceability and verifiability of the decision results. This process can form a more stable interpretive chain among multimodal information, making the relationship between conclusions and evidence clearer.
[0050] In one embodiment, step S10 above includes: S101 receives text data streams and image data streams through the data input interface; S102, perform format parsing on the text data stream to obtain text data; S103, The image data stream is parsed to obtain image data; S104, Establish the association between the text data and the image data; S105, the text data and image data that have a correlation relationship are treated as multimodal data.
[0051] In this embodiment, after text and image data streams enter the system, they are uniformly accessed through a data input interface. The data input interface can manifest as a server access channel, a message channel, or a batch access channel. The interface layer only handles streaming data transfer and arrival confirmation, without being limited to a physical form. To adapt to different sources, each text and image data stream is appended with a source marker, an arrival time marker, and a session marker during the access phase. The source marker distinguishes between the acquisition end and the service end, the time marker is used to establish the subsequent alignment order, and the session marker is used to maintain the same transaction scope across request links. These three types of markers do not change the data itself; they are only referenced as metadata in subsequent processing.
[0052] The format parsing of the text data stream is performed by the text parsing unit. The parsing unit first identifies character encoding and line break rules to establish a unified character sequence, and then segments the text data into processable segments based on sentence boundaries, paragraph boundaries, and entity clues. Sentence boundaries can be determined jointly based on punctuation and spacing patterns, paragraph boundaries are located through a combination of blank lines and structural cue words, and entity clues are extracted through dictionary mapping or contextual patterns. During parsing, time expressions, location expressions, and object titles are extracted simultaneously and referenced with source markers, time markers, and session markers to form text data with structural cues. This text data is subsequently accessed only in the form of unified encoding, unified boundaries, and complete meta-information, without restriction on language category or script type.
[0053] The format parsing of the image data stream is performed by the image parsing unit. The parsing unit first identifies the container encapsulation and pixel arrangement to obtain the pixel matrix and color space. Then, it reads the time information, spatial orientation, imaging parameters, and device information from the image metadata to form structured image data. For serialized image data streams, the parsing unit retains the frame number and acquisition interval for subsequent alignment with text time stamps. To facilitate subsequent spatial processing, the parsing unit establishes references between the image data and the pixel matrix size, color space type, and orientation without altering the image itself, ensuring consistent access to subsequent spatial representations.
[0054] The association between text and image data is established by the association construction unit. This unit uses session markers as the primary constraint, time markers as the alignment clues, and source markers and contextual cues to complete multi-dimensional association. The first step aggregates text and image data within the same transaction scope into a candidate set based on session markers. The second step uses time markers to find paired elements within adjacent windows in the candidate set. The third step, when multiple pairs of candidates exist, introduces contextual cues for semantic proximity matching. For example, cues such as "location direction," "part name," and "object title" provide mutual guidance to the metadata of the image content. In cases of one-to-many or many-to-one relationships, the association construction unit prioritizes time proximity and contextual consistency, reserving unresolved entries as a pending association set for later higher-level processing or resubmission. After the association is generated, the text and image data are bound as traceable pairs, with each holding a reference to the other and its corresponding marker mapping.
[0055] The multimodal encapsulation unit handles the unified output of related text and image data as multimodal data. This unit maintains a bidirectional index between text, image, and various tags, without compressing or cropping the content. The output is a structured object directly accessible to subsequent modules, containing one-to-one, one-to-many, or many-to-one mappings and reference tables. This structure also preserves source tags, time tags, session tags, and contextual cue mappings, ensuring direct reference to time sequence, source links, and event scope during subsequent feature extraction and fusion stages, thus avoiding redundant parsing and alignment. The encapsulated result is named "multimodal data" and maintains stable field naming and access paths internally for easy sharing across modules.
[0056] This embodiment performs format parsing and multidimensional association on text and image data streams, incorporating source markers, time markers, and session markers into unified management during the access phase. Multimodal data possesses stable structural representation and paired reference relationships before entering subsequent stages, reducing cross-modal alignment overhead and secondary processing burden, and decreasing the probability of mismatches and omissions. This provides a consistent entry point and traceable reference chain for subsequent feature extraction and fusion, joint encoding, and inference chain construction.
[0057] In one embodiment, step S20 above includes: S201, Extract semantic features of the text data using a text encoder; S202, extract the visual features of the image data using an image encoder; S203, The semantic features and the visual features are interactively fused through a cross-modal attention mechanism to obtain interactive fused features; S204, Perform feature enhancement processing on the interactive fusion features to obtain enhanced fusion features, and use the enhanced fusion features as multimodal fusion features.
[0058] In this embodiment, after receiving text data, the text encoder first performs unified encoding and boundary recognition to obtain a continuous sequence of word fragments and their corresponding positional information. The encoder projects the word fragments and positional information into the same semantic space using embedding mapping. Through a context modeling unit, it captures short-range and long-range dependencies in multi-layer semantic interactions and outputs sequentially arranged semantic features. These semantic features retain word-fragment level representations and sentence-segment level aggregation information, which are used as query or key-value sources in subsequent stages to participate in cross-modal interactions. To improve robustness, the text encoder maintains masking and regularization mechanisms to suppress the offset caused by noisy fragments, and stabilizes gradient propagation through normalization and residual connections, ensuring that the semantic features maintain the traceability of semantic boundaries and referential relationships even after multiple rounds of stacking.
[0059] After receiving image data, the image encoder performs scale alignment and color normalization. Based on the image size, it divides pixel regions into local receptive fields, extracting low-to-mid-level information such as edges, textures, and shapes. Then, it constructs a global receptive field through inter-layer aggregation. The visual features output by the encoder retain multi-scale representation in both channel and spatial dimensions, allowing key elements of the same object at different scales to be accessed simultaneously. To ensure spatial topological continuity, the image encoder maintains positional cues during feature generation, ensuring that visual features can establish a stable mapping with spatial vocabulary in subsequent cross-modal interactions.
[0060] The cross-modal attention mechanism establishes a bidirectional interaction between semantic and visual features. On one hand, the mechanism uses semantic features as queries and visual features as keys to locate the spatial region that best matches the textual expression, generating a language-guided visual response. On the other hand, it uses visual features as queries and semantic features as keys to backtrack to the textual cues closest to the candidate image region, generating a visually guided linguistic response. This dual-path design ensures that the interactive fusion features possess both semantic direction and spatial evidence, gradually suppressing irrelevant regions and redundant word phrases in multiple rounds of interaction. To avoid single-path dominance, the mechanism introduces weight balancing and gating aggregation to coordinate and merge the responses generated by language and visual guidance, outputting interactive fusion features that are aligned in both temporal order and spatial location.
[0061] Feature enhancement processing addresses the consistency and discriminative power of interactive fusion features. First, hierarchical normalization and residual stacking solidify the distribution stability after cross-modal alignment, reducing scale drift caused by different sources. Then, multi-scale aggregation facilitates information backflow between local and global windows, merging fine-grained localization information and long-range dependencies into a unified representation. Next, channel recalibration and sparse activation improve the responsiveness to key semantic and salient spatial channels, suppressing the influence of background and noise channels. Finally, consistency constraints maintain the correspondence between the linguistic and visual sides during multiple rounds of enhancement. The resulting enhanced fusion features, after structured encapsulation, are directly used as multimodal fusion features for subsequent semantic and spatial joint encoding, maintaining stable names and access paths and avoiding cumulative errors caused by repeated alignment and aggregation.
[0062] This embodiment constructs semantic and visual features using a text encoder and an image encoder, respectively. A cross-modal attention mechanism is then employed to achieve bidirectional interaction and output interactive fusion features. Feature enhancement processing is then used to obtain enhanced fusion features, which are input as multimodal fusion features into subsequent stages. This ensures that linguistic cues and spatial evidence undergo saliency screening and consistency correction before entering joint encoding. This reduces cross-modal mismatches and information redundancy, lowers the alignment burden of subsequent encoding and inference, and improves the stability and usability of representations in complex semantic and multi-scale visual contexts.
[0063] In one embodiment, step S30 above includes: S301, perform word-level semantic aggregation processing on the multimodal fusion features to extract local context information and obtain word-level aggregated features; S302, Perform sentence-level semantic aggregation processing on the word-level aggregation features to obtain sentence-level aggregation features; S303, Perform paragraph-level semantic aggregation processing on the sentence-level aggregation features to obtain paragraph-level aggregation features; S304, perform hierarchical attention compression processing on the paragraph-level aggregated features to obtain a joint feature representation.
[0064] In this embodiment, after the multimodal fusion features enter the joint encoding stage, they first undergo structured organization and index binding, maintaining bidirectional reference relationships with preceding semantic and spatial elements. This ensures that subsequent aggregations at all levels can directly access the original context and spatial location information. For word-level semantic aggregation processing, continuous semantic segments are first established based on word boundaries within the sentence. Semantic cues such as referentiality, limitation, negation, and temporal location are collected within adjacent windows through local context modeling. Then, positional cues and contextual cues jointly constrain the aggregation direction, forming a weighted aggregation based on neighborhood dependencies, outputting word-level aggregation features. This result retains the index table and positional mapping from word segments to aggregation units, facilitating subsequent higher-level aggregations to trace back to word-level criteria.
[0065] For sentence-level semantic aggregation processing, the method first determines intra-sentence boundaries and cross-sentence connection points based on sentence boundary markers and pause cues. Within a sentence, it aggregates event elements, entity relationships, and causal transitions. Within a cross-sentence scope, it identifies referential objects and temporal continuations, forming a joint expression that is consistent within sentences and coherent across sentences, and outputs sentence-level aggregation features. This result simultaneously maintains the citation relationship from the semantic core to the evidence fragments and retains potential ambiguities as candidates for paragraph-level processing and adjudication.
[0066] For paragraph-level semantic aggregation processing, guided by intra-paragraph thematic cues and inter-paragraph connection cues, sentence-level aggregation features are merged and compared. This process handles phenomena such as scene transitions, time progression, and perspective shifts, unifying the main and secondary storylines over a longer period and outputting paragraph-level aggregation features. To accommodate spatial representation, synchronization constraints are established with spatial positioning in multimodal fusion features. Locative words, location words, and region words appearing within a paragraph are bound to their corresponding spatial indices, enabling long-range semantics to carry stable spatial cues in paragraph-level expression.
[0067] For hierarchical attention compression, a multi-level weight allocation mechanism is constructed based on paragraph-level aggregated features. Weighted convergence is applied to word-level, sentence-level, and paragraph-level sources respectively. First, weights for the main theme and key evidence are extracted within each paragraph. Then, global sparse and dense aggregation are collaboratively compressed across paragraphs to obtain a joint feature representation. During compression, the index traceability structure is maintained. Each aggregation unit in the joint feature representation has a backtracking path, allowing retrieval from the global representation to paragraph-level, sentence-level, and word-level aggregated features, and further to the semantic and spatial elements of the multimodal fusion features. To ensure stable distribution, normalization and residual connections are introduced at each level of the output to suppress gradient and scale drift. Consistency constraints maintain the one-to-one correspondence between the semantic main theme and spatial index, ensuring the joint feature representation possesses both compressed global expressive power and refined retrieval capabilities down to word fragments and spatial locations.
[0068] This embodiment sequentially completes word-level, sentence-level, and paragraph-level aggregation features, and generates a joint feature representation using hierarchical attention compression. This enables local semantics, intra- and inter-sentence relationships, and paragraph-level topics to converge from the bottom up in a unified representation. At the same time, it retains the traceable index of multimodal fusion features, thereby reducing redundancy and ambiguity while maintaining a consistent semantic-space mapping. This facilitates flexible switching between globally compact representation and local evidence in the subsequent reasoning chain generation stage, improving long-range understanding stability and interpretable retrieval capabilities.
[0069] In one embodiment, step S40 above includes: S401, Input the joint feature representation into the language reasoning chain generator to generate a language reasoning unit sequence; S402, input the joint feature representation into the visual reasoning trajectory generator to generate a visual reasoning unit sequence; S403, using a cross-modal synchronization controller to process the language reasoning unit sequence and the visual reasoning unit sequence, generating a time-coordinated language reasoning unit sequence and a time-coordinated visual reasoning unit sequence; S404, combine the temporally coordinated language reasoning unit sequence into a language reasoning chain; S405, combine the time-coordinated visual reasoning unit sequence into a visual reasoning trajectory; S406, combine the language reasoning chain and the visual reasoning trajectory into a visual thinking chain.
[0070] In this embodiment, after the joint feature representation enters the language reasoning chain generator, it first completes the semantic clue deconstruction and causal clue sorting, mapping the event predicates, object names, temporal expressions, and spatial cues in the joint feature representation into a traceable sequence of language reasoning units. Without changing the internal referencing relationships of the joint feature representation, the generator establishes directed connections between language reasoning units based on contextual coherence and referential relationships, and retains a bidirectional index of the source fragment and spatial cues for each language reasoning unit to facilitate stable pairing with visual units later. To avoid fragment redundancy, the generator removes semantic overlap between adjacent language reasoning units, retaining action descriptions and condition descriptions that are distinctive in event progression, and ensures that the same object maintains consistent naming and consistent positional references in the sequence through constraints.
[0071] The joint feature representation is simultaneously input into the visual reasoning trajectory generator. Guided by spatial cues and saliency hints in the joint feature representation, the generator locates and determines the connectivity of visual candidate regions, outputting a sequence of visual reasoning units. Each visual reasoning unit carries three types of information: region location, appearance representation, and temporal position, and establishes a one-to-one or one-to-many reference with the spatial cues in the joint feature representation. The generator performs continuity determination on the regional connectivity across frames, removes unit fragments caused by short-term jitter, and merges the expressions of the same object at adjacent positions into stable segments to form a one-to-one temporal span with the action continuation on the language side.
[0072] The cross-modal synchronization controller receives sequences of linguistic reasoning units and visual reasoning units, and outputs temporally coordinated sequences of linguistic and visual reasoning units. The controller first uses time cues and event boundaries in the joint feature representation as a common time axis to align the start and end points of the two sequences. Then, based on the consistency of spatial cues, it fine-tunes any potential offsets between the two sequences, inserting empty spaces to bridge gaps or folding repetitions to remove stacking if necessary, without altering the original content. Finally, it prioritizes conflicting segments, preserving pairs that simultaneously satisfy source consistency, spatial consistency, and causal consistency. Segments that do not meet these conditions are marked as pending adjudication and accompanied by source evidence. After synchronization, the controller outputs a temporally coordinated sequence of linguistic and visual reasoning units, which are directly correlated in terms of time axis and spatial index.
[0073] The linguistic reasoning chain is obtained by combining a sequence of temporally coordinated linguistic reasoning units. The combination process follows an event-driven progression, connecting preconditions, actions, and results in chronological order, and establishing traceable branch labels at each branch point to ensure that the source of evidence can be located at any linguistic reasoning unit during subsequent interpretation generation stages. The visual reasoning trajectory is obtained by combining a sequence of temporally coordinated visual reasoning units. The combination process follows object continuity, organizing the starting position, movement path, and state changes as connected segments, and retaining dual references of spatial anchors and appearance anchors at path nodes to facilitate point-to-point association with action or state words in the linguistic reasoning chain.
[0074] The visualized thought chain is derived from a combination of a linguistic reasoning chain and a visual reasoning trajectory. The combination process employs a dual-path anchoring strategy: one type of anchoring uses actions or conditions in the linguistic reasoning chain as anchor points to find corresponding nodes in the visual reasoning trajectory; the other type uses state changes in the visual reasoning trajectory as anchor points to find corresponding expressions in the linguistic reasoning chain. Strong connections are established when both temporal and spatial consistency are satisfied; weak connections are established and marked for verification when only one is satisfied. The final visualized thought chain is stored in a graph structure, with edge types distinguishing between strong and weak connections. Nodes maintain backtracking pointers, allowing the entire chain to be traced back from any point to the joint feature representation, and then back to the multimodal fusion features and their source fragments, achieving a step-by-step localization from the global link to the original evidence.
[0075] This embodiment constructs language reasoning unit sequences and visual reasoning unit sequences in parallel on the same joint feature representation through a language reasoning chain generator and a visual reasoning trajectory generator. The cross-modal synchronization controller completes the temporal coordination under a unified time axis and spatial index. The two coordinated sequences are then combined into a language reasoning chain and a visual reasoning trajectory, and merged into a visual thinking chain. This makes the semantic advancement and spatial changes form a one-to-one evidence chain in structure, reducing ambiguity caused by cross-modal mismatch and temporal offset, and improving the traceability and consistency in subsequent interpretation generation and decision output.
[0076] In one embodiment, step S50 above includes: S501, Extract the semantic embedding representation of the language unit in the language reasoning chain; S502, Extract the visual unit visual embedding representation from the visual reasoning trajectory; S503, determine the semantic distance between the semantic embedding representation of the language unit and the visual embedding representation of the visual unit; S504, Based on the semantic distance, adjust the alignment relationship between the language inference chain and the visual inference trajectory to generate a semantically aligned language inference chain and a semantically aligned visual inference trajectory; S505, the semantically aligned linguistic reasoning chain and the semantically aligned visual reasoning trajectory are combined into an aligned visual thought chain.
[0077] In this embodiment, after the language inference chain and visual inference trajectory enter the adaptive constraint alignment stage, the semantic embedding representations of language units and the visual embedding representations of visual units are first extracted within a unified reference space. To ensure that the two types of embeddings can be directly compared, a normalized mapping with consistent sources is first established, binding the semantic range, temporal position, and spatial cues of language units to a joint index. Then, the appearance description, spatial anchor point, and temporal position of visual units are bound to the same joint index, forming a set of embeddings that can be traced back to each other. The acquisition of the semantic embedding representation of language units is based on the granularity of language inference chain nodes, retaining key components including action words, conditional words, entity references, and causal connections, and including references to source fragments and spatial cues. The acquisition of the visual embedding representation of visual units is based on the granularity of visual inference trajectory nodes, retaining key components including region appearance, pose changes, and trajectory continuity, and including references to spatial anchor points and temporal positions. Both types of embeddings maintain a unified scale description and index format at the output end, facilitating subsequent measurement by pair.
[0078] After the embedding set is ready, a visual candidate set that matches the temporal and spatial neighborhoods is constructed for the semantic embedding representation of each language unit to avoid noise introduced by irrelevant pairings. Constrained by temporal overlap, spatial proximity, and consistent source, visual units corresponding to the language units are screened out to form a one-to-many candidate list. Then, for each pair in the list, the semantic distance is determined. The determination of the semantic distance follows the principles of homologous scale and homodirectionality. First, the scales of the embedding vectors are aligned, and then the measurement results are given based on directional consistency and magnitude differences. Pairings with inconsistent sources or weak evidence are marked with low confidence. To prevent extreme values from dominating the overall alignment, robustness processing is also introduced in the determination of the semantic distance, providing suppression factors for abnormally deviating pairings and retaining the original measurement records for subsequent backtracking and verification.
[0079] Alignment adjustments are driven by semantic distance and constrained by both temporal and spatial constraints. The linguistic inference chain and visual inference trajectory are checked segment by segment. When the semantic distance of a pairing is consistently less than the neighborhood average and satisfies temporal overlap and spatial proximity, a stable correspondence is established and the segment is marked as a strong correspondence. When a segment only partially satisfies the constraints, a pending correspondence is established, and placeholders are inserted at both ends to bridge rhythmic differences. When a conflict occurs in a segment, structural adjustments such as rearrangement, merging, or pruning are triggered. Rearrangement corrects the sequential misalignment of linguistic and visual segments, merging eliminates fragmented segments caused by short-term jitter on the visual side, and pruning removes isolated nodes with insufficient continuity or inconsistent origins. Each structural adjustment synchronously updates the backtracking index, ensuring that the new correspondence can still be traced back level to the original evidence of the joint feature representation and multimodal fusion features.
[0080] After structural adjustments, the semantically aligned linguistic inference chain and the semantically aligned visual inference trajectory are output. The semantically aligned linguistic inference chain records the mapping item and time span corresponding to visual stability at each node. The semantically aligned visual inference trajectory records the action or state label corresponding to linguistic stability at each node, and provides identifiers for weak correspondences and those awaiting adjudication for interpretation and filtering in the next stage. Finally, the two semantically aligned sequences are combined into an aligned visual thought chain, storing the correspondences in a graph structure. Strong correspondences are represented by backbone edges, and weak correspondences and those awaiting adjudication are represented by auxiliary edges. Each node and edge in the graph retains a backtracking pointer and metric record, ensuring that any conclusion can be verified against the semantic embedding representation of the linguistic unit and the visual embedding representation of the visual unit. This combination process does not change the existing content; it only completes the pairing, labeling, and connectivity at the structural level, ensuring that the alignment relationships are simultaneously valid and directly searchable in the three dimensions of time, space, and semantics.
[0081] This embodiment extracts semantic embedding representations of language units and visual embedding representations of visual units within a unified reference space, filters out valid candidates using temporal and spatial constraints, and then uses semantic distance to drive structural adjustments such as rearrangement, merging, and pruning to obtain semantically aligned language reasoning chains and semantically aligned visual reasoning trajectories, which are then combined into aligned visual thought chains. This establishes stable and traceable cross-modal correspondences without altering the original evidence, significantly reduces ambiguity caused by mismatches and temporal drift, strengthens the consistent mapping between language progression and spatial changes, and provides directly referable evidence links and measurement records for the subsequent generation of interpretable decision outputs.
[0082] In one embodiment, step S60 above includes: S601, parse the language reasoning chain content in the aligned visual thinking chain and extract key reasoning steps; S602, parse the visual reasoning trajectory content in the aligned visual thinking chain and extract spatial relationship features; S603, integrate the key reasoning steps and the spatial relationship features to generate a preliminary structured report; S604, Perform logical consistency verification on the preliminary structured report and identify potential contradictions; S605, Based on the potential contradictions, revise the preliminary structured report to generate a revised report; S606, Perform a confidence analysis on the revised report and generate a confidence score; S607, Perform an anomaly level analysis on the revised report and generate an anomaly level identifier; S608, the confidence score, the anomaly level identifier, and the corrected report are correlated and integrated to obtain an interpretable decision output.
[0083] In this embodiment, when the aligned visual thought chain enters the interpretation generation stage, content parsing and evidence extraction are first completed. The language reasoning chain is expanded item by item, and verb phrases, conditional phrases, entity references, and causal connections are labeled on a traceable index, forming a continuous set of key reasoning steps; each item carries a time location and source fragment label, maintaining consistency with the alignment relationship. The visual reasoning trajectory is expanded according to time connected segments, and path nodes, posture changes, and object states are mapped to spatial anchor points, forming a set of spatial relationship features; each item retains a dual reference to the path segment and appearance anchor point, ensuring that the mapping with the language side can be directly retrieved. To avoid cross-domain ambiguity, the parsing process follows a unified reference coordinate and a unified naming table, ensuring consistency in reference at the string level and position level for different source fragments.
[0084] After parsing, semantic fusion and structured expression are performed. Key reasoning steps and spatial relationship features are aligned and verified on a unified timeline. Language actions and visual states within the same time span are grouped into entries. Each entry includes an action description, object reference, spatial location, time interval, and source evidence. Entries are linked by event progression relationships to generate a preliminary structured report. Entries are constructed following the single fact unit principle to avoid a single entry carrying multiple events, reducing the difficulty of subsequent consistency judgments. Entries that fail to meet temporal or spatial consistency requirements are marked with a "to be verified" tag, and corresponding weak connection indicators are retained within the entries.
[0085] Logical consistency verification was then conducted. The preliminary structured report was checked segment by segment under three constraints: chronological order, object continuity, and causal sequence. The chronological order check focused on whether the order of occurrence of items was consistent with the linguistic reasoning chain; the object continuity check focused on whether the trajectories of objects with the same name were interrupted or abruptly changed on the visual reasoning trajectory; and the causal sequence check focused on whether there were contradictions or omissions between the preconditions and the result. Any item that did not meet the constraints was recorded as a potential contradiction point, accompanied by alignment evidence and a backtracking index, for use in correction and adjustment.
[0086] In the revision and standardization phase, potential contradictions are addressed according to priority. Temporal contradictions are corrected by merging adjacent entries or splitting excessively long entries; object contradictions are addressed by backtracking the naming table to unify designations and correct trajectory connections; causal contradictions are resolved by supplementing missing entries or removing duplicate entries. The revision and standardization process consistently follows a strategy of prioritizing strong connections and subordinated weak connections within a visual thinking chain, prioritizing the retention of strongly connected entries and folding or replacing weakly connected entries without disrupting overall coherence. Upon completion, a revised report is output, with all entries accompanied by a final evidence backtracking pointer.
[0087] The revised report then undergoes confidence analysis and rating. The confidence score is derived from a combination of three metrics: evidence completeness, connectivity strength, and consistency pass rate. Evidence completeness reflects whether an item possesses both linguistic and visual evidence; connectivity strength reflects the ratio of strong to weak connections; and consistency pass rate reflects the percentage of constraints passed in logical verification. These three metrics are integrated under a unified scale to generate a confidence score, with numerical labels assigned to each item and the entire report. Anomaly rating analysis categorizes items based on the number of conflict remnants, weak connection density, and continuity gaps in key objects, outputting anomaly rating labels. The categorization process does not alter the report content; it only records the rating and its triggering basis at the report's metadata level. Finally, the confidence score, anomaly rating labels, and revised report are integrated to form a directly deliverable and interpretable decision output. The integrated result maintains a bidirectional index relationship with the aligned visual thought chain, allowing any conclusion to be traced back to the original evidence nodes in the linguistic and visual reasoning chains.
[0088] For example, the system mainly includes five modules: multimodal input understanding module; semantic-spatial joint coding module; visual thinking generation module; adaptive constraint mechanism module; and insurance decision interpretability generation module.
[0089] The system can be applied to risk assessment scenarios (such as auto insurance, agricultural insurance, and health insurance) as well as claims review and anti-fraud detection scenarios.
[0090] The multimodal input understanding module receives text descriptions (linguistic modality) and image evidence (visual modality) from insurance scenarios. Examples include: car insurance claim text + on-site accident photos, agricultural insurance disaster description + aerial imagery, and health insurance medical records + examination reports.
[0091] The module first extracts semantic and visual features using a two-stream Transformer structure:
[0092] in , representing the text feature sequence, is the text encoder's pair of features. The output of is a semantic feature representation; , representing a sequence of visual features, is the visual encoder's input to... The output of is a visual feature representation; This represents a text input sequence, which is a text description from an insurance scenario, such as a police report or medical record. It is generally a sequence composed of words or tokens. It represents a sequence of visual inputs, corresponding to image evidence, such as accident scene photos, aerial images, inspection report images, etc., and is usually represented by image patches, detected target areas, or pixel block sequences. This represents a text encoder, which processes text sequences. The encoding network (described in the text as the text branch in a two-stream Transformer) maps discrete word sequences into continuous semantic vector sequences; This represents a visual encoder for image sequences. The encoding network (visual branch of the two-stream Transformer) maps image patches or region features into a sequence of visual vectors.
[0093] Subsequently, cross-modal attention is used to achieve initial fusion of semantics and vision:
[0094] in, This represents a cross-modal attention mechanism, where one of the components is the query (usually text). The other side is key-value pairs (visual). In attention calculations, text is guided to "align" with and "focus on" relevant image regions, thereby achieving interactive fusion of language and vision. This yields a joint feature representation containing both linguistic and spatial information, providing multimodal context for subsequent reasoning.
[0095] The semantic-spatial joint coding module employs a hierarchical semantic aggregation strategy to capture spatial relationships and dynamic changes in insurance events. Word-level aggregation: Local contextual modeling of keywords (such as "left front", "collapsed", "damaged") in the crime report description; Sentence-level aggregation: Integrates event logic in a description using a multi-head attention mechanism; Paragraph-level aggregation: Integrates different descriptive paragraphs through hierarchical attention to form a semantic summary.
[0096] Overall representation of the compressor:
[0097] in These represent attention pooling functions at different levels, enabling information condensation and hierarchical association.
[0098] The Visual Reasoning Generation (VTG) module aims to generate a visual reasoning trajectory, such as a sketch or spatial relationship diagram, representing the vehicle's position, wall orientation, and collision point, while simultaneously outputting the linguistic reasoning chain (e.g., "vehicle veers left → hits wall → front bumper damaged"). The VTG module typically consists of two core parts: Visual thought generator: projects the hidden states of a language model into a visual feature space; outputs a series of visual tokens, which can be input into an image decoder (such as Diffusion or VQ-VAE) to generate a visual image; Formal definition:
[0099] in It is the hidden state vector when decoding to the i-th position, which is calculated by the upstream decoder based on the context, historical language tokens and conditional information. It is a shared representation that drives both language output and visual output. Indicates the effect on The visual projection matrix maps the hidden state from the decoder feature space to the representation space where the visual token is located. The dimension is generally "visual vector dimension × hidden state dimension". This represents the bias vector corresponding to the visual projection, which is compared with the vector during linear transformation. Adding them together yields the complete affine transformation before the visual output; This represents the representation of the i-th visual token.
[0100] Cross-modal synchronous controller: The control language is consistent with the timing of image generation; In generating the language token at step i At the same time, the visual token for step i is generated. The two maintain semantic pairing; Generation probability:
[0101] in, This indicates that the i-th language token (e.g., a word or subword) is represented by the same hidden state. Generated through the language output header; This represents the sequence of all language tokens generated up to the i-th position, i.e. ( , , ..., In autoregressive decoding, this is used as a historical condition input to determine the current... and ; This represents the sequence of all visual tokens generated up to the i-th position, i.e. ( , , ..., As a visual history trajectory, it is used to maintain the temporal continuity of visual output and its alignment with language; This represents the contextual representation or joint features passed in from the encoder, usually from a multimodal encoder or joint feature representation. It serves as a condition throughout the decoding process, globally controlling the generation of the language inference chain and the visual inference trajectory.
[0102] Throughout the decoding process, it serves as a condition for the global control of the generation of language reasoning chains and visual reasoning trajectories.
[0103] The implementation method is generally a "self-regressive decoding + dual-head decoding" structure.
[0104] Example explanation: Suppose the police report is: "The vehicle crashed into a wall while turning at the intersection to the left, and the front of the vehicle was damaged."
[0105] The system inference chain is generated as follows: Step 1: Detect "intersection to the left" → Infer spatial orientation; Step 2: Detecting "collision with wall" → generating collision relationships; Step 3: Infer "damage to the front of the vehicle" → map it to damage in the area in front.
[0106] Meanwhile, Visual Thought Generator generates visual annotations at each step: Step 1: Draw a road and a left-turning arrow in the image; Step 2: Add walls and collision points; Step 3: Highlight the front of the car.
[0107] Ultimately, a visual "reasoning trajectory" is obtained, allowing people to intuitively understand how the model arrives at its judgment step by step.
[0108] Among them, the adaptive constraint mechanism module: A common problem in multimodal models is that "left front" in text may be misinterpreted as "left side" in image space; or textual reasoning may suggest "damaged front of the car," but image attention may be focused on the rear of the car. This indicates that language and visual representation are "out of sync."
[0109] Thought Alignment Loss (TAL) is used to align language tokens and visual tokens in the semantic space at each step. For each language token Computational semantic embedding ; Visual tokens generated simultaneously Computer vision embedding ; Calculate the Euclidean distance (or cosine distance) between the two:
[0110] in, This represents the i-th language token, corresponding to the word or sub-word generated in the i-th step of the language inference chain; This represents the i-th visual token, which corresponds to the visual unit generated in the i-th step of the visual inference trajectory, such as a certain area of interest, trajectory point, etc. Language embedding vectors represent language tokens. The vector representation obtained after inputting into the language embedding map is located in a continuous vector space and is used to characterize the semantic state of the language at that time step. The visual embedding vector represents the visual token. The vector representation obtained after inputting into the visual embedding mapping is located in the same or comparable vector space as the language embedding, and is used to characterize the spatial / appearance state of the visual end at that time step. The Euclidean distance (squared) between the linguistic embedding and the visual embedding is represented. The smaller the distance, the closer the linguistic token and the visual token are in the embedding space, that is, the more consistent their semantic and visual meanings are. This indicates the number of paired tokens participating in the alignment calculation, which is the number of samples or time steps used for summation and is used for averaging.
[0111] This loss term allows the model to continuously adjust during training, ensuring that the generated text and images are semantically similar. The final result is: The language token "left front" will correspond to the visual token of the "top left area" in the image; "Impact" corresponds to the visual marker of "the point of contact between objects"; "Damage" corresponds to "abnormal areas on the surface".
[0112] During inference, this loss allows the model to learn to dynamically synchronize the rhythm of language and vision.
[0113] That is, when the model describes the progress of an event in language (such as "from turning → impact → stop"), the image content of the visual generation part will also change continuously, forming a dynamic trajectory similar to "mind animation", rather than a static frame.
[0114] Among them, the insurance decision interpretability generation module outputs a structured decision report based on the V-CoT reasoning results, including: linguistic reasoning chain (cause of accident, liability attribution, risk evolution); visual reasoning trajectory (spatial changes, damage distribution); confidence level and risk grade score.
[0115] The final result is a three-dimensional decision report consisting of text, images, and weighted explanations, making the claims process transparent and auditable. For example, the system can output: "Based on the accident description and image evidence, the model inference is 'left front tire blowout → vehicle veers to the right → collision with the wall'. The spatial relationship is consistent, the responsibility ratio is 80% primary responsibility, and the confidence level is 0.92." This embodiment analyzes and aligns the visualized thought chain under the same reference frame, extracts key reasoning steps and spatial relationship features, and integrates them into items. Then, it completes consistency verification and correction by using three types of constraints: time sequence, object continuity, and causal inheritance. Finally, it generates confidence scores and anomaly level labels for the corrected report and integrates them with the content for output. This achieves the integrated presentation and step-by-step backtracking of linguistic and visual evidence, keeps the explanatory text and evidence chain synchronized and searchable, reduces ambiguity and contradictions in conclusions, and improves the credibility and review efficiency of the delivered results.
[0116] In one embodiment, a decision-making device based on a visual thinking chain is provided, which corresponds one-to-one with the decision-making method based on a visual thinking chain in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the decision-making device based on a visual thinking chain of the present invention. The modules include a multimodal input module 10, a multimodal feature fusion module 20, a joint encoding module 30, a thinking chain generation module 40, a thinking chain alignment module 50, and a decision generation module 60. Detailed descriptions of each functional module are as follows: The multimodal input module 10 is used to receive multimodal data including text data and image data; The multimodal feature fusion module 20 is used to perform feature extraction and fusion processing on the text data and image data to obtain multimodal fused features; The joint encoding module 30 is used to perform semantic and spatial joint encoding processing on the multimodal fusion features to obtain a joint feature representation; The thought chain generation module 40 is used to generate a visual thought chain including a language reasoning chain and a visual reasoning trajectory based on the joint feature representation. The thought chain alignment module 50 is used to perform adaptive constraint alignment processing on the language reasoning chain and visual reasoning trajectory of the visualized thought chain to obtain the aligned visualized thought chain. The decision generation module 60 is used to generate interpretable decision outputs based on the aligned visual thought chain.
[0117] In one embodiment, the multimodal input module 10 is specifically used for: Receive text and image data streams through the data input interface; The text data stream is parsed to obtain text data; The image data stream is parsed to obtain the image data; Establish the association between the text data and the image data; The text data and image data that have a relationship are treated as multimodal data.
[0118] In one embodiment, the multimodal feature fusion module 20 is specifically used for: The semantic features of the text data are extracted using a text encoder; Visual features of the image data are extracted using an image encoder; The semantic features and the visual features are interactively fused through a cross-modal attention mechanism to obtain interactive fused features; The interactive fusion features are subjected to feature enhancement processing to obtain enhanced fusion features, which are then used as multimodal fusion features.
[0119] In one embodiment, the joint coding module 30 is specifically used for: The multimodal fusion features are subjected to word-level semantic aggregation processing to extract local contextual information and obtain word-level aggregated features; The word-level aggregation features are subjected to sentence-level semantic aggregation processing to obtain sentence-level aggregation features; The sentence-level aggregation features are subjected to paragraph-level semantic aggregation processing to obtain paragraph-level aggregation features; The paragraph-level aggregated features are subjected to hierarchical attention compression processing to obtain a joint feature representation.
[0120] In one embodiment, the thought chain generation module 40 is specifically used for: The joint feature representation is input into the language inference chain generator to generate a sequence of language inference units; The joint feature representation is input into the visual reasoning trajectory generator to generate a sequence of visual reasoning units; The language reasoning unit sequence and the visual reasoning unit sequence are processed using a cross-modal synchronization controller to generate a time-coordinated language reasoning unit sequence and a time-coordinated visual reasoning unit sequence. The sequence of time-coordinated language reasoning units is combined into a language reasoning chain; The sequence of time-coordinated visual reasoning units is combined into a visual reasoning trajectory; The linguistic reasoning chain and the visual reasoning trajectory are combined into a visual thinking chain.
[0121] In one embodiment, the thought chain alignment module 50 is specifically used for: Extract the semantic embedding representation of language units in the language inference chain; Extract the visual embedding representation of the visual unit in the visual reasoning trajectory; Determine the semantic distance between the semantic embedding representation of the language unit and the visual embedding representation of the visual unit; Based on the semantic distance, the alignment relationship between the language inference chain and the visual inference trajectory is adjusted to generate a semantically aligned language inference chain and a semantically aligned visual inference trajectory. The semantically aligned linguistic reasoning chain and the semantically aligned visual reasoning trajectory are combined into an aligned visual thought chain.
[0122] In one embodiment, the decision generation module 60 is specifically used for: Analyze the linguistic reasoning chain content in the aligned and visualized thought chain, and extract key reasoning steps; Analyze the visual reasoning trajectory content in the aligned visualized thought chain and extract spatial relationship features; By integrating the key reasoning steps and the spatial relationship features, a preliminary structured report is generated; The preliminary structured report is subjected to logical consistency verification to identify potential contradictions; The preliminary structured report is revised based on the potential contradictions to generate a revised report; A confidence analysis is performed on the revised report to generate a confidence score; Anomaly level analysis is performed on the revised report to generate anomaly level identifiers; The confidence score, the anomaly level identifier, and the corrected report are correlated and integrated to obtain an interpretable decision output.
[0123] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps of a decision-making method based on a visual thought chain.
[0124] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides decision-making and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of a decision-making method based on a visual thought chain.
[0125] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Receives multimodal data, including text and image data; Feature extraction and fusion processing are performed on the text data and image data to obtain multimodal fusion features; The multimodal fusion features are subjected to joint semantic and spatial encoding to obtain a joint feature representation; Based on the joint feature representation, a visual thought chain including a language reasoning chain and a visual reasoning trajectory is generated. Adaptive constraint alignment processing is performed on the language reasoning chain and visual reasoning trajectory of the visualized thinking chain to obtain the aligned visualized thinking chain. Based on the aligned visual thought chain, an interpretable decision output is generated.
[0126] In one embodiment, a non-volatile computer-readable storage medium is provided, which may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, it performs the following steps: Receives multimodal data, including text and image data; Feature extraction and fusion processing are performed on the text data and image data to obtain multimodal fusion features; The multimodal fusion features are subjected to joint semantic and spatial encoding to obtain a joint feature representation; Based on the joint feature representation, a visual thought chain including a language reasoning chain and a visual reasoning trajectory is generated. Adaptive constraint alignment processing is performed on the language reasoning chain and visual reasoning trajectory of the visualized thinking chain to obtain the aligned visualized thinking chain. Based on the aligned visual thought chain, an interpretable decision output is generated.
[0127] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0128] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0129] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0130] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
[0131] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.
Claims
1. A decision-making method based on visualized mind chains, characterized in that, The method comprises the following steps: receiving multi-modal data including text data and image data; performing feature extraction and fusion processing on the text data and image data to obtain multi-modal fusion features; performing semantic and spatial joint coding processing on the multi-modal fusion features to obtain joint feature representation; based on the joint feature representation, generating a visual thinking chain including a language reasoning chain and a visual reasoning track; performing adaptive constraint alignment processing on the language reasoning chain and the visual reasoning track of the visual thinking chain to obtain an aligned visual thinking chain; based on the aligned visual thinking chain, generating an interpretable decision output.
2. The decision making method based on the visualization mind chain according to claim 1, characterized in that, Receiving multi-modal data including text data and image data includes: receiving a text data stream and an image data stream through a data input interface; performing format analysis on the text data stream to obtain text data; performing format analysis on the image data stream to obtain image data; establishing an association between the text data and the image data; the text data and the image data with the association as multi-modal data.
3. The decision making method based on the visualization mind chain according to claim 1, characterized in that, Performing feature extraction and fusion processing on the text data and image data to obtain multi-modal fusion features includes: extracting semantic features of the text data through a text encoder; extracting visual features of the image data through an image encoder; interactively fusing the semantic features and the visual features through a cross-modal attention mechanism to obtain interactive fusion features; performing feature enhancement processing on the interactive fusion features to obtain enhanced fusion features, and taking the enhanced fusion features as multi-modal fusion features.
4. The decision making method based on the visualization mind chain according to claim 1, characterized in that, Performing semantic and spatial joint coding processing on the multi-modal fusion features to obtain joint feature representation includes: performing word-level semantic aggregation processing on the multi-modal fusion features to extract local context information to obtain word-level aggregated features; performing sentence-level semantic aggregation processing on the word-level aggregated features to obtain sentence-level aggregated features; performing paragraph-level semantic aggregation processing on the sentence-level aggregated features to obtain paragraph-level aggregated features; performing hierarchical attention compression processing on the paragraph-level aggregated features to obtain joint feature representation.
5. The method of claim 1, wherein the visualization-based mind mapping is based on a decision tree. Based on the joint feature representation, generating a visual thinking chain including a language reasoning chain and a visual reasoning track includes: inputting the joint feature representation into a language reasoning chain generator to generate a language reasoning unit sequence; inputting the joint feature representation into a visual reasoning track generator to generate a visual reasoning unit sequence; using a cross-modal synchronization controller to process the language reasoning unit sequence and the visual reasoning unit sequence to generate a time-coordinated language reasoning unit sequence and a time-coordinated visual reasoning unit sequence; combining the time-coordinated language reasoning unit sequence into a language reasoning chain; combining the time-coordinated visual reasoning unit sequence into a visual reasoning track; combining the language reasoning chain and the visual reasoning track into a visual thinking chain.
6. The method of claim 1, wherein the visualization-based mind mapping is based on a decision tree. Performing adaptive constraint alignment processing on the language reasoning chain and the visual reasoning track of the visual thinking chain to obtain an aligned visual thinking chain includes: extracting language unit semantic embedding representation in the language reasoning chain; extracting a visual unit visual embedding representation in the visual reasoning track; determining a semantic distance between the language unit semantic embedding representation and the visual unit visual embedding representation; adjusting an alignment relationship of the language reasoning chain and the visual reasoning track based on the semantic distance, to generate a semantically aligned language reasoning chain and a semantically aligned visual reasoning track; combining the semantically aligned language reasoning chain and the semantically aligned visual reasoning track into an aligned visual thinking chain.
7. The method of claim 1, wherein the visualization-based mind mapping is based on a decision tree. based on the aligned visual thinking chain, generating an interpretable decision output, including: parsing the language reasoning chain content in the aligned visual thinking chain, to extract key reasoning steps; parsing the visual reasoning track content in the aligned visual thinking chain, to extract spatial relationship features; fusing the key reasoning steps and the spatial relationship features, to generate a preliminary structured report; performing logical consistency verification on the preliminary structured report, to identify potential contradictions; based on the potential contradictions, correcting the preliminary structured report, to generate a corrected report; performing confidence analysis on the corrected report, to generate a confidence score; performing abnormal level analysis on the corrected report, to generate an abnormal level identifier; integrating the confidence score, the abnormal level identifier, and the corrected report, to obtain an interpretable decision output.
8. A decision device based on a visualized chain of thought, characterized in that The decision device based on a visual thinking chain includes: a multi-modal input module configured to receive multi-modal data including text data and image data; a multi-modal feature fusion module configured to perform feature extraction and fusion processing on the text data and image data, to obtain multi-modal fusion features; a joint encoding module configured to perform semantic and spatial joint encoding processing on the multi-modal fusion features, to obtain joint feature representations; a thinking chain generation module configured to generate a visual thinking chain including a language reasoning chain and a visual reasoning track based on the joint feature representations; a thinking chain alignment module configured to perform adaptive constraint alignment processing on the language reasoning chain and the visual reasoning track of the visual thinking chain, to obtain an aligned visual thinking chain; a decision generation module configured to generate an interpretable decision output based on the aligned visual thinking chain.
9. A computer device, comprising: The computer device includes a memory, a processor, and a decision program based on a visual thinking chain stored on the memory and executable on the processor, and the decision program based on a visual thinking chain, when executed by the processor, implements the steps of the decision method based on a visual thinking chain of any one of claims 1-7.
10. A non-transitory computer readable storage medium, comprising: The storage medium has a decision program based on a visual thinking chain stored thereon, and the decision program based on a visual thinking chain, when executed by the processor, implements the steps of the decision method based on a visual thinking chain of any one of claims 1-7.