A Dynamic Adaptive Cross-Modal Reasoning Retrieval Method for the Maintenance Field
Patent Information
- Application Number
- CN202610886015.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-08-14
AI Technical Summary
[0006]发明目的,提供一种面向维修领域的动态自适应跨模态推理检索方法,以期能够解决现有技术存在的至少部分问题
[0014] Beneficial effects: This invention achieves unified semantic alignment and dynamic adaptive retrieval of multimodal knowledge in the maintenance field, improving the accuracy of cross-modal retrieval, the adaptability to maintenance context, and the reliability of results.
Smart Images

Figure CN122570722A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent maintenance and knowledge retrieval technology, and in particular to a dynamic adaptive cross-modal reasoning retrieval method for the maintenance field. Background Technology
[0002] In the field of equipment maintenance and support, maintenance personnel typically need to consult a large number of technical documents to obtain fault diagnosis and maintenance operation guidance. Traditional maintenance knowledge retrieval methods mainly rely on keyword matching and document indexing. Maintenance personnel need to manually search for relevant information from scattered sources such as manuals, drawings, and logs, which is inefficient and prone to missing key information.
[0003] In recent years, with the popularization of multimodal data, maintenance knowledge bases have accumulated a large amount of unstructured data, including not only text-based maintenance manuals and technical logs, but also equipment 3D models (such as CAD drawings and STEP files), fault scene images, and abnormal operation audio. However, most existing retrieval technologies are designed for a single modality: text retrieval cannot handle the geometric structure information of 3D models, image retrieval struggles to understand the temporal features in audio, and the semantic gap between different modalities prevents maintenance personnel from obtaining comprehensive cross-modal search results.
[0004] The emergence of Retrieval Augmentation (RAG) technology has provided new ideas for knowledge retrieval. However, existing RAG solutions are mainly geared towards a single text modality. Even with preliminary attempts at multimodal approaches, they often employ fixed mappings of "text to image" or "text to audio," lacking a unified cross-modal semantic alignment mechanism. Furthermore, the maintenance process is dynamically evolving: from fault detection to component location and then to maintenance execution, the emphasis on information varies at different stages, while existing retrieval strategies are static and cannot be dynamically adjusted according to the maintenance context. Maintenance personnel with different experience levels also have significantly different needs for information granularity, and existing solutions lack personalized cognitive adaptation capabilities. More importantly, when text descriptions are inconsistent with the 3D model structure, or when audio features contradict image fault features, existing technologies cannot detect and resolve these multimodal conflicts, potentially outputting misleading search results.
[0005] Therefore, there is an urgent need for an intelligent retrieval method that can uniformly process multimodal data such as text, 3D models, images, and audio, and has the ability to dynamically perceive context, adapt cognition, and resolve conflicts, so as to improve the accuracy and efficiency of knowledge acquisition in the field of equipment maintenance. Summary of the Invention
[0006] The purpose of this invention is to provide a dynamic adaptive cross-modal reasoning retrieval method for the maintenance field, in order to solve at least some of the problems existing in the prior art.
[0007] Technical solution: A dynamic adaptive cross-modal reasoning retrieval method for the maintenance field, comprising the following steps:
[0008] Acquire multimodal raw data, which includes at least maintenance text, equipment 3D model, fault images, and abnormal audio.
[0009] Based on multimodal raw data, a unified semantic anchor space is constructed, which maps data from different modalities to the same vector space, generating a cross-modal semantic index library;
[0010] Based on the dynamic context of the maintenance process, a maintenance status graph is constructed, and the current maintenance status is inferred in real time. The retrieval strategy is switched dynamically with the cross-modal semantic index based on the current maintenance status.
[0011] Based on user profiles and real-time behavior, a cognitive adaptation retrieval layer is established to dynamically adjust the output granularity and presentation format of retrieval results.
[0012] Multimodal confidence conflict analysis is performed on search results from different modalities to detect and fuse semantic, spatiotemporal, or logical conflicts between modalities, generating fused results with confidence scores;
[0013] Based on the fusion results and output granularity, a personalized multimodal retrieval response is generated.
[0014] Beneficial effects: This invention achieves unified semantic alignment and dynamic adaptive retrieval of multimodal knowledge in the maintenance field, improving the accuracy of cross-modal retrieval, the adaptability to maintenance context, and the reliability of results. Attached Figure Description
[0015] Figure 1 This is a flowchart of the overall solution of the present invention.
[0016] Figure 2 This is a flowchart of the process for constructing a unified semantic anchor space according to the present invention.
[0017] Figure 3 This is a flowchart of the present invention for constructing a maintenance status diagram and dynamically switching retrieval strategies.
[0018] Figure 4 This is a flowchart of the cognitive adaptation retrieval layer established in this invention.
[0019] Figure 5 This is a flowchart of the multimodal belief conflict analysis of the present invention. Detailed Implementation
[0020] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0021] like Figures 1-5 As shown in this embodiment, the data processing flow of a dynamic adaptive cross-modal reasoning retrieval method for the maintenance field is described in detail, specifically including:
[0022] S1: Obtain multimodal raw data, which includes at least maintenance text, equipment 3D model, fault images, and abnormal audio.
[0023] In one aspect of this application, a multimodal raw data acquisition operation is first performed. The acquired data covers at least four dimensions of maintenance information sources: maintenance text, equipment 3D model, fault images, and abnormal audio.
[0024] Specifically, the sources of maintenance documents include, but are not limited to: equipment maintenance manuals, PDF technical documents, Word maintenance logs, and scanned copies. These documents contain key information such as descriptions of fault symptoms, cause analysis, troubleshooting steps, and parts lists.
[0025] Furthermore, the sources of the equipment's 3D models include: CAD design drawings, STEP format 3D exchange files, and Web3D interactive models (such as 3Dhtml format). These models record the equipment's geometry, component hierarchy, and assembly constraints.
[0026] In a preferred embodiment, the fault images are collected from the repair site or testing environment, specifically including: on-site photos of the faulty part, instrument panel readings, close-ups of burnt circuit boards, etc. Each image is accompanied by metadata such as the shooting time, equipment identification, and fault label.
[0027] For example, the abnormal audio is collected as sound signals of equipment in normal operation and abnormal conditions, such as abnormal engine noise, gear grinding noise, bearing squealing noise, etc. The audio data is recorded at a high sampling rate (e.g., 44.1kHz) and saved as uncompressed WAV format to preserve complete time and frequency domain information.
[0028] Through the above collection steps, a diverse and modal-rich raw dataset is obtained, laying the data foundation for subsequent unified semantic processing.
[0029] S2: Based on multimodal raw data, a unified semantic anchor space is constructed, which maps data from different modalities to the same vector space, generating a cross-modal semantic index library.
[0030] According to one aspect of this application, after obtaining the multimodal raw data, a unified semantic anchor space construction operation is performed to map the data of different modalities to the same vector space, thereby forming a semantic index library that can be retrieved across modalities.
[0031] Specifically, for maintenance text data, a fine-grained semantic role labeling technique is first employed to extract "subject-action-object" triples from each text segment. For example, from "replacing a worn bearing," the triples are extracted as (subject: operator, action: replacement, object: bearing). Subsequently, the extracted triples are input into a pre-trained BERT encoder to generate a 256-dimensional text semantic anchor vector.
[0032] Furthermore, for the equipment 3D model data, the hierarchical tree of the model is first parsed to extract the name, parent-child assembly relationship, and geometric features of each component node. The geometric features are obtained from the voxelized representation of the model using a 3D convolutional neural network. These geometric features are then concatenated with the descriptive text of the component nodes (such as "main shaft bearing" and "end cap bolt"), and mapped through a fully connected layer to generate a 256-dimensional 3D semantic anchor vector.
[0033] In a preferred embodiment, for fault image data, a visual Transformer model is first used to segment salient regions in the image, such as fault locations, instrument dials, or burn marks. For each salient region, scene description text output by a visual language model (such as BLIP) (e.g., "capacitor bulge on circuit board") is fused to generate a 256-dimensional image semantic anchor vector.
[0034] For example, for anomalous audio data, the original audio signal is first pre-emphasized, framed, and windowed. Then, a Mel spectrogram is generated through short-time Fourier transform. This spectrogram is input into a convolutional-transformer hybrid model, which simultaneously extracts joint features in the time and frequency domains and identifies the audio event type (such as "gear friction sound"). The final output is a 256-dimensional audio semantic anchor vector.
[0035] In this embodiment, after generating the anchor vectors for each modality, a contrastive learning training framework is further constructed. Clearly corresponding multimodal pairs are collected from maintenance documents as positive samples, such as text snippets describing "bearing wear," images showing bearing wear, audio clips recording bearing noise, and semantic vectors of bearing nodes in the 3D model. These four constitute a positive sample quadruple. Vectors from different sources are randomly combined as negative samples. A mapping network is trained using the InfoNCE contrastive loss function to minimize the vector distance between positive sample pairs and maximize the distance between negative sample pairs. After training convergence, the semantic anchor vectors of all modalities reside in the same metric space, forming a cross-modal semantic index. This index supports retrieval from any modality to any modality, such as "searching for 3D nodes by audio" or "searching for maintenance text by image."
[0036] S3: Based on the dynamic context of the maintenance process, construct a maintenance status graph and infer the current maintenance status in real time. Switch retrieval strategies dynamically with the cross-modal semantic index based on the current maintenance status.
[0037] In this application, in order to achieve dynamic retrieval oriented towards the maintenance process, a maintenance status graph is constructed, and the current maintenance stage is inferred in real time based on the graph, and then the retrieval strategy is dynamically switched according to the current status.
[0038] Specifically, the standard state nodes in the maintenance process are first defined. These nodes include at least: fault detection, component location, disassembly operation, maintenance execution, and verification testing. Each node represents a specific stage of the maintenance operation. Subsequently, directed edges are established between the nodes, with each edge accompanied by a trigger condition. For example, the trigger condition for the transition from "fault detection" to "component location" is "a definite fault phenomenon has been detected." This forms the initial maintenance state diagram.
[0039] Furthermore, data is collected from historical maintenance logs, including user operation sequences (such as query keywords, clicked search results, and confirmation steps), timestamps for each operation, and the final labeled actual maintenance stage. Using this data, a Long Short-Term Memory (LSTM) network is trained as a state inferrer. The input to this LSTM model is the feature vector of the most recent operations in the current session, and the output is the most likely maintenance state and its probability.
[0040] In a preferred embodiment, during runtime, each time a user query or interaction is received, the state inferrer is invoked to obtain the current maintenance state (e.g., "part location"). Then, based on the current maintenance state, the corresponding priority retrieval modality is read from a pre-built state-policy mapping table. This mapping table is defined as follows: if the current state is "fault detection," audio and image modalities are retrieved first, because abnormal noises and visual fault features are the earliest clues; if the current state is "part location," 3D modalities and knowledge graphs are retrieved first to quickly find the location and relationships of the faulty part; if the current state is "disassembly operation" or "maintenance execution," text modalities (maintenance steps) and 3D highlights (indicating disassembly locations) are retrieved first; if the current state is "verification test," all modalities are integrated for a comprehensive retrieval to confirm the maintenance effect.
[0041] For example, when the state inferrer outputs "fault detection" with a probability higher than 0.8, the retrieval strategy automatically switches to: first, performing a joint audio-image retrieval in the generated cross-modal semantic index library, calculating the similarity between the audio feature vector and the image feature vector and the vectors in the library respectively, and returning the most matching abnormal sound sample and the corresponding fault image. Only when the audio or image retrieval results are empty is text retrieval used as a fallback.
[0042] The above design enables the synchronous evolution of retrieval strategies and maintenance processes, avoiding information redundancy or missing information in dynamic maintenance scenarios caused by static retrieval.
[0043] S4: Based on user profiles and real-time behavior, establish a cognitive adaptation retrieval layer to dynamically adjust the output granularity and presentation format of retrieval results.
[0044] In this application, there are significant differences in the knowledge level and operating habits of different maintenance personnel. Therefore, it is necessary to establish a cognitive adaptation retrieval layer to achieve personalized and adaptive result presentation.
[0045] Specifically, the system first collects users' historical behavioral data, including: the length of each query (in words), the time spent on each document, the frequency of interaction with the 3D model (e.g., rotations, zooms), and the number of error responses (e.g., clicking the "irrelevant" button). Based on these characteristics, the K-means clustering algorithm is used to divide users into three categories: experts, intermediate, and beginners. Simultaneously, based on whether users prefer viewing images or reading text during interaction, their preference type (visual preference or text preference) is labeled.
[0046] Furthermore, during a single maintenance session, the system monitors the user's immediate behavior in real time. An "insufficient information" condition is identified when one of the following occurs: the user spends more than a preset threshold (e.g., 10 seconds) on a search result; or the user immediately initiates a second query after viewing the result with a keyword similarity higher than 0.7. Once insufficient information is determined, a granularity enhancement function is invoked. This function automatically expands the originally output brief information (e.g., only the part name and highlighted position) into detailed content, including: a step-by-step operation sequence with illustrations, a fault audio waveform diagram with timeline markers, and a disassembly path animation in a 3D model.
[0047] In a preferred embodiment, multimodal output cards are dynamically assembled based on user category and current granularity state. The specific rules are as follows:
[0048] For users categorized as "newbies", the output cards include: a text description of the operation steps (each step accompanied by a screenshot), a highlighted display of the operated component in the 3D model, a waveform diagram of the fault audio and playback controls, and a fault cause and spare parts list derived from the knowledge graph.
[0049] For users categorized as "experts", the output cards are simplified to include only: 3D node highlighting (directly locating the faulty component), fault code (such as "BRG-001"), and minimum action set (such as "remove end cap → replace bearing"), without any redundant descriptions.
[0050] For "intermediate" users, the output cards fall somewhere in between and can be dynamically adjusted during the session based on their real-time behavior.
[0051] For example, if a user is marked as "visual preference" and is "intermediate", even if their query text is short, the system will prioritize placing 3D model screenshots and fault images in the output card, with text steps presented in short sentences below the images.
[0052] Through the cognitive adaptation retrieval layer, this solution can effectively match the user's cognitive load and information needs, thereby improving maintenance efficiency and user experience.
[0053] S5: Perform multimodal confidence conflict analysis on the search results from different modalities, detect and fuse semantic conflicts, spatiotemporal conflicts or logical conflicts between modalities, and generate fused results with confidence scores.
[0054] In this application, there may be inconsistencies or even contradictions between the search results of different modalities. Therefore, a multimodal confidence conflict resolution mechanism is introduced to detect conflicts and generate credible fusion results.
[0055] Specifically, initial confidence scores are first calculated for the retrieval results of four modalities: text, 3D, image, and audio. The confidence score for the text modality is based on a weighted average of the BM25 score and vector similarity; the confidence score for the 3D modality is based on the matching degree of component names and the cosine similarity of geometric features; the confidence score for the image modality uses the classification probability output by the visual language model; and the confidence score for the audio modality uses the softmax output value of the audio event classifier. This results in a four-channel confidence vector, denoted as C_text, C_3d, C_img, and C_audio.
[0056] Furthermore, the retrieval results from the four modalities (including their respective semantic anchor vectors and metadata) are input into a pre-trained Siamese Transformer network. This network outputs two quantities: the conflict probability p_conf (ranging from 0 to 1) and the conflict type. There are three conflict types: semantic conflict (e.g., the text describes "the bearing is removable," but the 3D model shows the bearing node as "non-removable"); spatiotemporal conflict (e.g., the audio indicates "gear noise," but the image shows the gear is intact); and logical conflict (e.g., the knowledge graph infers "replace the bearing," but the text steps require "remove the end cap first," contradicting the preconditions in the knowledge graph).
[0057] In a preferred embodiment, an external maintenance rule base is invoked based on the detected conflict type. This rule base contains conflict resolution priorities predefined by domain experts. For example, for semantic conflicts, the rule base stipulates that "the structural attributes of the 3D model shall prevail, unless the text source is the official maintenance manual"; for spatiotemporal conflicts, it stipulates that "the one with higher confidence among the real-time collected audio and images shall prevail, and the user shall be prompted to verify on-site."
[0058] For example, a weighted fusion formula is used to correct the four-channel confidence vector to obtain the final confidence score C_final: C_final = ( w_text * C_text + w_3d * C_3d + w_img * C_img + w_audio * C_audio ) / (1 + alpha * p_conf); where w_text, w_3d, w_img, and w_audio are the dynamic weights of each modality, all initially set to 0.25; alpha is the conflict penalty coefficient, ranging from 0 to 1, and is set to 0.5 in this embodiment. When the conflict probability p_conf is high, the denominator increases, and the final confidence score is correspondingly reduced.
[0059] Furthermore, when the conflict probability p_conf exceeds a preset threshold (e.g., 0.7), a conflict warning text is generated, such as: "Note: The audio indicates abnormal gear noise, but no obvious abnormality is seen in the gear in the picture. It is recommended to check the inside of the gearbox or verify the source of the audio." Finally, the corrected C_final value and the conflict warning text are encapsulated together to form the fused result.
[0060] By using conflict resolution steps, the misleading influence of conflicting multimodal information on maintenance decisions is effectively avoided, thus improving the reliability of search results.
[0061] S6: Generate personalized multimodal retrieval responses based on the fusion results and output granularity.
[0062] In this application, a personalized multimodal retrieval response is generated based on the fusion results and the determined output granularity, tailored to the current user and the current maintenance scenario.
[0063] Specifically, the fused output is used as the final candidate set of search content, sorted from high to low according to the C_final value. Then, based on the determined user category (novice, intermediate, expert) and the current granularity (brief / detailed), an appropriate number of results are selected from the candidate set for assembly.
[0064] Furthermore, the assembly process follows these principles: If the user category is "Beginner" and the granularity is "Detailed," a multimodal response containing complete information is generated: the top text step area displays step-by-step instructions, the right area displays a highlighted view of the 3D model (interactively rotatable), and the lower area sequentially displays fault image thumbnails, an audio waveform player, and a reasoning path diagram from the knowledge graph. If the user category is "Expert" and the granularity is "Simplified," a minimalist response is generated: only the highlighted faulty component in the 3D model, the fault code (e.g., "ERR-202"), and a single, concise instruction (e.g., "Replace bearing SKF-6204"). If the user category is "Intermediate" or the granularity is a dynamically adjusted intermediate value, a compromise response is generated, retaining an "Expand Details" button for users to actively expand.
[0065] In one alternative implementation, the final response also includes a visual indicator to indicate the presence of a conflict warning. When a conflict warning text is generated, the text is displayed at the top of the response interface in a yellow or red alert bar, along with suggested actions (e.g., "Click to review audio source").
[0066] For example, suppose that in a query's fusion results, both the text and 3D modalities point to "replacing the spindle bearing," but the audio modal has low confidence and detects a spatiotemporal conflict. The system generates the following response: the spindle bearing location is highlighted in red in the 3D model; the text area displays "Recommend replacing the bearing (confidence 0.85)"; and a yellow warning bar is displayed at the top: "Audio sample does not match the image; manual verification of the source of the abnormal noise is recommended." Users can click the warning bar to view detailed conflict explanations.
[0067] According to a further improvement of this embodiment, after generating a response, the system records the user's subsequent actions (such as clicking a highlighted component, playing audio, confirming a step, or reporting an error) as implicit feedback. This feedback data will be used periodically to fine-tune the mapping network in S2, the state inferrer in S3, and the modal dynamic weights in S5, enabling the continuous evolution of the entire retrieval framework.
[0068] According to another aspect of this application, constructing a unified semantic anchor space includes:
[0069] S2.1: Extract triples from the maintenance text and generate text semantic anchor vectors through semantic role labeling and encoder.
[0070] In one aspect of this application, semantic triple extraction is first performed on the acquired maintenance text data. Specifically, the maintenance text includes fault description paragraphs in equipment maintenance manuals, maintenance procedures in PDF technical documents, fault records in Word logs, etc. A fine-grained semantic role labeling model (e.g., a Transformer-based SRL framework) is used to analyze each sentence, identifying the predicate verb and its associated agent and object, forming triples in the form of "subject-action-object". For example, for the sentence "maintenance personnel replace worn bearings", the extracted triples are (subject: maintenance personnel, action: replacement, object: bearing). For passive sentences or imperative sentences without a clear subject, "system" or "operator" is used as the default subject. After triple extraction, each element in the triple is converted into a word embedding and concatenated into a fixed-length vector sequence. This sequence is input into a pre-trained BERT encoder, and the output of the last layer [CLS] position is taken as the semantic representation of the entire triple. This vector representation has a dimension of 256, which is the text semantic anchor vector. Through the above processing, each piece of maintenance text is transformed into a semantic point in a high-dimensional space, providing a unified text feature representation for subsequent cross-modal alignment.
[0071] S2.2: Parse the hierarchical structure tree from the equipment 3D model, extract the component nodes and their geometric features and descriptive text, and generate 3D semantic anchor vectors.
[0072] According to a further improvement of this embodiment, for the collected equipment 3D model data (such as CAD drawings, STEP files, or Web3D models), a model parsing operation is first performed. Specifically, the assembly tree structure of the model is read, which records the parent-child relationships and assembly constraints between parts. Starting from the root node (final assembly), each part node is recursively traversed to extract it. Each node contains at least the following information: part name (e.g., "spindle bearing", "end cap bolt"), node ID, parent node reference, and a list of child nodes. Further, for each part node, its geometric features are extracted. In a preferred embodiment, the 3D model is voxelized, i.e., converted into a fixed-size voxel mesh (e.g., 64×64×64), and then input into a pre-trained 3D convolutional neural network, outputting a 128-dimensional geometric feature vector. Simultaneously, the descriptive text of the part nodes (e.g., explanatory text obtained from model metadata or associated documents) is mapped to another 128-dimensional text feature vector through a lightweight text encoder. The geometric feature vector and the text feature vector are concatenated to obtain the original 256-dimensional features. Finally, the original features are input into a fully connected layer for nonlinear transformation, outputting the final 3D semantic anchor vector. In this way, each component node is encoded as a point in the semantic space, which simultaneously contains information about its geometric shape and semantic description.
[0073] S2.3: Segment salient regions from fault images, fuse scene descriptions output by visual language models, and generate image semantic anchor vectors.
[0074] In one alternative implementation, for fault image data (such as photos of the fault scene, dashboard images, and close-ups of circuit boards), a salient region segmentation operation is first performed. Specifically, a visual Transformer-based image segmentation model is used to divide the input image into multiple semantic regions. This model can automatically identify key regions in the image, such as the charred circuit board area, the tire tread area with cracks, and the abnormal reading area on the dashboard. Each salient region is cropped and scaled to a uniform size (e.g., 224×224 pixels). Then, each region image is input into two parallel processing branches: the first branch uses a visual Transformer to extract the visual features of the image, outputting a 128-dimensional visual vector; the second branch uses a visual language model (such as BLIP or CLIP) to generate a scene description text for the region, such as "the capacitor on the circuit board is bulging with black charred marks." This description text is further converted into another 128-dimensional semantic vector by a text encoder. The visual vector and the semantic vector are concatenated to obtain the original 256-dimensional region features. If multiple salient regions exist within a single image, a weighted average is calculated using the feature vectors of these regions. The weights are determined by the region area or the salientity score output by the model. This results in the image's semantic anchor vector. This vector not only reflects the visual content of the image but also incorporates semantic information from natural language, facilitating subsequent alignment with the text modality.
[0075] S2.4: Extract Mel spectrograms from abnormal audio, identify audio event types using a convolution-transformer hybrid model, and generate audio semantic anchor vectors.
[0076] According to one aspect of this application, for abnormal audio data (such as engine noise, gear grinding noise), audio preprocessing and spectrum transformation are first performed. Specifically, the original audio signal is read in at a sampling rate of 44.1 kHz and pre-emphasized filtering is applied to compensate for high-frequency loss. Then, the signal is framed, with each frame having a length of 1024 sampling points, a frame shift of 512 sampling points, and a Hamming window is applied to reduce spectral leakage. A short-time Fourier transform is performed on each frame to calculate the power spectrum, and then the frequency axis is mapped to the Mel scale through a set of Mel filter banks (typically 80 filters). The logarithm is then taken to obtain the Mel spectrogram. The size of this spectrogram is the number of time frames multiplied by the number of Mel frequency bands, and it can be regarded as a two-dimensional image. This spectrogram is input into a convolutional-Transformer hybrid model. The front end of this model consists of several convolutional layers for extracting local time-frequency patterns; the back end consists of a Transformer encoder for capturing long-term temporal dependencies. The model's output consists of two parts: first, a classification probability distribution of audio event types, such as categories like "gear friction," "bearing squeal," and "normal operation"; and second, a global feature vector output from the Transformer. This global feature vector is reduced to 256 dimensions through a fully connected layer to obtain the audio semantic anchor vector. Simultaneously, the classification probability distribution can serve as one of the inputs for subsequent confidence calculations. Through this processing, each audio segment is transformed into a point in the semantic space, reflecting the event type and temporal-frequency structure features of the audio.
[0077] S2.5: The contrastive learning loss function is used to train the mapping network so that the semantic anchor vectors of all modalities are located in the same metric space, thus obtaining a cross-modal semantic index library.
[0078] In this embodiment, after generating the semantic anchor vectors for the four modalities, they need to be mapped to the same metric space to achieve cross-modal retrieval. Specifically, a contrastive learning training framework is constructed. First, a batch of multimodal data pairs with known matching relationships are collected from maintenance documents as positive samples. For example, a complete maintenance case may include: a text paragraph describing "bearing wear", fault images showing bearing wear, an audio clip recording bearing noise, and anchor vectors of bearing nodes in a 3D model. These four modal data point to the same maintenance entity, constituting a quadruple positive sample. Modal vectors from different entities are randomly combined to form negative samples. The training objective is to learn a mapping network (for each modality, it can be an independent linear projection layer or a small MLP) such that the vectors of different modalities in the positive samples are as close as possible in the projected space, while the vectors in the negative samples are as far apart as possible. The loss function used is the InfoNCE contrastive loss, mathematically expressed as: L = -log( exp(sim(u,v) / tau) / sum(exp(sim(u, v_neg) / tau) ), where sim represents cosine similarity, tau is the temperature parameter, u is the projection vector of the current anchor sample, v represents the projection vector of another modality sample that forms a positive sample with u, and v_neg represents the vector that forms a negative sample with u, typically from any modality projection vector of a different maintenance entity than u. After training, the original semantic anchor vectors of all modalities are projected through their respective mapping networks to obtain a final vector with a unified dimension (e.g., 256 dimensions). These vectors, along with their source metadata (e.g., text source, 3D node ID, image path, audio timestamp), are stored in a vector database to form a cross-modal semantic index. This index supports query vector input from any modality and returns results from any modality with similar semantics.
[0079] According to another aspect of this application, constructing a maintenance status diagram and dynamically switching retrieval strategies includes:
[0080] S3.1: Define the state nodes of the maintenance process, establish directed edges with trigger conditions, and form the initial maintenance state graph.
[0081] According to one aspect of this application, the maintenance process is first formally modeled. Specifically, a set of standard state nodes is defined, including at least five: fault detection, component location, disassembly operation, maintenance execution, and verification test. Each state node represents a distinguishable stage in the maintenance process. Then, directed edges are established between the nodes to represent legal transition paths between states. For example, there is a directed edge from "fault detection" to "component location," triggered by the condition "a clear fault phenomenon has been detected"; the trigger condition from "component location" to "disassembly operation" is "the faulty component has been uniquely identified"; the trigger condition from "disassembly operation" to "maintenance execution" is "disassembly completed"; the trigger condition from "maintenance execution" to "verification test" is "maintenance action has been performed"; and "verification test" can return to "fault detection" (if verification fails) or terminate. Furthermore, some skip edges can be defined, such as direct edges from "fault detection" to "maintenance execution" when the fault type is known and location is not required. The trigger condition for each edge is stored in the form of a logical expression, such as "the user confirms the fault type is a known common fault AND a standard maintenance solution exists." This forms an initial maintenance state diagram, which is stored in the system as a directed graph data structure for subsequent state inference and policy binding.
[0082] S3.2: Collect user operation sequences and timestamps from historical maintenance logs, train a long short-term memory network as a state inferrer, and output the current maintenance status in real time.
[0083] In a preferred embodiment, to achieve runtime state inference, a time-series model needs to be trained using historical maintenance logs. Specifically, a large amount of historical maintenance session data is collected. Each record includes: user ID, session ID, timestamp, user operation type (e.g., "enter query", "click search results", "view 3D model", "confirm step completion", "mark fault"), operation object (e.g., query keyword, document ID, 3D node ID), and manually labeled actual maintenance state at that moment (as a label). Each operation is represented as a feature vector, for example, using one-hot encoding of the operation type, combined with the TF-IDF vector of the keyword. The operation sequence in the same session is arranged in chronological order as the input sequence. A Long Short-Term Memory (LSTM) network is trained, with the operation feature sequence as input and the state probability distribution at each time step as output. The cross-entropy loss function is used for optimization. After training, the LSTM model can be deployed as a state inferrer. During runtime, whenever a user performs a new operation (e.g., submits a query), the system inputs the feature vector of that operation and its preceding operations (e.g., the last 5 steps) into the LSTM, and the model outputs the most likely maintenance state at the current moment and its confidence level. For example, if the output probability is "fault detection: 0.9, component location: 0.1", then the current state is determined to be "fault detection". This inference result will be used to drive subsequent retrieval strategy switching.
[0084] S3.3: Based on the output of the state inferrer, read the corresponding priority retrieval mode from the preset state-policy mapping table.
[0085] In this embodiment, once the current maintenance status is obtained, the system needs to dynamically adjust the allocation of retrieval resources based on this status. Specifically, a status-policy mapping table is pre-constructed, where each row maps a status node to a priority retrieval mode list. The mapping rules are as follows:
[0086] When the status is "Fault Detection", the priority search modalities are, in order: audio, then image. This is because in the early stages of a fault, unusual noises and visual anomalies are the most direct symptom clues.
[0087] When the status is "Component Location", the priority retrieval modalities are as follows: 3D modality, then knowledge graph. The 3D model can help quickly locate the faulty component, while the knowledge graph provides the relationships between components.
[0088] When the status is "Disassembly Operation" or "Maintenance Execution," the priority retrieval modal is as follows: text modal, then 3D highlighting. Text steps provide detailed operation instructions, while 3D highlighting indicates the specific operation location.
[0089] When the status is "Verification Test", the priority retrieval modality is: all modalities are fused, that is, text, 3D, image, and audio are retrieved in parallel, and the results are comprehensively compared.
[0090] During runtime, after the state inferr outputs the current state, the system reads the corresponding priority list. The retrieval module first attempts to search in the constructed cross-modal semantic index using the highest priority modality; if the number or confidence level of the search results is lower than a preset threshold, subsequent modalities are used in descending order of priority. For example, in the "fault detection" state, if the confidence level of all abnormal noise samples returned by the audio retrieval is lower than 0.6, the system automatically switches to image retrieval. Through this mechanism, the retrieval strategy is adaptively adjusted according to the maintenance process.
[0091] According to another aspect of this application, establishing a cognitive adaptation retrieval layer includes:
[0092] S4.1: Collect users' historical query length, document dwell time, 3D model interaction frequency, and error feedback number, and classify users and label preferences through clustering.
[0093] In a preferred embodiment, to achieve personalized retrieval, it is first necessary to construct user profiles. Specifically, behavioral data generated by each user during historical usage is collected, including but not limited to: (a) average query length (in words), reflecting the user's expressive ability; (b) average document dwell time (in seconds), reflecting the user's reading speed; (c) 3D model interaction frequency (e.g., the number of rotation and zoom operations on the 3D view per session), reflecting the user's dependence on spatial information; and (d) the number of error feedbacks (e.g., clicking the "irrelevant" button or submitting the "result error" marker), reflecting the user's professional level and patience. After collecting this data, a four-dimensional feature vector is generated for each user. The K-means clustering algorithm (preset K=3) is used to divide all users into three clusters. By interpreting the behavioral characteristics of users in each cluster, the cluster is labeled as "novice," "intermediate," or "expert." For example, the expert cluster is characterized by short queries, short dwell time, medium interaction frequency, and few error feedbacks; the novice cluster is characterized by long queries, long dwell time, high interaction frequency, and many error feedbacks. In addition, the click ratio of image-based results to text-based results is calculated. If the click ratio of images exceeds 70%, it is marked as "visual preference"; otherwise, it is marked as "text preference". This user profile is stored in the user configuration file and loaded at the start of each session.
[0094] S4.2: Monitor user dwell time and secondary query behavior in real time during maintenance sessions. When insufficient information is detected, call the granularity improvement function to expand the output content.
[0095] According to a further improvement of this embodiment, in addition to static user profiles, the system also senses user behavior in the current session in real time to adjust the information granularity. Specifically, during the session, the system continuously monitors the following two types of behavior: The first type is the time a user spends on a search result. This is recorded by front-end JavaScript tracking, showing the time between the display of the result and the next interaction (such as clicking, scrolling, or querying). If the dwell time exceeds a preset threshold (e.g., 10 seconds), it is determined that the user is carefully reading or feeling confused, and is considered "insufficient information". The second type is secondary query behavior: If a user initiates a new query within a short period of time (e.g., within 30 seconds) after viewing the search results, and the keyword similarity between the old and new queries (calculated using Jaccard similarity or word vector cosine similarity) is higher than 0.7, it is determined that the original search result failed to meet the user's needs, and is also considered "insufficient information". Once insufficient information is determined, the system calls the granularity improvement function. The input of this function is the current output granularity level (e.g., 0 for concise, 1 for detailed), and the output is the new granularity level (at least one level higher). After the granularity is improved, the system re-searches or reorganizes the output content, expanding the original simplified results that may only contain part names and highlighted positions into detailed results that include graphic and textual step sequences, audio waveform diagrams, and 3D disassembly animations.
[0096] S4.3: Assemble multimodal output cards based on user category and current granularity status.
[0097] In one optional implementation, the final output is in the form of a multimodal card, the specific content of which is determined by the user category and the current granularity state. Specifically, three card templates are defined:
[0098] For "new" users, regardless of the current granularity, a complete card is output by default. This card includes: a text step-by-step area on the left or top (with step-by-step instructions, each accompanied by a screenshot or icon); a 3D model view in the central area, where the parts to be operated are highlighted in red and can be dragged and rotated with the mouse; and below, in order, a thumbnail of the fault image (click to enlarge), an audio waveform player (with play / pause buttons), and a simplified diagram of the knowledge graph reasoning path (e.g., "Bearing wear → causing abnormal noise → replacement recommended").
[0099] For "Expert" users, the default output is a minimal card. This card contains only: highlighted nodes in the 3D model (directly located), fault codes (e.g., "ERR-107"), and a minimal set of actions in one line (e.g., "Replace bearing SKF-6204"). If the current granularity is boosted due to insufficient information, the Expert user's card can be expanded into a text list containing key steps, but no images or audio will be displayed.
[0100] For "intermediate" users, a standard card is output, positioned between beginner and expert levels: it includes a text summary of the steps (not the full steps), 3D highlighting, and a clickable "More Details" button, which dynamically loads images and audio content upon user click. If the current granularity is elevated, the details are automatically expanded. Furthermore, if the user profile indicates a "visual preference," images and 3D views are prioritized even in the standard card, with text steps relegated to a secondary position. This assembly strategy achieves personalized and context-adaptive presentation of search results.
[0101] According to another aspect of this application, multimodal confidence conflict resolution includes:
[0102] S5.1: Calculate the initial confidence scores for the text modality, 3D modality, image modality, and audio modality retrieval results respectively to obtain a four-channel confidence vector.
[0103] In one aspect of this application, for retrieval results from different modalities, an initial confidence score needs to be calculated for each. Specifically: For text modal retrieval results, the BM25 algorithm is used to calculate the keyword matching score between the query and the text fragment, and the cosine similarity score between the query vector and the text anchor vector is also calculated. The two scores are weighted and summed with weights of 0.4 and 0.6, respectively, to obtain the text confidence score C_text. For 3D modal retrieval results, the string matching degree between the part names mentioned in the query and the 3D node names (based on edit distance normalization) and the cosine similarity between the query vector and the 3D anchor vector are calculated. These are also weighted and summed to obtain C_3d. For image modal retrieval results, a visual language model (such as CLIP) is used to calculate the matching probability between the query text and the image. If the query itself is an image (image search), the similarity between the anchor vectors of the two images is calculated. The output probability or similarity is used as C_img. For audio modality retrieval results, the probability of the corresponding category (such as "gear friction") in the softmax value output by the audio event classifier is used as C_audio; if the query is the text description "abnormal noise", then the similarity between the text anchor and the audio anchor is calculated.
[0104] The four confidence levels are combined into a four-dimensional vector [C_text, C_3d, C_img, C_audio], with each component taking a value between 0 and 1.
[0105] S5.2: Input the retrieval results of each modality into the pre-trained Siamese Transformer network, and output the conflict probability and conflict type.
[0106] In a further improvement to this embodiment, a Siamese Transformer network is introduced to detect inconsistencies between multimodal results. Specifically, the input to this network consists of semantic anchor vectors of retrieval results from four modalities and their associated metadata (such as temporal information in text, structural constraints in 3D models, spatial locations in images, and timestamps in audio). This information is concatenated into a multimodal feature sequence and input into the Siamese Transformer network. This network employs a pre-trained cross-modal encoder (e.g., an architecture based on ViLBERT or UNITER) and is fine-tuned on maintenance domain data. The network output consists of two branches: the first branch outputs a scalar p_conf, representing the conflict probability, ranging from 0 to 1; the second branch outputs a three-class softmax probability distribution, corresponding to three conflict types: semantic conflict (conceptual contradiction, such as detachable versus non-detachable), spatiotemporal conflict (inconsistent time or spatial location, such as audio indicating the right side but the image showing the left side), and logical conflict (causal or sequential contradiction, such as disassembling A first versus disassembling B first). The training data comes from manually labeled conflict cases. At runtime, inputting the search results into the network will yield p_conf and the dominant conflict type.
[0107] S5.3: Based on the conflict type, call the external maintenance rule base, and use a weighted fusion formula to correct the four-channel confidence vector to obtain the final confidence score.
[0108] In a preferred embodiment, the system intervenes by invoking an external maintenance rule base based on the detected conflict type. This rule base stores conflict resolution priorities preset by domain experts. For example:
[0109] Regarding semantic conflicts, the rule base stipulates: "The structural attributes of the 3D model shall prevail, unless the text source is the official maintenance manual and the version is updated."
[0110] Regarding spatiotemporal conflicts, the rule is: "The one with higher confidence among the real-time collected audio and images shall be used, and the conflict shall be marked."
[0111] For logical conflicts, the rule is: "Based on the causal relationship order in the knowledge graph, the preceding node in the graph takes priority."
[0112] According to the rules, the weights of each modality—w_text, w_3d, w_img, and w_audio—are adjusted. The initial weights are all 0.25. If the rule indicates priority for the 3D modality, w_3d is increased to 0.4, and the others are decreased to 0.2. Then, the final confidence score C_final is calculated using the following weighted fusion formula: C_final = (w_text * C_text + w_3d * C_3d + w_img * C_img + w_audio * C_audio) / (1 + alpha * p_conf); where alpha is the conflict penalty coefficient, which is 0.5 in this embodiment. When p_conf is high, the denominator increases, and C_final is lowered, reflecting the negative impact of conflicts on the reliability of the results.
[0113] S5.4: When the conflict probability exceeds the preset threshold, generate a conflict warning text and encapsulate the final confidence score and the conflict warning text together into the fusion result.
[0114] According to one aspect of this application, after calculating the final confidence score, it is also necessary to determine whether a warning needs to be issued to the user. Specifically, a conflict probability threshold T_conf is set, for example, 0.7. When p_conf ≥ T_conf, the system automatically generates a conflict warning text. This text is automatically concatenated based on the conflict type and the modalities involved. For example:
[0115] If there is a semantic conflict, the warning text will be: "There is a contradiction between the text description and the 3D model structure (the bearing is not removable). It is recommended to refer to the actual product and first refer to the latest version of the maintenance manual."
[0116] If it is a time-space conflict, the warning text will be: "The audio indicates an abnormal noise on the left, but the image shows that the right-side component is damaged. Please verify the direction of the sound source on-site."
[0117] If there is a logical conflict, the warning text will be: "The reasoning path conflicts with the order of the text steps (disassemble A first or disassemble B first). Please check the current assembly status."
[0118] If p_conf is below the threshold, no warning text is generated. Finally, the C_final value, optional conflict warning text, and dominant conflict type (if any) are encapsulated together into a fusion result object. This fusion result is then passed to the subsequent output generation module. Through this conflict resolution process, the system can honestly reveal inconsistencies in multimodal information to the user, avoiding blind trust in the fusion result.
[0119] In a preferred embodiment, it further includes:
[0120] S5.5: Dynamically adjust the weights of each modality based on historical feedback data.
[0121] In one optional implementation, to achieve adaptive optimization of modality weights, implicit user feedback on search results is recorded, and the weights of each modality are dynamically adjusted accordingly. Specifically, each time a user interacts with the search results, if the user performs a positive action (such as clicking the "helpful" button, confirming a step, staying for a long time and successfully completing a repair operation), the result is recorded as "correct"; if the user performs a negative action (such as clicking "useless," submitting error feedback, or modifying the query twice consecutively on the same issue), it is recorded as "incorrect." The average confidence score of each modality in correct cases and the average confidence score in incorrect cases are periodically calculated (e.g., every 100 interactions or daily). Weights are adjusted using an incremental update method: for modality i, if its average confidence score in correct cases is higher than its average confidence score in incorrect cases, the weight of that modality is increased; otherwise, it is decreased. The update formula is: w_i_new = w_i_old + beta * (acc_i_correct - acc_i_error), where beta is the learning rate (e.g., 0.05), and acc_i_correct and acc_i_error are the average confidence scores of the modality in correct and incorrect cases, respectively. After the update, all weights are normalized so that their sum is 1. Through this mechanism, the system can gradually learn the confidence distribution of each modality under different user groups or different maintenance scenarios, further improving the accuracy of the fusion results.
[0122] This invention has the following advantages: First, by constructing a unified semantic anchor space, it enables arbitrary semantic cross-referencing among four modalities: text, 3D models, fault images, and abnormal audio. This breaks the limitation of traditional cross-modal retrieval, which only supports fixed mappings such as "text to image," allowing maintenance personnel to quickly locate fault information in the most natural way (e.g., using audio clips to retrieve the corresponding 3D component location), thus improving the recall rate and retrieval flexibility of multimodal knowledge. Second, based on maintenance status diagrams and real-time status inference, it can dynamically identify the current maintenance stage (e.g., fault detection, component location, disassembly execution) and automatically switch to the most suitable retrieval modality combination accordingly. This avoids the problem of outputting lengthy text steps in the initial fault screening stage or missing image clues during the operation stage, ensuring that the retrieval strategy is highly synchronized with the maintenance process. Third, by introducing a cognitive adaptation retrieval layer, it dynamically adjusts the output granularity based on user profiles (expert / novice, visual / text preferences) and real-time behavior (dwell time, secondary queries). Novices can receive detailed guidance with illustrations, while experts only receive highlighted nodes and the minimum action set, effectively reducing cognitive load and improving maintenance efficiency for users of different skill levels. Fourth, the multimodal confidence conflict resolver can proactively detect and quantify contradictions between text, 3D models, images, and audio (e.g., a structure that cannot be disassembled but the text requires disassembly). Through conflict penalty fusion and warning text output, it avoids misleading information interfering with maintenance decisions, enhancing the reliability and credibility of the retrieval results. In summary, this invention achieves systematic innovation in four dimensions: cross-modal semantic alignment, dynamic context awareness, personalized cognitive adaptation, and multimodal consistency assurance, providing efficient, accurate, and reliable intelligent knowledge retrieval capabilities for the maintenance and support of complex equipment.
[0123] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A dynamic adaptive cross-modal reasoning retrieval method for the maintenance field, characterized in that, include: Acquire multimodal raw data, which includes at least maintenance text, equipment 3D model, fault images, and abnormal audio. Based on multimodal raw data, a unified semantic anchor space is constructed, which maps data from different modalities to the same vector space, generating a cross-modal semantic index library; Based on the dynamic context of the maintenance process, a maintenance status graph is constructed, and the current maintenance status is inferred in real time. The retrieval strategy is switched dynamically with the cross-modal semantic index based on the current maintenance status. Based on user profiles and real-time behavior, a cognitive adaptation retrieval layer is established to dynamically adjust the output granularity and presentation format of retrieval results. Multimodal confidence conflict analysis is performed on search results from different modalities to detect and fuse semantic, spatiotemporal, or logical conflicts between modalities, generating fused results with confidence scores; Based on the fusion results and output granularity, a personalized multimodal retrieval response is generated.
2. The method according to claim 1, characterized in that, Constructing a unified semantic anchor space includes: Triples are extracted from the maintenance text, and semantic anchor vectors are generated through semantic role labeling and encoder. The hierarchical structure tree is parsed from the equipment 3D model, and the component nodes, their geometric features, and descriptive text are extracted to generate 3D semantic anchor vectors. Segment salient regions from faulty images, fuse scene descriptions output by visual language models, and generate image semantic anchor vectors. Mel spectrograms are extracted from abnormal audio, and audio event types are identified through a convolution-transformer hybrid model to generate audio semantic anchor vectors. A contrastive learning loss function is used to train a mapping network, so that the semantic anchor vectors of all modalities are located in the same metric space, resulting in a cross-modal semantic index library.
3. The method according to claim 1, characterized in that, Constructing a maintenance status graph and dynamically switching retrieval strategies includes: Define the state nodes of the maintenance process. The state nodes should include at least fault detection, component location, disassembly operation, maintenance execution and verification test, and establish directed edges with trigger conditions to form an initial maintenance state graph. Collect user operation sequences and timestamps from historical maintenance logs, train a long short-term memory network as a state inferr, and output the current maintenance status in real time; Based on the output of the state inferrer, the corresponding priority retrieval modalities are read from the preset state-policy mapping table. In the fault detection stage, audio and image modalities are retrieved first; in the component localization stage, 3D modalities and knowledge graphs are retrieved first; in the disassembly and execution stage, text modalities and 3D highlights are retrieved first; and in the verification and testing stage, all modalities are integrated.
4. The method according to claim 1, characterized in that, Establishing a cognitive adaptation retrieval layer includes: Collect users' historical query length, document dwell time, 3D model interaction frequency, and error feedback number, and classify users into experts, intermediate, or beginners through clustering, and label visual preferences or text preferences; During the maintenance session, the system monitors the user's dwell time and secondary query behavior in real time. When insufficient information is detected, the granularity enhancement function is called to expand the output content from brief information to a step sequence with pictures and text and audio prompts. Based on the user category and current granularity, assemble multimodal output cards. For novice users, generate cards containing text steps, 3D highlighted screenshots, fault audio waveforms, and graph reasoning paths. For expert users, only 3D node highlights, fault codes, and minimal action sets are output.
5. The method according to claim 1, characterized in that, Multimodal belief conflict resolution includes: The initial confidence scores of the retrieval results for text modality, 3D modality, image modality, and audio modality are calculated respectively to obtain a four-channel confidence vector; The retrieval results of each modality are input into a pre-trained Siamese Transformer network, which outputs the conflict probability and conflict type, including semantic conflict, spatiotemporal conflict and logical conflict. Based on the conflict type, an external maintenance rule base is invoked, and a weighted fusion formula is used to correct the four-channel confidence vector to obtain the final confidence score. When the probability of conflict exceeds a preset threshold, a conflict warning text is generated, and the final confidence score and the conflict warning text are encapsulated together in the fusion result.
6. The method according to claim 5, characterized in that, The method also includes: The weights of each modality are dynamically adjusted based on historical feedback data: when the search results of a certain modality are correct multiple times consecutively after user confirmation, the weight of that modality is increased; when the results of that modality are rejected by the user or conflict occurs, its weight is decreased.
7. The method according to claim 1, characterized in that, After generating a personalized multimodal search response, it also includes: Collect implicit user feedback on search responses, including likes, dislikes, secondary queries, or confirmation of successful steps. By periodically fine-tuning the mapping network in the unified semantic anchor space, the state inferr in the maintenance state diagram, and the modal dynamic weights in the multimodal confidence conflict resolution using implicit feedback, the system can achieve continuous evolution.
8. The method according to claim 1, characterized in that, Maintenance documents include equipment maintenance manuals, PDF technical documents, Word logs, and scanned copies; equipment 3D models include CAD drawings, STEP files, and Web3D interactive models; fault images include photos of the fault scene, instrument panel images, and close-ups of circuit boards; abnormal audio includes engine noises and gear grinding sounds.