A digital resource retrieval system and method based on multi-modal and AI agent

CN122654337BActive Publication Date: 2026-09-29LANGLANG CULTURE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611141288.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-29
Estimated Expiration
2046-07-30

AI Technical Summary

Technical Problem

[0005]为解决上述技术问题,本发明提供一种基于多模态与AI智能体的数字资源检索系统及方法,用于解决多模态数据因特征表达方式异构而难以在统一语义空间中进行表征与比对,导致跨模态检索准确率低的问题

Benefits of technology

本发明包括查询解析模块,用于获取用户输入的多模态查询数据,对多模态查询数据进行预处理、特征提取及跨模态语义对齐,生成多源查询向量集;意图增强模块,用于调用用户分类智能体识别用户类别,根据用户类别对多源查询向量集进行语义增强,生成增强查询向量;语义扩展模块,用于将增强查询向量与预设知识图谱进行匹配,确定目标节点,基于目标节点沿预设知识图谱进行语义扩散,生成扩展查询向量;智能体协同模块,用于基于智能体协同机制,对扩展查询向量执行并行检索得到候选检索结果集,对候选检索结果集执行融合和排序得到初步排序结果,对初步排序结果执行置信度修正得到修正排序结果;重排序输出模块,用于根据预设场景适配因子对修正排序结果进行多维重排序,结合用户类别输出数字资源检索结果,从而可以实现了面向多模态教育场景的用户感知、语义扩展与多智能体协同检索,提升了数字资源检索的准确率与个性化适配程度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654337B_ABST
    Figure CN122654337B_ABST
Patent Text Reader

Abstract

The application provides a kind of digital resource retrieval system and method based on multi-modal and AI intelligent agent, it is related to electric digital data processing technical field, its method includes obtaining multi-modal query data and pre-processing, feature extraction and cross-modal semantic alignment, generates query vector set;Identify user category and accordingly the query vector set is semantically enhanced, and generates enhanced query vector;The enhanced query vector is matched with knowledge graph to determine the target node, and the extended query vector is generated along the semantic diffusion of knowledge graph;Based on the agent collaborative mechanism, the extended query vector is sequentially executed, and the parallel retrieval, fusion and sorting and confidence correction are carried out, and the modified sorting result is obtained;According to the scene adaptation factor, the modified sorting result is reordered, and the digital resource retrieval result is output in combination with the user category, so that user perception, semantic expansion and multi-agent collaborative retrieval for multi-modal education scene can be realized, and the accuracy and personalized adaptation degree of digital resource retrieval are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a digital resource retrieval system and method based on multimodal and AI intelligent agents. Background Technology

[0002] With the rapid advancement of educational informatization, the preschool education field has accumulated a massive amount of digital resources, including text lesson plans, teaching pictures, classroom videos, images of children's artwork, and teaching audio, among other modalities. How to efficiently and accurately retrieve the digital resources needed by users from this heterogeneous and massive amount of multimodal data has become a key technical problem that urgently needs to be solved in the field of educational digitalization.

[0003] Existing digital resource retrieval systems primarily employ keyword matching or single-modal feature-based retrieval methods. For example, traditional solutions typically only build inverted indexes for text-based resources, performing precise or fuzzy matching based on user-input keywords. For non-textual modal data such as images and videos, they rely on manually labeled tags for classification and retrieval, which is costly and lacks comprehensive tag coverage. In recent years, some solutions have attempted to introduce multimodal fusion technology, constructing a retrieval index by simply concatenating or weighting the features of different modal data. However, due to the heterogeneous feature representations of text, images, audio, and video modal data, simple concatenation cannot effectively capture the deep semantic relationships between cross-modal data, resulting in a "semantic gap" between query intent and candidate resources, leading to insufficient relevance and accuracy of retrieval results.

[0004] Therefore, it is necessary to provide a digital resource retrieval system and method based on multimodal and AI intelligent agents to solve the above-mentioned technical problems. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a digital resource retrieval system and method based on multimodal and AI agents, which solves the problem that multimodal data is difficult to represent and compare in a unified semantic space due to heterogeneous feature expression methods, resulting in low cross-modal retrieval accuracy.

[0006] This invention provides a digital resource retrieval system based on multimodal computing and AI agents, the system comprising: The query parsing module is used to obtain multimodal query data input by the user, preprocess the multimodal query data, extract features and perform cross-modal semantic alignment to generate a multi-source query vector set; The intent enhancement module is used to invoke a user classification agent to identify user categories, perform semantic enhancement on the multi-source query vector set based on the user categories, and generate enhanced query vectors. The semantic extension module is used to match the enhanced query vector with a preset knowledge graph, determine the target node, and perform semantic diffusion along the preset knowledge graph based on the target node to generate an extended query vector; The agent collaboration module is used to perform parallel retrieval on the extended query vector based on the agent collaboration mechanism to obtain a candidate retrieval result set, perform fusion and sorting on the candidate retrieval result set to obtain a preliminary sorting result, and perform confidence correction on the preliminary sorting result to obtain a corrected sorting result. The reordering output module is used to perform multidimensional reordering of the corrected sorting results according to a preset scenario adaptation factor, and output digital resource retrieval results in combination with the user category.

[0007] Preferably, the query parsing module is used to acquire multimodal query data input by the user, preprocess the multimodal query data, extract features, and perform cross-modal semantic alignment to generate a multi-source query vector set, specifically including: The text processing unit is used to perform word segmentation and entity recognition on the text modal data in the multimodal query data, and to extract text modal features using a language model; The image processing unit is used to denoise and normalize the image modal data in the multimodal query data, and to extract image modal features using a visual model. An audio processing unit is used to perform noise reduction and speech activity detection on the audio modal data in the multimodal query data, extract acoustic features, and extract semantic features through speech recognition, wherein the acoustic features and the semantic features constitute audio modal features; The video processing unit is used to extract keyframes and segment scenes from the video modal data in the multimodal query data, extract video visual features and video audio features, and generate video modal features. The cross-modal alignment unit is used to construct positive and negative sample pairs between modalities using a contrastive learning framework, minimize the contrastive loss between modalities, and map the text modal features, image modal features, audio modal features, and video modal features to a unified semantic space to generate the multi-source query vector set.

[0008] Preferably, the intent enhancement module is used to invoke a user classification agent to identify user categories, and perform semantic enhancement on the multi-source query vector set based on the user categories to generate enhanced query vectors, specifically including: The user classification unit is used to invoke the user classification agent to obtain the user feature information of the user and identify the user category based on the user feature information; wherein, the user category includes teacher category, parent category and child category; The intent completion unit is used to infer missing modality information based on existing modality information in the multimodal query data, and to perform intent semantic completion on the multi-source query vector set based on the missing modality information to generate a completed query vector set. The attention fusion unit is used to calculate the completion query vector set through an attention mechanism, generate confidence weights, and perform weighted fusion of the completion query vector set based on the confidence weights to generate a fused query vector. The semantic enhancement unit is used to select a corresponding differentiated semantic enhancement strategy according to the user category, and enhance the fused query vector based on the differentiated semantic enhancement strategy to generate the enhanced query vector.

[0009] Preferably, the semantic enhancement unit is used to select a corresponding differentiated semantic enhancement strategy according to the user category, and enhance the fused query vector based on the differentiated semantic enhancement strategy to generate the enhanced query vector, specifically including: A terminology standardization unit is used to perform normalization processing on the fusion query vector in response to the user category being the teacher category, mapping the semantic vector representation of non-standard educational terms in the fusion query vector to the semantic space of standard educational terms, and generating the enhanced query vector. A colloquial mapping unit is used to perform colloquial deconstruction and normalization processing on the fusion query vector in response to the user category being the parent category, mapping the semantic vector representation of colloquial expressions in the fusion query vector to the semantic space of the standard educational terms, and generating the enhanced query vector. The preschool semantic adaptation unit is used to perform the normalization process and the semantic simplification process based on the preschool cognitive characteristics on the fused query vector in response to the user category being the preschool category, and generate the enhanced query vector.

[0010] Preferably, the semantic expansion module is used to match the enhanced query vector with a preset knowledge graph to determine the target node, and perform semantic diffusion along the preset knowledge graph based on the target node to generate an expanded query vector, specifically including: The graph storage unit is used to store the preset knowledge graph, which is constructed from entities in the education field and the relationships between entities. The graph embedding unit is used to perform embedding training on the entities and relationships in the preset knowledge graph using a graph embedding model to obtain the node vectors corresponding to the entities. A node matching unit is used to calculate the similarity between the enhanced query vector and the node vector, and select the node with the highest similarity as the target node; The semantic diffusion unit is used to perform semantic diffusion along the association relationship in the preset knowledge graph, starting from the target node, to obtain the associated nodes associated with the target node; The query expansion unit is used to fuse the node vector of the target node, the node vector of the associated node, and the enhanced query vector to generate the expanded query vector.

[0011] Preferably, the agent collaboration module is used to perform parallel retrieval on the extended query vector based on the agent collaboration mechanism to obtain a candidate retrieval result set, perform fusion and sorting on the candidate retrieval result set to obtain a preliminary sorting result, and perform confidence correction on the preliminary sorting result to obtain a corrected sorting result, specifically including: The routing allocation unit is configured with an AI routing agent, which is used to parse the query features of the extended query vector and determine the retrieval engine to be invoked based on the query features; The retrieval execution unit is used to distribute the extended query vector to the determined retrieval engine, trigger the retrieval engine to perform retrieval operations in parallel, and obtain the candidate retrieval result set; The multi-source fusion unit is equipped with an AI fusion intelligent agent, which is used to perform multi-source fusion and sorting of the retrieval results in the candidate retrieval result set using a ranking learning strategy, and generate the preliminary ranking result. The reflection and correction unit is equipped with an AI reflection and correction agent to evaluate the confidence level of the preliminary ranking result. If the confidence level is lower than a preset confidence threshold, iterative retrieval and correction of the preliminary ranking result is triggered to obtain the corrected ranking result; otherwise, the preliminary ranking result is used as the corrected ranking result.

[0012] Preferably, the reflection and correction unit is configured with an AI reflection and correction agent to evaluate the confidence level of the preliminary ranking result. If the confidence level is lower than a preset confidence threshold, iterative retrieval and correction of the preliminary ranking result is triggered to obtain the corrected ranking result; otherwise, the preliminary ranking result is used as the corrected ranking result. Specifically, this includes: A confidence assessment unit is used to obtain the assessment index of the preliminary ranking result, calculate the confidence level of the preliminary ranking result based on the assessment index, and generate a confidence assessment result; when the confidence level meets the preset confidence threshold, the preliminary ranking result is used as the corrected ranking result. The query correction unit is used to generate a corrected retrieval strategy based on the confidence assessment result when the confidence level is lower than the preset confidence threshold. An iterative retrieval unit is used to send the revised retrieval strategy to the AI ​​routing agent to trigger the agent's collaborative module to re-execute the retrieval, fusion, and sorting according to the revised retrieval strategy, and generate an updated preliminary sorting result; An iterative control unit is used to evaluate the confidence level of the updated preliminary ranking result. When the confidence level of the updated preliminary ranking result is lower than the preset confidence threshold, iterative retrieval is repeatedly triggered until the confidence level of the updated preliminary ranking result meets the preset confidence threshold, and the updated preliminary ranking result is used as the corrected ranking result.

[0013] Preferably, the reordering output module is used to perform multidimensional reordering of the corrected sorting results according to a preset scenario adaptation factor, and output digital resource retrieval results in combination with the user category, specifically including: The scenario factor acquisition unit is used to acquire the preset scenario adaptation factors, which include at least the teaching scenario type, resource usage stage, and interaction mode requirements. The multidimensional re-sorting unit is used to generate multidimensional sorting weights based on the preset scenario adaptation factor and the user category, and to perform weighted sorting on the corrected sorting result based on the multidimensional sorting weights to obtain the re-sorting result. The result output unit is used to perform resource adaptation filtering on the re-ranking results according to the user category, and output the filtered re-ranking results as the digital resource retrieval results.

[0014] A digital resource retrieval method based on multimodal and AI agents, the method comprising: The system acquires multimodal query data input by the user, performs preprocessing, feature extraction, and cross-modal semantic alignment on the multimodal query data, and generates a multi-source query vector set. The user classification agent is invoked to identify user categories, and semantic enhancement is performed on the multi-source query vector set based on the user categories to generate enhanced query vectors; The enhanced query vector is matched with a preset knowledge graph to determine the target node. Based on the target node, semantic diffusion is performed along the preset knowledge graph to generate an extended query vector. Based on the intelligent agent collaboration mechanism, parallel retrieval is performed on the extended query vector to obtain a candidate retrieval result set, fusion and sorting are performed on the candidate retrieval result set to obtain a preliminary sorting result, and confidence correction is performed on the preliminary sorting result to obtain a corrected sorting result; The corrected sorting results are re-sorted in multiple dimensions according to the preset scenario adaptation factors, and the digital resource retrieval results are output in combination with the user category.

[0015] Compared with related technologies, the digital resource retrieval system and method based on multimodal and AI intelligent agents provided by this invention has the following beneficial effects: This invention includes a query parsing module for acquiring multimodal query data input by the user, preprocessing the multimodal query data, extracting features, and performing cross-modal semantic alignment to generate a multi-source query vector set; an intent enhancement module for calling a user classification agent to identify user categories, and semantically enhancing the multi-source query vector set based on user categories to generate enhanced query vectors; a semantic expansion module for matching the enhanced query vectors with a preset knowledge graph to determine target nodes, and semantically diffusing along the preset knowledge graph based on the target nodes to generate expanded query vectors; an agent collaboration module for performing parallel retrieval on the expanded query vectors based on an agent collaboration mechanism to obtain a candidate retrieval result set, performing fusion and sorting on the candidate retrieval result set to obtain a preliminary ranking result, and performing confidence correction on the preliminary ranking result to obtain a corrected ranking result; and a re-ranking output module for performing multi-dimensional re-ranking on the corrected ranking result according to a preset scenario adaptation factor, and outputting digital resource retrieval results in conjunction with user categories. This enables user perception, semantic expansion, and multi-agent collaborative retrieval for multimodal education scenarios, improving the accuracy and personalization of digital resource retrieval.

[0016] This invention preprocesses, extracts features, and performs cross-modal semantic alignment on multimodal query data (text, images, audio, video, etc.) through a query parsing module, generating a multi-source query vector set in a unified semantic space. This effectively eliminates the semantic gap between different modalities and improves the system's compatibility and semantic understanding capabilities for diverse user inputs. The invention also uses an intent enhancement module to invoke a user classification agent to identify user categories such as teachers, parents, and children. Based on these user categories, the multi-source query vector set undergoes intent completion, attention fusion, and differentiated semantic enhancement, achieving precise adaptation to different users' cognitive levels and expression habits, making the search results more aligned with the real needs and comprehension abilities of various users. Finally, the invention uses a semantic extension module to perform node matching between enhanced query vectors and a pre-defined knowledge graph, and performs semantic diffusion along the relationships within the knowledge graph. This fuses target nodes, related nodes, and enhanced query vectors to generate extended query vectors, achieving knowledge-driven expansion of user query intent and overcoming the shortcomings of insufficient semantic coverage in traditional keyword retrieval. This invention configures AI routing agents, AI fusion agents, and AI reflection and correction agents through an intelligent agent collaboration module, driving multiple search engines to perform parallel searches. It performs multi-source fusion and ranking of candidate search result sets, and triggers iterative correction based on confidence assessment, forming a closed-loop mechanism of retrieval, evaluation, and correction. This effectively overcomes the shortcomings of traditional single-search result quality uncontrollability, significantly improving the accuracy and robustness of search results. Furthermore, this invention uses a re-ranking output module to perform multi-dimensional re-ranking and resource-adaptive filtering of the corrected ranking results based on preset scenario adaptation factors and user categories. This achieves multi-dimensional adaptive output based on teaching scenarios, resource usage stages, and interaction mode requirements, further enhancing the accuracy and personalized service capabilities of digital resource retrieval. Attached Figure Description

[0017] Figure 1 This is a system block diagram of a digital resource retrieval system based on multimodal and AI intelligent agents according to the present invention; Figure 2 This is a flowchart of a digital resource retrieval method based on multimodal and AI agents according to the present invention. Detailed Implementation

[0018] The present invention will be further described below with reference to the accompanying drawings and embodiments. Example

[0019] like Figure 1 As shown, a digital resource retrieval system based on multimodal computing and AI agents is disclosed, the system comprising: The query parsing module is used to obtain multimodal query data input by the user, preprocess the multimodal query data, extract features and perform cross-modal semantic alignment to generate a multi-source query vector set; Multimodal query data refers to search input submitted by users in the form of text, images, audio, or video. Preprocessing includes data cleaning operations such as word segmentation, noise reduction, and normalization. Feature extraction refers to the encoding process of extracting semantic features from various modal data. Cross-modal semantic alignment uses a contrastive learning mechanism to map features from different modalities to a unified semantic space, eliminating modal heterogeneity. The multi-source query vector set is a vector set composed of the semantic vectors corresponding to the aligned modal data.

[0020] A digital resource retrieval system is an information processing tool that performs unified semantic understanding and localization of heterogeneous digital content such as images, audio, video, and text. Through cross-modal perception and semantic alignment, the system understands the user intent implied in multimodal query inputs. It leverages knowledge graphs for semantic expansion and a multi-agent collaborative retrieval correction mechanism to achieve accurate retrieval of massive amounts of digital resources and can output retrieval results adapted to different users' cognitive levels and usage scenarios.

[0021] Understandably, after completing cross-modal semantic alignment, the query parsing module does not immediately compress all modal features into a single vector. Instead, it outputs a set of vectors that retain the independent semantic representations of each modality. This design aims to provide fine-grained operational space for the subsequent intent enhancement module. If fusion occurs too early at this stage, modal boundaries will disappear, and the intent completion unit will be unable to distinguish between existing modal information and missing modal information, making it difficult to perform semantic inference for the missing modality. Simultaneously, the attention fusion unit also needs to maintain the independence of the semantic vectors of each modality to perform differentiated weighting based on the confidence of different modalities.

[0022] The intent enhancement module is used to invoke a user classification agent to identify user categories, perform semantic enhancement on the multi-source query vector set based on the user categories, and generate enhanced query vectors. The user classification agent is an intelligent decision-making unit that dynamically identifies a user's cognitive level and identity attributes based on user characteristic information. The enhanced query vector is a fused query vector generated from a multi-source query vector set through intent completion and attention fusion, and is the output vector representation after differentiated semantic enhancement according to different user categories.

[0023] It is understandable that semantic enhancement is not simply result filtering, but rather refers to the differentiated reshaping of query intent in the vector space based on the cognitive characteristics and expression habits of different user categories. This includes operations such as terminology normalization mapping, colloquial deconstruction and translation, and semantic simplification and adaptation, ultimately generating enhanced query vectors that are more in line with the understanding of that type of user.

[0024] The semantic extension module is used to match the enhanced query vector with a preset knowledge graph, determine the target node, and perform semantic diffusion along the preset knowledge graph based on the target node to generate an extended query vector; The pre-defined knowledge graph is a structured semantic network built upon entities in the education field and the relationships between them. Entities encompass teaching elements such as knowledge points, courses, and resource types, while relationships include teaching logic such as membership, predecessors and successors, and relevance. The target node is the entity node in the pre-defined knowledge graph that best matches the semantics of the enhanced query vector, located through vector similarity calculation. The extended query vector is a semantic vector generated by fusing the target node vector, the associated node vector, and the enhanced query vector.

[0025] Understandably, relying solely on keywords or shallow semantics in user queries for retrieval often fails to cover all concepts and terms related to the query intent within the knowledge system, easily leading to insufficient semantic coverage of the retrieval results. The semantic expansion module addresses this by matching enhanced query vectors with nodes in a pre-defined knowledge graph to locate the target node closest to the user's query semantics. Starting from this target node, it then expands semantically along the relationships within the knowledge graph to obtain related nodes with pedagogical logical connections to the target node. Based on this, the node vectors of the target node, the node vectors of related nodes, and the enhanced query vector are fused. This results in an expanded query vector that not only retains the user's original query intent but also incorporates structured related knowledge from the knowledge graph. This proactive expansion of query semantics before retrieval execution effectively improves retrieval recall and semantic coverage breadth.

[0026] The agent collaboration module is used to perform parallel retrieval on the extended query vector based on the agent collaboration mechanism to obtain a candidate retrieval result set, perform fusion and sorting on the candidate retrieval result set to obtain a preliminary sorting result, and perform confidence correction on the preliminary sorting result to obtain a corrected sorting result. The intelligent agent collaboration mechanism refers to a working mode in which multiple AI agents with independent decision-making capabilities work together to complete retrieval tasks according to the principle of division of labor and cooperation. This includes AI routing agents, AI fusion agents, and AI reflection and correction agents. The candidate retrieval result set is a summary collection of results returned by multiple retrieval engines after parallel execution of retrieval on the extended query vector. It includes resource lists and relevance scores independently generated by each engine. The preliminary ranking result is an ordered resource list generated after multi-source fusion and ranking of the candidate retrieval result set, including the comprehensive relevance score and ranking of each candidate resource. The revised ranking result is a quality-optimized ranking result obtained after evaluating the confidence level of the preliminary ranking result and triggering query correction and iterative retrieval when the confidence level is insufficient.

[0027] Understandably, relying solely on a single retrieval operation often limits the quality of search results due to the coverage limitations and ranking biases of a single engine, making it difficult to consistently meet query demands of varying complexity. The agent collaboration module, by introducing a multi-agent collaborative retrieval architecture, forms a hierarchical processing mechanism from coarse screening to fine ranking and then to a quality closed loop, significantly improving the recall completeness and ranking accuracy of search results.

[0028] The reordering output module is used to perform multidimensional reordering of the corrected sorting results according to a preset scenario adaptation factor, and output digital resource retrieval results in combination with the user category.

[0029] Understandably, the preset scenario adaptation factor is a ranking criterion used to adjust the ranking results based on dynamic factors such as the type of teaching scenario, the stage of resource usage, and the needs of the interaction mode in the current search. It ensures that the same resource presents a differentiated order in different usage scenarios, maintaining consistency between the ranking results and the user's current actual needs. Multidimensional re-ranking is not simply replacing the previous ranking results; rather, it adds weighted adjustments based on the relevance of the revised ranking results, incorporating user adaptation and scenario adaptation dimensions, ultimately outputting digital resource search results that are both relevant and applicable.

[0030] The query parsing module is used to acquire multimodal query data input by the user, preprocess the multimodal query data, extract features, and perform cross-modal semantic alignment to generate a multi-source query vector set, specifically including: The text processing unit is used to perform word segmentation and entity recognition on the text modal data in the multimodal query data, and to extract text modal features using a language model; The image processing unit is used to denoise and normalize the image modal data in the multimodal query data, and to extract image modal features using a visual model. An audio processing unit is used to perform noise reduction and speech activity detection on the audio modal data in the multimodal query data, extract acoustic features, and extract semantic features through speech recognition, wherein the acoustic features and the semantic features constitute audio modal features; The video processing unit is used to extract keyframes and segment scenes from the video modal data in the multimodal query data, extract video visual features and video audio features, and generate video modal features. The cross-modal alignment unit is used to construct positive and negative sample pairs between modalities using a contrastive learning framework, minimize the contrastive loss between modalities, and map the text modal features, image modal features, audio modal features, and video modal features to a unified semantic space to generate the multi-source query vector set.

[0031] The language model is a pre-trained language model fine-tuned based on educational corpora, adapted to the text expression habits and professional terminology system in preschool education scenarios. The visual model is a visual Transformer model that extracts global and local region features of images through a self-attention mechanism. Acoustic features are parameters in audio signals reflecting the physical attributes of the speaker, including timbre, speech rate, and emotional dimension, used to aid in understanding non-semantic information in audio. Semantic features are language representations extracted after the speech signal is converted into text, used to express the semantic content in the audio. Video visual features are spatial semantic descriptions extracted from video keyframes, while video audio features are acoustic and semantic information extracted from the video audio track; these are fused after temporal alignment to form video modal features. The contrastive learning framework constructs positive and negative sample pairs between modalities to shorten the feature distance between different modalities under the same semantic meaning and widen the feature distance between modalities under different semantic meanings, aligning the four modal features in a unified semantic space.

[0032] In practical applications, after a user submits query data containing at least two modalities (text, image, audio, or video) through a terminal device, the query parsing module first distributes the data to the corresponding processing unit according to the modality type. The text processing unit performs word segmentation, part-of-speech tagging, and named entity recognition on the input text, extracting keywords such as names of people, places, course names, and teaching objectives. The processed text sequence is then input into a pre-trained language model fine-tuned based on educational corpora. This model outputs a fixed-dimensional text semantic vector as the text modality feature. The image processing unit performs Gaussian filtering for noise reduction and bicubic interpolation for size normalization on the input image, scaling the image to a fixed resolution. The normalized image is then input into a visual Transformer model. This model segments the image into several image patches, performs linear projection on each patch, adds positional encoding, and outputs global and local image features after multi-head self-attention computation, collectively constituting the image modality feature.

[0033] The audio processing unit performs spectral subtraction noise reduction on the input audio to remove background noise interference. It then removes silent segments through speech activity detection, retaining only the effective speech segments. Mel-frequency cepstral coefficients are extracted from the effective speech segments as acoustic features. These segments are then fed into a speech recognition model to transcribe the audio into text, which is then semantically encoded to obtain semantic features. The acoustic and semantic features are concatenated to form the audio modality features. The video processing unit decodes the input video, extracts keyframes based on scene transition detection results, and separates the accompanying audio from the video audio track. Spatial visual features are extracted from the keyframe sequence, and acoustic and semantic features are extracted from the accompanying audio. These two types of features are aligned by timestamps and then fused to generate the video modality features.

[0034] The cross-modal alignment unit receives the four modal features mentioned above and uses a contrastive learning framework for semantic alignment. This framework configures an independent encoder for each modal feature, encoding each modal feature into a fixed-dimensional vector. During the training phase, a batch of sample pairs containing the four modalities is sampled from the training set. Different modal features belonging to the same semantic content are constructed into positive sample pairs, and different modal features belonging to different semantic contents are constructed into negative sample pairs. By minimizing the inter-modal contrastive loss, the distance between positive sample pairs in the semantic space is made closer, while the distance between negative sample pairs is made farther. After the above mapping, the four modal features are in the same semantic space, forming a multi-source query vector set.

[0035] The intent enhancement module is used to invoke a user classification agent to identify user categories, and perform semantic enhancement on the multi-source query vector set based on the user categories to generate enhanced query vectors, specifically including: The user classification unit is used to invoke the user classification agent to obtain the user feature information of the user and identify the user category based on the user feature information; wherein, the user category includes teacher category, parent category and child category; The intent completion unit is used to infer missing modality information based on existing modality information in the multimodal query data, and to perform intent semantic completion on the multi-source query vector set based on the missing modality information to generate a completed query vector set. The attention fusion unit is used to calculate the completion query vector set through an attention mechanism, generate confidence weights, and perform weighted fusion of the completion query vector set based on the confidence weights to generate a fused query vector. The semantic enhancement unit is used to select a corresponding differentiated semantic enhancement strategy according to the user category, and enhance the fused query vector based on the differentiated semantic enhancement strategy to generate the enhanced query vector.

[0036] The user characteristic information includes at least voiceprint features, input method type, and interaction behavior pattern. The teacher category refers to a group with professional educational backgrounds who use structured teaching terminology. The parent category refers to a non-professional group that uses colloquial language. The preschool category refers to a group whose cognitive abilities are still developing and who often express themselves in unstructured ways, such as incomplete speech or doodles. Existing modal information refers to the modal types and corresponding semantic content already included in the query data actually submitted by the user. Missing modal information refers to the semantic content carried by modal data related to the current query intent but not directly provided by the user. The completed query vector set is a set of multimodal vectors generated after semantic inference of missing modalities. The attention mechanism is a weighting strategy for calculating the importance of each modal vector in the completed query vector set. The confidence weight is the proportion coefficient of each modal vector in the fusion process. The fused query vector is a single vector generated by weighted summation of each modal vector based on the confidence weight. Differentiated semantic enhancement strategy is a set of rules that select different processing logics based on user category, used to further enhance the semantic dimension that matches the cognitive characteristics of the current user category on the basis of fused query vector.

[0037] In practical applications, after the intent enhancement module receives the multi-source query vector set, the user classification unit first obtains the input method type, voiceprint label, and interaction behavior features. The input method type is collected and labeled by the user terminal, the voiceprint label is obtained by comparing the voiceprint features with the registered voiceprint database, and the interaction behavior features are generated by the system based on the user's historical operation records. These three types of features are encoded and concatenated into a user feature vector. The cosine similarity of this vector with the center of the teacher, parent, and child categories is calculated respectively, and the category with the highest similarity is output as the user category.

[0038] The intent completion unit receives a multi-source query vector set, extracts the effective dimension count and vector magnitude of each modality vector, and determines whether the data corresponding to each modality is complete. If a user submits only a single modality of voice input, such as a child user submitting only voice input, then the other modality vectors are all zero vectors. The intent completion unit determines that the audio modality is an existing modality, and the other modalities are missing modalities. It retrieves the text semantic vector, image semantic vector, and video semantic vector that are closest to the audio modality semantic vector from the cross-modal association matrix, and fills them into the vector positions corresponding to the missing modalities, forming a completion query vector set containing semantic vectors from all four modalities. During the training phase, the cross-modal completion network uses complete multimodal sample pairs, randomly masking the input of some modalities, and trains the network to reconstruct the semantic features of the masked modalities from the visible modalities.

[0039] The attention fusion unit receives the complete query vector set and maps each modality vector to the same dimensional space via linear projection. For each modality vector, the corresponding query vector, key vector, and value vector are calculated using three sets of projection matrices. Taking the query vector of any modality as a reference, its dot product similarity with the key vectors of all modalities is calculated. After normalization, the attention weight of that modality to each modality is obtained. Then, the value vectors of each modality are weighted and summed to obtain the updated query vector after fusing global information for that modality. After the above calculations are performed on each modality, all updated query vectors are average pooled to output the fused query vector.

[0040] The semantic enhancement unit is used to select a corresponding differentiated semantic enhancement strategy according to the user category, and enhance the fused query vector based on the differentiated semantic enhancement strategy to generate the enhanced query vector, specifically including: A terminology standardization unit is used to perform normalization processing on the fusion query vector in response to the user category being the teacher category, mapping the semantic vector representation of non-standard educational terms in the fusion query vector to the semantic space of standard educational terms, and generating the enhanced query vector. A colloquial mapping unit is used to perform colloquial deconstruction and normalization processing on the fusion query vector in response to the user category being the parent category, mapping the semantic vector representation of colloquial expressions in the fusion query vector to the semantic space of the standard educational terms, and generating the enhanced query vector. The preschool semantic adaptation unit is used to perform the normalization process and the semantic simplification process based on the preschool cognitive characteristics on the fused query vector in response to the user category being the preschool category, and generate the enhanced query vector.

[0041] Non-standard educational terminology includes abbreviations, local expressions, and non-standard teaching language used by teachers in daily lesson preparation. The semantic space of standard educational terminology is a standardized vector space constructed based on the standard terminology in general textbooks. Colloquial expressions refer to the everyday conversational language used by parents when describing their educational needs. Early childhood cognitive characteristics include the range of attention span, the dominance of concrete thinking, and the cognitive threshold of young children.

[0042] In practical applications, after receiving the fused query vector, the semantic enhancement unit selects the corresponding processing branch based on the category label output by the user classification unit. When the user category is teacher, the terminology standardization unit loads a pre-built deep semantic correction network, inputting the fused query vector as a whole into the network. This network consists of multiple fully connected layers stacked with non-linear activation functions. During the training phase, it is supervised learning using query vectors labeled with the correspondence between non-standard and standard educational terms. The network parameters encode the global transformation relationship from the semantic distribution pattern of non-standard educational terms to the semantic space of standard educational terms. During inference, the fused query vector flows through each layer of the network. Each layer performs linear combination and non-linear activation on the fused query vector, re-encoding all semantic information of the vector layer by layer, and outputting an enhanced query vector.

[0043] When the user category is "parent," the colloquial mapping unit loads the colloquial semantic mapping network, using the same mechanism to transfer the overall semantics of the fused query vector from the everyday spoken language space to the standard educational terminology space. When the user category is "child," the child semantic adaptation unit first normalizes the vector through a deep semantic correction network, then inputs the intermediate vector into the child cognitive adaptation mapping network. It performs cognitive adaptation recoding on all dimensions, applying global attenuation to dimensions carrying abstract concepts and global enhancement to basic cognitive dimensions, outputting an enhanced query vector. The entire transformation process does not involve local component localization or replacement; the network achieves context-aware overall semantic transfer through global dimension recoding. The colloquial semantic mapping network and the deep semantic correction network use the same network structure (3 fully connected layers), differing only in their training data: the training samples for the colloquial semantic mapping network are pairs of correspondences between everyday spoken expressions and standard educational terms. The child cognitive adaptation mapping network also uses a 3-layer fully connected layer structure, with training samples consisting of pairs of correspondences between typical child expressions and simplified educational terms.

[0044] Understandably, after obtaining cross-modal global semantics through attention fusion, user category adaptation is then performed, ensuring that enhancement operations are based on complete multimodal intent rather than single-modal information. This approach ensures that differentiated semantic enhancement for teachers, parents, or young children is built upon a comprehensive understanding of the query content, avoiding intent bias caused by single-modal enhancement. The fused single vector, after user category adaptation, generates an enhanced query vector that possesses both cross-modal semantic integrity and aligns with the cognitive characteristics and expression habits of specific user groups, enabling personalized adaptation of search results based on a comprehensive understanding of the query intent.

[0045] The semantic expansion module is used to match the enhanced query vector with a preset knowledge graph to determine the target node, and perform semantic diffusion along the preset knowledge graph based on the target node to generate an expanded query vector, specifically including: The graph storage unit is used to store the preset knowledge graph, which is constructed from entities in the education field and the relationships between entities. The graph embedding unit is used to perform embedding training on the entities and relationships in the preset knowledge graph using a graph embedding model to obtain the node vectors corresponding to the entities. A node matching unit is used to calculate the similarity between the enhanced query vector and the node vector, and select the node with the highest similarity as the target node; The semantic diffusion unit is used to perform semantic diffusion along the association relationship in the preset knowledge graph, starting from the target node, to obtain the associated nodes associated with the target node; The query expansion unit is used to fuse the node vector of the target node, the node vector of the associated node, and the enhanced query vector to generate the expanded query vector.

[0046] Graph embedding models are a technique that maps entities and relations in a knowledge graph to a low-dimensional, dense vector space. During construction, a knowledge graph containing entities and relations is first built, and then trained using algorithms such as TransE. The model optimizes the loss function to ensure that these vectors preserve the semantic relationships between entities.

[0047] In practical applications, after the semantic extension module receives the enhanced query vector, the graph storage unit loads a pre-defined knowledge graph. The entity types in the graph include teaching target entities, knowledge point entities, resource type entities, age group entities, and course entities, and the relationship types include belonging to, applicable to, belonging to, and associated with. The graph embedding unit loads pre-trained graph embedding model parameters, mapping each entity in the graph from its one-hot encoding space to node vectors in a low-dimensional continuous vector space. The dimension of each node vector is consistent with the dimension of the enhanced query vector. The node matching unit receives the enhanced query vector, traverses all node vectors corresponding to all entities in the pre-defined knowledge graph, calculates the cosine similarity between the enhanced query vector and each node vector, and maintains the current highest similarity value and its corresponding node identifier during the traversal. After the traversal, the node with the highest similarity is selected as the target node.

[0048] The semantic diffusion unit starts with the target node, obtains its entity type, determines the corresponding diffusion depth and direction from a pre-defined diffusion strategy table based on the entity type, and performs a multi-hop traversal along the association relationships in the knowledge graph. Each node visited during a hop is added to the diffusion result list. Traversal stops when the diffusion depth is reached, and all nodes in the diffusion result list except the starting point are marked as associated nodes. The query expansion unit obtains the node vector of the target node and the node vectors of each associated node, calculates the mean of all associated node vectors as the associated semantic vector, and weights and concatenates the enhanced query vector, the target node's node vector, and the associated semantic vector to generate and output the expanded query vector.

[0049] The agent collaboration module is used to perform parallel retrieval on the extended query vector based on the agent collaboration mechanism to obtain a candidate retrieval result set, perform fusion and sorting on the candidate retrieval result set to obtain a preliminary sorting result, and perform confidence correction on the preliminary sorting result to obtain a corrected sorting result. Specifically, it includes: The routing allocation unit is configured with an AI routing agent, which is used to parse the query features of the extended query vector and determine the retrieval engine to be invoked based on the query features; The retrieval execution unit is used to distribute the extended query vector to the determined retrieval engine, trigger the retrieval engine to perform retrieval operations in parallel, and obtain the candidate retrieval result set; The multi-source fusion unit is equipped with an AI fusion intelligent agent, which is used to perform multi-source fusion and sorting of the retrieval results in the candidate retrieval result set using a ranking learning strategy, and generate the preliminary ranking result. The reflection and correction unit is equipped with an AI reflection and correction agent to evaluate the confidence level of the preliminary ranking result. If the confidence level is lower than a preset confidence threshold, iterative retrieval and correction of the preliminary ranking result is triggered to obtain the corrected ranking result; otherwise, the preliminary ranking result is used as the corrected ranking result.

[0050] The AI ​​routing agent is a retrieval routing unit based on joint decision-making using rules and thresholds. It has a pre-built query feature-retrieval engine mapping table. This table categorizes queries into semantic, keyword, and graph-related types based on semantic distribution density and keyword matching degree in the query features, and determines the retrieval engine to be invoked based on the classification results. The AI ​​fusion agent is a ranking model trained using a ranking learning strategy. During the training phase, it takes query-resource sample pairs labeled with relevance tags as input, extracting multi-dimensional features such as semantic similarity, knowledge graph association strength, resource quality, and user preferences. It uses the LambdaMART algorithm to iteratively optimize the loss function until convergence. During deployment, model parameters are loaded for inference scoring. The AI ​​reflection and correction agent is a decision-making unit built based on pre-set confidence assessment logic. This confidence assessment logic is jointly defined by the threshold difference between the highest and lowest scores in the initial ranking results and a diversity index threshold.

[0051] In practical applications, the agent collaboration module receives extended query vectors. The routing allocation unit loads the AI ​​routing agent, which has a pre-built query feature-retrieval engine mapping table. It parses the semantic distribution density of the extended query vector, statistically analyzes the distribution of activation values ​​in each dimension of the vector, calculates the proportion of non-zero dimensions in the vector as the semantic distribution density, extracts keyword feature vectors from the vector, and calculates the similarity between the keyword feature vectors and a pre-built keyword feature codebook as the keyword matching degree. The keyword feature codebook is constructed by collecting, deduplicating, and encoding terms from standard lesson plans, curriculum outlines, and teaching guidance documents in the field of preschool education. The semantic distribution density and keyword matching degree are matched against each rule in the query feature-retrieval engine mapping table; upon a match, a list of corresponding retrieval engine identifiers is output.

[0052] The retrieval execution unit expands the query vector and distributes it to each retrieval engine in the retrieval engine identifier list. Each retrieval engine independently performs retrieval operations based on its own index structure and aggregates the returned results into a candidate retrieval result set. The retrieval engines include a first retrieval engine based on vector similarity, a second retrieval engine based on keyword inverted index, and a third retrieval engine based on knowledge graph structure matching. The multi-source fusion unit loads the AI ​​fusion agent, receives the candidate retrieval result set, extracts semantic similarity features, knowledge graph association strength features, resource quality features, and user historical preference features for each candidate retrieval result, concatenates these features into a feature vector, and inputs it into a trained ranking model. The model outputs a comprehensive relevance score after forward propagation. Based on the comprehensive relevance score, the candidate retrieval result set is sorted in descending order to generate a preliminary ranking result.

[0053] The reflection and correction unit is configured with an AI reflection and correction agent to evaluate the confidence level of the preliminary ranking result. If the confidence level is lower than a preset confidence threshold, iterative retrieval and correction of the preliminary ranking result is triggered to obtain the corrected ranking result; otherwise, the preliminary ranking result is used as the corrected ranking result. Specifically, this includes: A confidence assessment unit is used to obtain the assessment index of the preliminary ranking result, calculate the confidence level of the preliminary ranking result based on the assessment index, and generate a confidence assessment result; when the confidence level meets the preset confidence threshold, the preliminary ranking result is used as the corrected ranking result. The query correction unit is used to generate a corrected retrieval strategy based on the confidence assessment result when the confidence level is lower than the preset confidence threshold. An iterative retrieval unit is used to send the revised retrieval strategy to the AI ​​routing agent to trigger the agent's collaborative module to re-execute the retrieval, fusion, and sorting according to the revised retrieval strategy, and generate an updated preliminary sorting result; An iterative control unit is used to evaluate the confidence level of the updated preliminary ranking result. When the confidence level of the updated preliminary ranking result is lower than the preset confidence threshold, iterative retrieval is repeatedly triggered until the confidence level of the updated preliminary ranking result meets the preset confidence threshold, and the updated preliminary ranking result is used as the corrected ranking result.

[0054] The evaluation metric is the standard deviation of all comprehensive relevance scores in the initial ranking results. The pre-set reliability threshold is set before system deployment based on business scenario requirements and dynamically adjusted during system operation; it is typically set to 0.7. The confidence assessment results include the distribution characteristics of the confidence value and the comprehensive relevance score. The revised search strategy is a search adjustment instruction generated when the confidence level is insufficient, containing a reselection scheme for the search engine combination and constraints on the search scope.

[0055] In practical applications, after receiving the preliminary ranking results, the reflection and correction unit first uses the confidence assessment unit to iterate through the comprehensive relevance scores of each search result in the preliminary ranking, calculating the standard deviation of all comprehensive relevance scores as an evaluation indicator. The confidence assessment unit inputs the calculated standard deviation into a preset confidence mapping function. This function uses a sigmoid function to map the original standard deviation to a numerical range consistent with a preset confidence threshold, outputting a standardized confidence value. It also analyzes the distribution characteristics of the comprehensive relevance scores, packaging the distribution characteristics and the confidence value together to generate the confidence assessment result. When the confidence value is greater than or equal to the preset confidence threshold, it indicates that the scores of each candidate resource are significantly different and the ranking quality is reliable. The preliminary ranking result is then directly marked as the corrected ranking result and output. When the confidence score is less than the preset confidence threshold, it indicates that the scores of each candidate resource are too concentrated and the discrimination is insufficient. The query correction unit analyzes the distribution characteristics of the comprehensive relevance score in the confidence evaluation results, identifies the dimensions with weak semantic coverage in the current search, and generates a correction search strategy accordingly. This strategy includes search engine selection instructions and search scope constraints set for the semantic features of low-discrimination resources.

[0056] The iterative retrieval unit encapsulates the revised retrieval strategy into a call request and sends it to the input interface of the AI ​​routing agent. After parsing the revised retrieval strategy content, the AI ​​routing agent re-determines the combination of retrieval engines to be called according to the retrieval engine selection instruction. The retrieval execution unit and the multi-source fusion unit sequentially perform parallel retrieval and fusion sorting to generate updated preliminary sorting results. After reading the updated preliminary sorting results, the iterative control unit calls the confidence evaluation unit again to calculate the confidence value. If the confidence value is still less than the preset confidence threshold, the query correction and iterative retrieval process is triggered repeatedly until the confidence value meets the preset confidence threshold or reaches the preset maximum number of iterations. The current preliminary sorting result is then output as the revised sorting result.

[0057] The re-sorting output module is used to perform multi-dimensional re-sorting of the corrected sorting results according to a preset scenario adaptation factor, and output digital resource retrieval results in conjunction with the user category, specifically including: The scenario factor acquisition unit is used to acquire the preset scenario adaptation factors, which include at least the teaching scenario type, resource usage stage, and interaction mode requirements. The multidimensional re-sorting unit is used to generate multidimensional sorting weights based on the preset scenario adaptation factor and the user category, and to perform weighted sorting on the corrected sorting result based on the multidimensional sorting weights to obtain the re-sorting result. The result output unit is used to perform resource adaptation filtering on the re-ranking results according to the user category, and output the filtered re-ranking results as the digital resource retrieval results.

[0058] The preset scenario adaptation factors are a set of sorting adjustment parameters pre-configured by the system based on the current search context before the search is executed. These parameters include three dimensions: teaching scenario type, resource usage stage, and interaction mode requirements. Teaching scenario type refers to the category of teaching activity to which the current search belongs, including group teaching, area activities, and individual instruction. Resource usage stage refers to the expected timing of using the search results, including pre-class preparation, in-class presentation, and post-class extension. Interaction mode requirements refer to the interaction methods of the user's terminal device, including touchscreen browsing, voice interaction, and point-and-read interaction. The multi-dimensional sorting weight is a weight vector obtained by multiplying the values ​​of each dimension in the preset scenario adaptation factors with the user preference coefficient corresponding to the user category and then normalizing the result. Resource adaptation filtering filters out unsuitable resource types from the re-sorting results based on the user category; for example, plain text resources are filtered out from the re-sorting results output by users in the preschool category.

[0059] In practical applications, after receiving the revised ranking results, the re-ranking output module first reads the teaching scenario type, resource usage stage, and interaction mode requirements from the current search context, encoding the values ​​of these three dimensions into a scenario adaptation factor vector. The multi-dimensional re-ranking unit multiplies the scenario adaptation factor vector with the preference coefficients corresponding to the user category tags item by item to obtain a multi-dimensional ranking weight vector. It then iterates through each candidate resource in the revised ranking results, multiplying the comprehensive relevance score of each candidate resource with the corresponding weight value in the multi-dimensional ranking weight vector; the product is the final ranking score for that candidate resource. Based on the final ranking scores, all candidate resources are re-ranked to generate the re-ranking results. The result output unit filters out unsuitable resource types from the re-ranking results according to the user category. The output for the teacher category retains all resource types, the output for the parent category filters out professional lesson plan resources, and the output for the preschool category filters out plain text resources, retaining only image and audio / video resources. The result output unit then outputs the filtered re-ranking results to the user terminal as the digital resource retrieval results. Example

[0060] like Figure 2 As shown, a digital resource retrieval method based on multimodal and AI agents is described, the method comprising: S1, acquire multimodal query data input by the user, preprocess the multimodal query data, extract features and perform cross-modal semantic alignment to generate a multi-source query vector set; S2, invoke the user classification agent to identify user categories, and perform semantic enhancement on the multi-source query vector set according to the user categories to generate enhanced query vectors; S3, match the enhanced query vector with the preset knowledge graph to determine the target node, and perform semantic diffusion along the preset knowledge graph based on the target node to generate an extended query vector; S4, based on the intelligent agent collaboration mechanism, parallel retrieval is performed on the extended query vector to obtain a candidate retrieval result set, fusion and sorting are performed on the candidate retrieval result set to obtain a preliminary sorting result, and confidence correction is performed on the preliminary sorting result to obtain a corrected sorting result; S5. The corrected sorting results are re-sorted in multiple dimensions according to the preset scenario adaptation factor, and the digital resource retrieval results are output in combination with the user category.

[0061] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0062] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0063] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

Claims

1. A digital resource retrieval system based on multi-modal and AI agent, characterized in that, The system comprises: a query analysis module, configured to acquire multi-modal query data input by a user, perform preprocessing, feature extraction and cross-modal semantic alignment on the multi-modal query data, and generate a multi-source query vector set; an intent enhancement module, configured to invoke a user classification agent to identify a user category, perform semantic enhancement on the multi-source query vector set according to the user category, and generate an enhanced query vector; a semantic expansion module, configured to match the enhanced query vector with a preset knowledge graph, determine a target node, perform semantic diffusion along the preset knowledge graph based on the target node, and generate an expanded query vector; an agent coordination module, configured to perform parallel retrieval on the expanded query vector based on an agent coordination mechanism to obtain a candidate retrieval result set, perform fusion and sorting on the candidate retrieval result set to obtain a preliminary sorting result, and perform confidence correction on the preliminary sorting result to obtain a corrected sorting result; a reordering output module, configured to perform multi-dimensional reordering on the corrected sorting result according to a preset scene adaptation factor, and output a digital resource retrieval result in combination with the user category; the intent enhancement module is configured to invoke a user classification agent to identify a user category, perform semantic enhancement on the multi-source query vector set according to the user category, and generate an enhanced query vector, and specifically comprises: a user classification unit, configured to invoke the user classification agent, acquire user feature information of the user, and identify the user category based on the user feature information; wherein the user category includes a teacher category, a parent category and a baby category; an intent completion unit, configured to infer missing modal information from existing modal information in the multi-modal query data, perform intent semantic completion on the multi-source query vector set based on the missing modal information, and generate a completed query vector set; an attention fusion unit, configured to calculate the completed query vector set through an attention mechanism, generate a confidence weight, weight and fuse the completed query vector set based on the confidence weight, and generate a fused query vector; a semantic enhancement unit, configured to select a corresponding differential semantic enhancement strategy according to the user category, perform enhancement processing on the fused query vector based on the differential semantic enhancement strategy, and generate the enhanced query vector; the semantic enhancement unit is configured to select a corresponding differential semantic enhancement strategy according to the user category, perform enhancement processing on the fused query vector based on the differential semantic enhancement strategy, and generate the enhanced query vector, and specifically comprises: a term standardization unit, configured to, in response to the user category being the teacher category, perform normalization processing on the fused query vector, map semantic vector representations of non-standard educational terms in the fused query vector to a semantic space of standard educational terms, and generate the enhanced query vector; a colloquial mapping unit, configured to, in response to the user category being the parent category, perform colloquial deconstruction and normalization processing on the fused query vector, map semantic vector representations of colloquial expressions in the fused query vector to the semantic space of standard educational terms, and generate the enhanced query vector. The semantic adaptation unit for young children is configured to perform the normalization processing and semantic simplification processing based on the cognitive characteristics of young children on the fusion query vector to generate the enhanced query vector in response to the user category being the young child category. The reordering output module is configured to perform multi-dimensional reordering on the modified ranking result according to a preset scene adaptation factor, and output a digital resource retrieval result in combination with the user category, and specifically includes: The scene factor acquisition unit is configured to acquire the preset scene adaptation factor, and the preset scene adaptation factor at least includes a teaching scene type, a resource use stage, and an interaction mode requirement. The multi-dimensional reordering unit is configured to generate a multi-dimensional ranking weight according to the preset scene adaptation factor and the user category, and perform weighted reordering on the modified ranking result based on the multi-dimensional ranking weight to obtain a reordering result. The result output unit is configured to perform resource adaptation filtering on the reordering result according to the user category, and output the filtered reordering result as the digital resource retrieval result.

2. The digital resource retrieval system based on multi-modal and AI agent according to claim 1, characterized in that, The query analysis module is configured to acquire multi-modal query data input by a user, pre-process, extract features, and align cross-modal semantics of the multi-modal query data to generate a multi-source query vector set, and specifically includes: The text processing unit is configured to perform word segmentation and entity recognition on text modal data in the multi-modal query data, and extract text modal features by using a language model. The image processing unit is configured to perform denoising and size normalization on image modal data in the multi-modal query data, and extract image modal features by using a visual model. The audio processing unit is configured to perform noise reduction and speech activity detection on audio modal data in the multi-modal query data, extract acoustic features, and extract semantic features by speech recognition, and the audio modal features are composed of the acoustic features and the semantic features. The video processing unit is configured to perform key frame extraction and scene segmentation on video modal data in the multi-modal query data, extract video visual features and video audio features, and generate video modal features. The cross-modal alignment unit is configured to use a contrast learning framework to construct positive and negative sample pairs between modalities, minimize the contrast loss between modalities, map the text modal features, the image modal features, the audio modal features, and the video modal features to a unified semantic space, and generate the multi-source query vector set.

3. The digital resource retrieval system based on multi-modal and AI agent of claim 1, wherein, The semantic expansion module is configured to match the enhanced query vector with a preset knowledge graph, determine a target node, perform semantic diffusion along the preset knowledge graph based on the target node, and generate an expanded query vector, and specifically includes: The graph storage unit is configured to store the preset knowledge graph, and the preset knowledge graph is constructed by entities in the education field and the association relationships between the entities. The graph embedding unit is configured to use a graph embedding model to perform embedding training on the entities and the association relationships in the preset knowledge graph to obtain node vectors corresponding to the entities. The node matching unit is configured to calculate the similarity between the enhanced query vector and the node vectors, and select a node with the highest similarity as the target node. a semantic diffusion unit, configured to perform semantic diffusion along the association relationship in the preset knowledge graph starting from the target node, to obtain an associated node associated with the target node; a query expansion unit, configured to fuse the node vector of the target node, the node vector of the associated node, and the enhanced query vector, to generate the expanded query vector.

4. The digital resource retrieval system based on multi-modal and AI agent of claim 1, wherein, The agent cooperation module is configured to perform parallel retrieval on the expanded query vector based on an agent cooperation mechanism to obtain a candidate retrieval result set, perform fusion and sorting on the candidate retrieval result set to obtain a preliminary sorting result, and perform confidence correction on the preliminary sorting result to obtain a corrected sorting result. Specifically, the agent cooperation module includes: a routing allocation unit configured with an AI routing agent, configured to analyze query features of the expanded query vector, and determine a retrieval engine to be called according to the query features; a retrieval execution unit, configured to distribute the expanded query vector to the determined retrieval engine, to trigger the retrieval engine to perform retrieval operations in parallel, and obtain the candidate retrieval result set; a multi-source fusion unit configured with an AI fusion agent, configured to perform multi-source fusion and sorting on retrieval results in the candidate retrieval result set using a sorting learning strategy, to generate the preliminary sorting result; a reflection correction unit configured with an AI reflection and correction agent, configured to evaluate the confidence of the preliminary sorting result. If the confidence is lower than a preset confidence threshold, iterative retrieval is triggered to correct the preliminary sorting result, to obtain the corrected sorting result. Otherwise, the preliminary sorting result is taken as the corrected sorting result.

5. The digital resource retrieval system based on multi-modal and AI agent of claim 4, wherein, The reflection correction unit is configured with an AI reflection and correction agent, configured to evaluate the confidence of the preliminary sorting result. If the confidence is lower than a preset confidence threshold, iterative retrieval is triggered to correct the preliminary sorting result, to obtain the corrected sorting result. Otherwise, the preliminary sorting result is taken as the corrected sorting result. Specifically, the reflection correction unit includes: a confidence evaluation unit, configured to obtain evaluation indexes of the preliminary sorting result, calculate the confidence of the preliminary sorting result according to the evaluation indexes, and generate a confidence evaluation result; and when the confidence meets the preset confidence threshold, take the preliminary sorting result as the corrected sorting result; a query correction unit, configured to, when the confidence is lower than the preset confidence threshold, generate a correction retrieval strategy according to the confidence evaluation result; an iterative retrieval unit, configured to send the correction retrieval strategy to the AI routing agent, to trigger the agent cooperation module to perform retrieval, fusion, and sorting again according to the correction retrieval strategy, to generate an updated preliminary sorting result; an iterative control unit, configured to evaluate the confidence of the updated preliminary sorting result. When the confidence of the updated preliminary sorting result is lower than the preset confidence threshold, iterative retrieval is repeatedly triggered until the confidence of the updated preliminary sorting result meets the preset confidence threshold, and the updated preliminary sorting result is taken as the corrected sorting result.

6. A digital resource retrieval method based on multi-modal and AI agent, characterized in that, The method is applied to the digital resource retrieval system based on multi-modal and AI intelligent agent according to any one of claims 1-5, and the method comprises: Obtaining multi-modal query data input by a user, preprocessing, feature extraction and cross-modal semantic alignment of the multi-modal query data, and generating a multi-source query vector set; Calling a user classification intelligent agent to identify a user category, performing semantic enhancement on the multi-source query vector set according to the user category, and generating an enhanced query vector; Matching the enhanced query vector with a preset knowledge graph, determining a target node, performing semantic diffusion along the preset knowledge graph based on the target node, and generating an extended query vector; Based on the intelligent agent cooperation mechanism, performing parallel retrieval on the extended query vector to obtain a candidate retrieval result set, performing fusion and sorting on the candidate retrieval result set to obtain a preliminary sorting result, and performing confidence correction on the preliminary sorting result to obtain a corrected sorting result; According to a preset scene adaptation factor, the modified sorting result is multi-dimensionally reordered, and the digital resource retrieval result is output in combination with the user category.

Citation Information

Patent Citations

  • Intention classification method and device based on vector retrieval and context awareness and medium

    CN120448929A

  • Infant teaching resource intelligent searching and recommending system based on AI large model

    CN120849700A