Method and device for retrieving material information

By extracting multi-dimensional structured information from the material library and building a label system, the problem of low correlation and accuracy of search results in the existing technology is solved, and the accurate understanding and efficient search of user intentions are achieved, and the quality of search results is improved.

CN120045693APending Publication Date: 2025-05-27SHANGHAI IQIYI NEW MEDIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145081.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, the material content in the material library is complex and the amount of information is huge, which makes users spend a lot of time and energy to sort and extract. The existing search system cannot accurately understand the user's intentions, resulting in low relevance and accuracy of the search results.

Method used

By extracting multi-dimensional structured information from each text material in the material library, and building a tag system based on this information, mapping the search requests input by the user into the tag system, determining the correlation information corresponding to the hit tag, and sorting the correlation information through multi-feature fusion to output material information that best meets user needs.

Benefits of technology

It realizes an accurate understanding of user intentions, improves the relevance and accuracy of search results, reduces the time cost of users to find required information, and improves the accuracy of text material retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045693A_ABST
    Figure CN120045693A_ABST
Patent Text Reader

Abstract

The invention provides a material information retrieval method and device. The method comprises the following steps: extracting multi-dimensional structured information from each text material in a material library; according to the multi-dimensional structured information and each text material, a label system of the material library is constructed, each structured information corresponds to a label, and each text material corresponds to a label; mapping a retrieval request input by a user into the tag system, and determining associated information corresponding to a hit tag; and sorting the associated information in a multi-feature fusion mode, and outputting the material information of which the sorting position is located in front of the set position. According to the method, the text material retrieval accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data retrieval, and in particular to a method and device for retrieval of material information. Background Art

[0002] Users often need to retrieve the required content from the complex data in the material library. Currently, the material content in the material library is complex and the amount of information is huge. Users need to spend a lot of time and energy to organize and extract it. Moreover, most existing retrieval systems are based on keyword matching and cannot accurately understand the user's intentions, resulting in low relevance and accuracy of the retrieval results, which leads to inaccurate retrieval results. Summary of the invention

[0003] In order to solve the above technical problem or at least partially solve the above technical problem, the present application provides a method and device for retrieving material information.

[0004] In a first aspect, the present application provides a method for retrieving material information, the method comprising:

[0005] Extract multi-dimensional structured information from each text material in the material library;

[0006] Constructing a label system of the material library according to the multi-dimensional structured information and each text material, wherein each structured information corresponds to a label and each text material corresponds to a label;

[0007] Mapping the search request input by the user to the tag system, and determining the associated information corresponding to the hit tag;

[0008] The associated information is sorted by means of multi-feature fusion, and the material information whose sorting position is before the set position is output.

[0009] Optionally, constructing a label system of the material library according to the multi-dimensional structured information and each text material includes:

[0010] Determine the semantic label corresponding to the structured information of each dimension;

[0011] Determining a material tag of the text material according to a source of the text material;

[0012] Using the semantic label and the material label as the overall label of the text material;

[0013] The label system of the material library is obtained by integrating the overall labels of multiple text materials.

[0014] Optionally, after constructing a label system of the material library according to the multi-dimensional structured information and each text material, the method further includes:

[0015] Determine a vector representation of the structured information and a vector index of the vector representation, wherein the vector index is used to perform vector similarity calculation;

[0016] The semantic label of the structured information is integrated into the vector index, and the vector index and the semantic label are stored collaboratively.

[0017] Optionally, mapping the search request input by the user to the tag system and determining the associated information corresponding to the hit tag includes:

[0018] Extracting the user intention in the search request and rewriting the search request, wherein the rewriting of the search request is used to optimize and adjust the search content;

[0019] Mapping the user intention to the tag system to determine a hit tag;

[0020] Determine the vector index of the structured information under the hit label constraint;

[0021] Calculate the vector similarity between the vector index of the rewritten retrieval request and the vector index under the label constraint;

[0022] The structured information whose similarity calculation score exceeds a set threshold is used as the associated information of the search request.

[0023] Optionally, sorting the associated information by means of multi-feature fusion includes:

[0024] Determining a feature combination of the associated information, wherein the feature combination includes a similarity calculation score, timeliness, and a quality score of a corresponding text material;

[0025] Adjusting the weight of each feature in the feature combination based on the user's historical behavior and preferences;

[0026] Performing feature-weighted summation on the associated information to determine a comprehensive score of the associated information;

[0027] The associated information is sorted in descending order of comprehensive scores.

[0028] Optionally, the multi-dimensional structured information includes: story summary, detailed introduction, character portraits and story chain.

[0029] Optionally, before extracting multi-dimensional structured information from each text material in the material library, the method includes:

[0030] Get multiple text materials to be processed;

[0031] Preprocessing the text material to be processed, wherein the preprocessing includes text denoising, rule segmentation and paragraph division;

[0032] The preprocessed text material is stored in the material library according to the specified format.

[0033] In a second aspect, the present application provides a device for retrieving material information, the device comprising:

[0034] An extraction module is used to extract multi-dimensional structured information from each text material in the material library;

[0035] A construction module, used to construct a label system of the material library according to the multi-dimensional structured information and each text material, wherein each structured information corresponds to a label and each text material corresponds to a label;

[0036] A mapping module, used to map the search request input by the user to the tag system and determine the associated information corresponding to the hit tag;

[0037] The sorting module is used to sort the related information by means of multi-feature fusion, and output the material information whose sorting position is before the set position.

[0038] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0039] Memory, used to store computer programs;

[0040] The processor is used to implement any of the steps of the material information retrieval method when executing the program stored in the memory.

[0041] In a fourth aspect, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, any of the steps of the material information retrieval method is implemented.

[0042] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:

[0043] The method provided in the embodiment of the present application extracts multi-dimensional structured information from text materials, and constructs a label system for the structured information of each dimension and the overall text material, so that the label system contains both labels for each detail dimension (structured information) and labels for the overall dimension (overall text material). Once the search request input by the user hits the relevant label, the system can accurately understand the user's query intention through detailed and multi-level labels, and then determine the related information related to the search request under the constraints of the hit label, ensuring that the related information found accurately meets the user's needs. Finally, the system outputs the text material that best meets the user's needs by performing feature fusion sorting on the related information, thereby improving the accuracy of text material retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0046] Figure 1 A schematic diagram of the hardware environment of a material information retrieval method provided in an embodiment of the present application;

[0047] Figure 2 A flowchart of a method for retrieving material information provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of a retrieval system architecture for material information provided in an embodiment of the present application;

[0049] Figure 4 A schematic diagram of the structure of a material information retrieval device provided in an embodiment of the present application;

[0050] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0052] In the subsequent description, the suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present application and have no specific meaning. Therefore, "module" and "component" can be used interchangeably.

[0053] In order to solve the problem of inaccurate material query in the prior art, this application constructs tags through multi-dimensional structured information in text materials to achieve precise retrieval and improve the accuracy of text material retrieval.

[0054] This application is applicable to multiple application scenarios, such as film and television creation, literary research, news reporting, and legal case analysis.

[0055] In order to solve the problem mentioned in the background technology, according to one aspect of an embodiment of the present application, an embodiment of a method for retrieving material information is provided.

[0056] Optionally, in the embodiment of the present application, the above-mentioned method for retrieving material information can be applied to: Figure 1 In the hardware environment composed of the terminal 101 and the server 103 shown in FIG. Figure 1 As shown, the server 103 stores a material library and a label system. The server 103 is connected to the terminal 101 through a network. The user enters a search request on the terminal, and the server searches for the material information required by the user in the material library according to the search request entered by the user. A database 105 can be set on the server or independently of the server to provide data storage services for the server 103. The above network includes but is not limited to: a wide area network, a metropolitan area network or a local area network, and the terminal 101 includes but is not limited to a PC, a mobile phone, a tablet computer, etc.

[0057] A method for retrieving material information in an embodiment of the present application may be executed by the server 103 or by the terminal 101 .

[0058] The following will describe in detail a method for retrieving material information provided by an embodiment of the present application in conjunction with a specific implementation method. Figure 2 As shown, the specific steps are as follows:

[0059] Step 201: extracting multi-dimensional structured information from each text material in the material library;

[0060] Step 202: constructing a label system of the material library according to the multi-dimensional structured information and each text material, wherein each structured information corresponds to a label and each text material corresponds to a label;

[0061] Step 203: Map the search request input by the user to the tag system, and determine the associated information corresponding to the hit tag;

[0062] Step 204: sort the associated information by means of multi-feature fusion, and output the material information whose sorting position is before the set position.

[0063] The system has a library containing a variety of text materials, which can be case reports, novels, movie scripts, news reports, academic papers and other forms of text materials. In order to better manage and retrieve these materials, the system uses large-scale language model technology based on deep learning to process each text material and extract multi-dimensional structured information.

[0064] Structured information refers to data with clear meaning and format, usually including but not limited to story summary, detailed introduction, character portrait, and event chain. Specifically, the story summary summarizes the core plot and main events of the material; the detailed introduction provides detailed background information and plot development; the character portrait describes the main characters and their characteristics that appear in the material; and the event chain lists the main events that occurred in the material and their sequence. Through this multi-level information extraction method, the system can more comprehensively understand and organize each text material.

[0065] The system further labels each extracted structured information to determine its semantic label. For example, story summary labels may include romance, suspense, science fiction, etc.; character type labels include heroes, villains, supporting roles, etc.; emotional labels involve love, friendship, hatred, fear, etc.; theme labels include growth, revenge, redemption, adventure, etc. In addition, the system also labels the text material according to its source, such as marking the text as a case report, script or novel, and further subdividing the case report into specific categories such as fraud, theft or intentional injury.

[0066] The tags of the above structured information and the tags of the text material itself together constitute the tag system of the entire material library. The tag system is a classification framework composed of multiple tags, which is used to describe and organize all the materials in the material library. Each tag represents a specific theme or attribute, and some tags directly correspond to specific structured information. Such a tag system not only helps users quickly locate the materials of interest, but also supports efficient search and recommendation functions.

[0067] When the user enters a search request in the terminal, the system first parses the user's query intent, and then matches the parsed intent with the tags in the tag system. For example, if the user queries "science fiction novels about time travel", the system will match the tags "novel", "science fiction" and "time travel". The system first screens the content that matches the "novel" tag as the source of the material, and on this basis further determines the structured information corresponding to the "science fiction" and "time travel" tags, and uses the structured information as related information. Each piece of related information has multiple features, such as relevance score, timeliness, quality score of the corresponding text material, etc. The system sorts the related information by integrating these features, and finally outputs the material information whose sorting position is before the set position to provide the most relevant search results to the user.

[0068] For example, there is a novel named "xx" in the material library, and its structured information is as follows: the story summary is "In the future world, humans have invented time travel technology, and the protagonist has solved a series of historical puzzles through time travel"; the details describe the background setting and plot development in detail; the character portraits include "Dr. Li, a genius physicist; Engineer Wang, responsible for building a time machine"; the event chain records "time travel experiment started--> parallel universe discovered--> resource depletion problem solved". The system has marked the novel with tags such as "novel", "science fiction", and "time travel". When the user enters the query "science fiction novel about time travel", the system will match these tags, and select "xx" that meets the conditions from the material library, further extract the related information in its structured information, and after multi-feature fusion sorting, give priority to displaying the relevant information of this novel to the user. In this way, users can quickly find high-quality text materials that meet their needs.

[0069] In this application, multi-dimensional structured information is extracted from text materials, and a label system is constructed for the structured information of each dimension and the overall text material, so that the label system contains both labels for each detail dimension (structured information) and labels for the overall dimension (overall text material). Once the search request input by the user hits the relevant label, the system can accurately understand the user's query intention through detailed and multi-level labels, and then determine the related information related to the search request under the constraints of the hit label, ensuring that the related information found accurately meets the user's needs. Finally, the system outputs the text material that best meets the user's needs by feature fusion sorting of the related information, thereby improving the accuracy of text material retrieval.

[0070] As an optional implementation, in step 201, constructing a label system of a material library according to multi-dimensional structured information and each text material includes the following steps:

[0071] Step S11: Determine the semantic label corresponding to the structured information of each dimension;

[0072] Step S12: determining a material label of the text material according to the source of the text material;

[0073] Step S13: taking the semantic label and the material label as the overall label of the text material;

[0074] Step S14: Obtaining a label system of the material library by integrating the overall labels of multiple text materials.

[0075] When building a labeling system for the material library, we first need to determine the semantic labels corresponding to the structured information of each dimension. This means identifying and marking the core features of the extracted multi-dimensional structured information such as story summaries, detailed introductions, character portraits, and event chains. For example, story summaries may be labeled with labels such as "science fiction", "suspense", or "love"; character portraits can be labeled with role type labels such as "hero", "villain", or "supporting role"; emotional labels include "love", "friendship", "hatred", or "fear"; theme labels include "growth", "revenge", "redemption", or "adventure". These labels not only help the system understand the core elements of the text content, but also provide a basis for subsequent retrieval and recommendation.

[0076] Next, determine the material label based on the source of the text material. Text materials can be text materials in various forms such as case reports, novels, movie scripts, news reports or academic papers. For each type of material, the system will perform specific classification and labeling. For example, case reports can be subdivided into labels such as "fraud", "theft" or "intentional injury"; novels can be labeled with labels such as "science fiction", "suspense" or "history" according to their main plot or style. This source-based labeling not only refines the classification of materials, but also enables users to more accurately locate the required type of material. By combining semantic labels and material labels, an overall label for each text material is formed, ensuring that each material has a comprehensive and detailed description framework.

[0077] After integrating the overall tags of multiple text materials, a tag system for the entire material library can be constructed. This tag system is a classification framework composed of multiple tags, which is used to describe and organize all the materials in the material library. Each tag represents a specific theme or attribute and corresponds to specific structured information. In this way, the tag system can not only help users quickly find materials of interest, but also support efficient search and recommendation functions. For example, when a user queries "science fiction novels about time travel", the system can quickly match relevant tags (such as "science fiction", "time travel") through the tag system and filter out text materials that meet the conditions. This tag-based retrieval method greatly improves the efficiency and accuracy of information retrieval and reduces the time cost for users to find the information they need.

[0078] This labeling system based on multi-dimensional structured information and text material sources improves retrieval accuracy. Through detailed and multi-level labels, the system can more accurately understand and match user query intentions.

[0079] As an optional implementation, after constructing the label system of the material library according to the multi-dimensional structured information and each text material, the following contents are also included:

[0080] Step S21: determining a vector representation of the structured information and a vector index of the vector representation, wherein the vector index is used to perform vector similarity calculation;

[0081] Step S22: Integrate the semantic tags of the structured information into the vector index, and store the vector index and the semantic tags together.

[0082] In the process of processing and optimizing the material library, the received structured information needs to be processed first. The system receives the structured information obtained from the extraction stage, including story summaries, character portraits, event chains, etc. This information is the basis for subsequent processing. In order to improve the semantic expression ability of the text, the system uses a dual-tower neural network to encode and fine-tune the text and annotate some relevant data. In this process, contrastive learning technology is applied to further enhance the semantic expression ability of the vector. In this way, each structured information fragment can be accurately represented as a high-dimensional vector, capturing its deep semantic features in the vector space. Then, using the fine-tuned trained model, the system vectorizes each information fragment, extracts independent vector representations for the story, character, event, and original information, and ensures that these vectors not only retain the semantic information of the original text, but also can perform efficient similarity calculations in the vector space.

[0083] In order to support fast similarity calculation, the system uses vector database technology (such as Elasticsearch, Vespa, etc.) to build an efficient vector retrieval index structure. This index structure is particularly suitable for efficient retrieval on large-scale data sets, ensuring that users can quickly find the most relevant text materials. At the same time, the system integrates semantic tags into the metadata index to form a multimodal, multi-dimensional retrieval index system. In this way, semantic tags are stored in conjunction with vector indexes to ensure that relevant content can be quickly located during retrieval.

[0084] In addition, to ensure the real-time and consistency of the system, the semantic vector generation and index construction module ensures that the vector index can be updated in real time to adapt to the addition of new data and the modification of old data. This enables the system to dynamically respond to the latest content changes and maintain the integrity and consistency of the index. Maintaining the integrity and consistency of the index is crucial to ensuring the accuracy and efficiency of retrieval. Through real-time updates and consistency maintenance, the system can provide the latest and most accurate retrieval results at any point in time, thereby improving the user experience.

[0085] Finally, the system determines the vector representation of each structured information and its corresponding vector index for vector similarity calculation, and integrates the semantic tags of the structured information into the vector index, and stores the vector index and semantic tags together to ensure that the most suitable content can be quickly located and returned during retrieval. This method not only improves the retrieval accuracy, but also enhances the flexibility and response speed of the system.

[0086] As an optional implementation, in step 203, mapping the search request input by the user to the tag system, and searching for associated information related to the search request from the structured information under the hit tag constraint includes the following contents:

[0087] Step S31: extracting the user intention in the search request and rewriting the search request, wherein the rewriting of the search request is used to optimize and adjust the search content;

[0088] Step S32: Map the user intention to the tag system and determine the hit tag;

[0089] Step S33: determining the vector index of the structured information under the hit tag constraint;

[0090] Step S34: performing vector similarity calculation between the vector index of the rewritten search request and the vector index under the tag constraint;

[0091] Step S35: The structured information whose similarity calculation score exceeds the set threshold is used as the associated information of the search request.

[0092] When processing a user's search request, the system first obtains and preprocesses the search request input by the user, removes noise and other interference factors, and ensures the accuracy of subsequent steps. Then, a deep learning model (such as a large language model) is used to semantically understand the user input, analyze the context and language of the input, and extract the user's intention. The system then rewrites the search request. The purpose of rewriting is to optimize and adjust the search content so that the query is more in line with the system's index structure and improve the relevance of the search results. The system may add, delete, or replace keywords in the query to improve the accuracy of the search. This process not only retains the user's original intention, but also enhances the expressiveness of the query. For example, if a user enters "science fiction about time travel", the system may rewrite it as "science fiction containing time travel plots" or "novels involving time travel and future technology". This rewriting optimizes the query content and improves the accuracy and comprehensiveness of the search.

[0093] The system maps the user's intent to a pre-built tag system to determine the hit tag. The tag system covers multiple dimensions such as material type (such as novels, scripts), story type (such as science fiction, suspense), character type (such as heroes, villains), and emotional tags (such as love, friendship). Through this mapping, the system can quickly locate tags related to user intent. For example, the user intent "science fiction novels about time travel" will be mapped to tags such as "science fiction", "time travel", and "novel". These tags provide clear directions and constraints for subsequent searches, ensuring the relevance and accuracy of the search results.

[0094] After determining the hit label constraints, the system uses the same model as when building the offline material library to vectorize the rewritten query and generate a vector representation of the query. Then, the vector similarity between the structured information of each material in different dimensions (story content, character portrait, event chain) and the retrieval request is calculated in the vector database. This method can quickly find the material most relevant to the user query in a high-dimensional space. In order to further ensure the quality of the retrieval results, the system sets a relevance threshold, and only materials with a similarity score exceeding this threshold will be considered relevant. In this way, the system can screen out high-quality materials that best meet user needs.

[0095] Finally, the system calculates the vector similarity between the vector index of the rewritten retrieval request and the vector index under the label constraint, and outputs the structured information whose similarity score exceeds the set threshold as the associated information of the retrieval request. For example, when a user queries "science fiction novels about time travel", the system will give priority to those novels with high similarity scores and the labels "science fiction" and "time travel". This method not only improves retrieval accuracy, but also enhances user experience, supports flexible queries, and ensures efficient retrieval on large-scale data sets.

[0096] In this application, through semantic understanding and query rewriting, the system can more accurately capture the user's intention and optimize the query content, thereby improving the accuracy of the retrieval results; combined with the tag system and vector similarity calculation, the system can provide high-quality search results to ensure that the search results meet user needs.

[0097] As an optional implementation, in step 204, sorting the associated information by multi-feature fusion includes the following:

[0098] Step S41: determining a feature combination of the associated information, wherein the feature combination includes a similarity calculation score, timeliness, and a quality score of a corresponding text material;

[0099] Step S42: adjusting the weight of each feature in the feature combination based on the user's historical behavior and preferences;

[0100] Step S43: performing feature weighted summation on the associated information to determine a comprehensive score of the associated information;

[0101] Step S44: Sort the associated information in descending order of comprehensive scores.

[0102] After determining the relevant information related to the user's search request, the system needs to further process this information to ensure that the search results finally displayed to the user are the high-quality content that best meets their needs. First, the system determines the feature combination of the relevant information, which includes the similarity calculation score, timeliness, and the quality score of the corresponding text material. The similarity calculation score reflects the semantic similarity between each material and the user's query; timeliness takes into account the temporal relevance of the material, which is especially important for news reports or real-time updated content; the quality score is an assessment of the overall quality of the text material, which may be based on ratings, comments or other indicators.

[0103] In order to more accurately meet the needs of users, the system will adjust the weight of each feature in the feature combination based on the user's historical behavior and preferences. For example, if the user often pays attention to the latest scientific and technological developments, the system may increase the weight of timeliness; if the user prefers high-quality classic works, the system will increase the weight of the quality score accordingly. This personalized adjustment enables the system to provide more customized search results based on the user's unique preferences. By analyzing the user's historical behavior (such as past query records, click behavior, etc.) and clear preference settings, the system can dynamically adjust the weight of each feature, thereby improving the relevance and satisfaction of the search results.

[0104] After adjusting the feature weights, the system performs a weighted summation of the associated information to determine the comprehensive score of each piece of associated information. This process generates a comprehensive score by weighted summation of each feature, which reflects the performance of each material in multiple dimensions. For example, a material with a high similarity score but low timeliness may have a lower comprehensive score than another material with a slightly lower similarity but higher timeliness when the timeliness weight is higher. In this way, the system is able to balance various factors in multiple dimensions to ensure that the final displayed results not only meet the user's query intent, but also meet their personalized needs.

[0105] Finally, the system sorts the related information in descending order of comprehensive scores, and displays the sorting results to the user in the form of a list or card. This method ensures that the most relevant and high-quality materials are displayed first, improving the user experience. For example, when a user queries "science fiction novels about time travel", the system will not only find the closest text fragment in the vector space, but also conduct a comprehensive evaluation based on factors such as timeliness and quality score. In the end, the system will give priority to those novels with the highest comprehensive scores, because it not only highly matches the user's query semantically, but also performs well in timeliness and quality. This multi-dimensional sorting mechanism not only improves the accuracy and relevance of the retrieval results, but also enhances the flexibility and responsiveness of the system, ensuring that users can quickly find the text materials that best meet their needs.

[0106] To further enhance the user experience, the system provides a result filtering function, allowing users to refine the displayed results according to their needs. For example, users can filter by material type (such as novels, scripts, news reports, etc.), release time (latest release, release within a specific time period), etc. This flexible filtering mechanism enables users to locate the required content more accurately and reduce the interference of irrelevant information.

[0107] In addition, the system also supports interactive feedback functions to encourage users to interact with search results. Users can express their interest and evaluation of a piece of material by liking, commenting, and collecting it. The system will record these feedback data and use them to optimize subsequent sorting strategies. This dynamic adjustment mechanism based on user feedback not only enhances the system's personalized service capabilities, but also continuously improves and optimizes the quality of search results, ensuring that users always get the content that best suits their interests and preferences.

[0108] In this application, by comprehensively considering multiple dimensions such as similarity calculation scores, timeliness and quality scores, the system can more comprehensively evaluate the relevance of each piece of structural information, thereby improving the accuracy of retrieval results; personalized feature weight adjustment enables the system to provide customized search results based on the user's historical behavior and preferences to meet the diverse needs of different users; the system can adjust feature weights and update sorting results in real time to ensure that the addition of the latest data and the modification of old data will not affect the accuracy and efficiency of the retrieval.

[0109] As an optional implementation, before extracting multi-dimensional structured information from each text material in the material library, the method includes the following contents:

[0110] Step S51: obtaining a plurality of text materials to be processed;

[0111] Step S52: preprocessing the text material to be processed, wherein the preprocessing includes text denoising, rule segmentation and paragraph division;

[0112] Step S53: storing the preprocessed text material in a material library according to a specified format.

[0113] In the process of building a material library, it is first necessary to obtain multiple text materials to be processed. These materials can be from multiple data sources and have different structures. Next, the system preprocesses these text materials to be processed to ensure the accuracy and efficiency of subsequent analysis and processing. Preprocessing includes several key steps: First, text denoising, removing irrelevant characters (such as extra spaces, special characters, etc.), this step can clean up the noise in the text and improve data quality. Secondly, rule segmentation, segmenting the text according to specific rules (such as sentence boundaries, paragraph boundaries), so that each fragment can be processed independently. Finally, paragraph division, dividing long texts into logical paragraphs, is convenient for subsequent structured information extraction and semantic understanding. Through these preprocessing steps, the system can convert the original text into a form that is easier to process and analyze, laying a solid foundation for high-quality information extraction. After completing the preprocessing, the system stores the preprocessed text materials in the material library (such as Mongodb database) in a specified format (such as json format). This standardized storage method not only helps to maintain the consistency and integrity of the data, but also improves the efficiency of subsequent processing.

[0114] In this application, through preprocessing steps such as text denoising, rule segmentation, and paragraph division, the system can effectively clean up the noise in the original text and ensure the data quality of subsequent processing. High-quality data is the basis for accurate information extraction and semantic understanding, and helps to improve the overall performance of the system. The preprocessed text materials are stored in the material library in the specified format. This standardized storage method improves the efficiency of subsequent processing.

[0115] This solution aims to achieve efficient and accurate text material retrieval and recommendation through two core parts: offline material library construction and online retrieval. Figure 3 A schematic diagram of the system architecture.

[0116] Part 1: Offline material library construction.

[0117] 1.Material collection and preprocessing module.

[0118] The system first collects a large number of creative materials from multiple sources (such as manual uploads, automatic crawling or API interfaces), including but not limited to criminal case reports, scripts and novels. The system supports the input of multi-source heterogeneous data and standardizes and uniformly processes them.

[0119] Perform text denoising on each material to remove irrelevant characters. Segment the text according to specific rules (such as sentence boundaries and paragraph boundaries) so that each segment can be processed independently and divide long texts into logical paragraphs to facilitate subsequent structured information extraction and semantic understanding.

[0120] 2. Structured information extraction module.

[0121] With the support of a large language model, the system scores the quality of each material. This step ensures that the system can identify high-quality materials.

[0122] The system extracts key structured information from text materials, including story summaries, detailed introductions, character portraits, event chains, etc., and scores the quality of the text materials. The scoring criteria include: the originality of the materials, the completeness of the materials, the richness of the materials, etc. The system labels the extracted structured information, covering multiple dimensions such as story type, character type, and emotional tags. The system also labels the text materials according to their sources.

[0123] 3.Semantic vector generation and index construction module.

[0124] Based on the labeled data set, the system uses a dual-tower neural network architecture to fine-tune the text encoding. In this way, the system can extract independent vector representations of the story, characters, events, and original text information. The application of contrastive learning technology further enhances the semantic expression ability of these vectors, ensuring that the deep semantic features of the text can be accurately captured in high-dimensional space.

[0125] The system builds an efficient vector retrieval index structure, then collaboratively stores the label system and vector index, and supports real-time index updates and maintenance, ensuring that the addition of the latest data and the modification of old data will not affect the accuracy and efficiency of retrieval.

[0126] Part II: Online Search.

[0127] 1. Intelligent intention recognition module.

[0128] The system uses a large language model to analyze the user's search request. Through semantic understanding and context analysis, the system can accurately identify the user's intention. Then the user's intention is mapped to a preset tag system and the user's query is rewritten. The rewriting process optimizes the query content to better match the system's index and retrieval mechanism, ensuring that the user's search intention can be accurately understood.

[0129] 2. Multi-dimensional correlation retrieval module.

[0130] The system determines the hit tag based on the user's intent, and then performs a search within the tag constraint. It calculates the vector similarity of the structural information under the hit tag and the rewritten search request. Only the structural information with a similarity score exceeding the threshold will be considered relevant, thereby filtering out the related information that best meets the user's needs.

[0131] 3. Intelligent sorting and output module.

[0132] The sorting algorithm based on multi-feature fusion comprehensively considers multi-dimensional features such as relevance score, timeliness, and material quality. The system adjusts the weight of each feature according to the user's personalized preferences to ensure the personalization and accuracy of the sorting results. Then the comprehensive score of each structural information is calculated, and the materials are sorted according to the comprehensive score, with materials with high scores being displayed first.

[0133] The sorted materials are displayed to users in the form of lists or cards, and users can further filter the results as needed (such as by material type, time, etc.). Users can also interact with the results, such as liking, commenting, and collecting, and the system will record these feedback data to optimize subsequent sorting strategies.

[0134] The system can accurately understand user needs, efficiently retrieve relevant materials, and present them to users in the most optimized manner, thereby improving retrieval efficiency.

[0135] Based on the same technical concept, the embodiment of the present application also provides a retrieval device for material information, such as Figure 4 As shown, the device comprises:

[0136] An extraction module 401 is used to extract multi-dimensional structured information from each text material in the material library;

[0137] A construction module 402 is used to construct a label system of the material library according to the multi-dimensional structured information and each text material, wherein each structured information corresponds to a label and each text material corresponds to a label;

[0138] A mapping module 403 is used to map the search request input by the user into a tag system and determine the associated information corresponding to the hit tag;

[0139] The sorting module 404 is used to sort the related information by means of multi-feature fusion, and output the material information whose sorting position is before the set position.

[0140] Optionally, the construction module 402 is used to:

[0141] Determine the semantic label corresponding to the structured information of each dimension;

[0142] Determine the material label of the text material according to the source of the text material;

[0143] Semantic tags and material tags are used as overall tags for text materials;

[0144] The label system of the material library is obtained by integrating the overall labels of multiple text materials.

[0145] Optionally, the device is also used for:

[0146] Determine a vector representation of the structured information and a vector index of the vector representation, wherein the vector index is used to perform vector similarity calculation;

[0147] The semantic tags of structured information are integrated into the vector index, and the vector index and semantic tags are stored collaboratively.

[0148] Optionally, the mapping module 403 is used to:

[0149] Extracting user intent from a search request and rewriting the search request, wherein the rewriting of the search request is used to optimize and adjust the search content;

[0150] Map user intent to the tag system and determine the hit tag;

[0151] Determine the vector index of the structured information under the hit label constraint;

[0152] Calculate the vector similarity between the vector index of the rewritten retrieval request and the vector index under the label constraint;

[0153] The structured information whose similarity calculation score exceeds the set threshold is used as the associated information of the retrieval request.

[0154] Optionally, the sorting module 404 is used to:

[0155] Determine a feature combination of the associated information, wherein the feature combination includes a similarity calculation score, timeliness, and a quality score of the corresponding text material;

[0156] Adjust the weight of each feature in the feature combination based on the user's historical behavior and preferences;

[0157] Perform feature-weighted summation on the associated information to determine the comprehensive score of the associated information;

[0158] Sort the related information in descending order according to the comprehensive scores.

[0159] Optionally, the multi-dimensional structured information includes: story summary, detailed introduction, character portrait and story chain.

[0160] Optionally, the device is also used for:

[0161] Get multiple text materials to be processed;

[0162] Preprocessing the text material to be processed, wherein the preprocessing includes text denoising, rule segmentation and paragraph division;

[0163] The preprocessed text material is stored in the material library according to the specified format.

[0164] Based on the same technical concept, an embodiment of the present invention further provides an electronic device, such as Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503 and a communication bus 505, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 505.

[0165] Memory 503, used for storing computer programs;

[0166] The processor 501 is used to implement the above steps when executing the program stored in the memory 503. The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0167] The communication interface is used for communication between the above electronic device and other devices.

[0168] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0169] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0170] In another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0171] In another embodiment of the present invention, a computer program product including instructions is provided, which enables the computer to execute any one of the methods in the above embodiments when the computer program product is executed on the computer.

[0172] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.

[0173] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0174] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for retrieving material information, characterized in that: The method comprises: Extract multi-dimensional structured information from each text material in the material library; Constructing a label system of the material library according to the multi-dimensional structured information and each text material, wherein each structured information corresponds to a label and each text material corresponds to a label; Mapping the search request input by the user to the tag system, and determining the associated information corresponding to the hit tag; The associated information is sorted by means of multi-feature fusion, and the material information whose sorting position is before the set position is output.

2. The method according to claim 1, characterized in that The label system of the material library is constructed according to the multi-dimensional structured information and each text material, including: Determine the semantic label corresponding to the structured information of each dimension; Determining a material tag of the text material according to a source of the text material; Using the semantic label and the material label as the overall label of the text material; The label system of the material library is obtained by integrating the overall labels of multiple text materials.

3. The method according to claim 1, characterized in that After constructing the label system of the material library according to the multi-dimensional structured information and each text material, the method further includes: Determine a vector representation of the structured information and a vector index of the vector representation, wherein the vector index is used to perform vector similarity calculation; The semantic label of the structured information is integrated into the vector index, and the vector index and the semantic label are stored collaboratively.

4. The method according to claim 3, characterized in that Mapping the search request input by the user to the tag system and determining the associated information corresponding to the hit tag includes: Extracting the user intention in the search request and rewriting the search request, wherein the rewriting of the search request is used to optimize and adjust the search content; Mapping the user intention to the tag system to determine a hit tag; Determine the vector index of the structured information under the hit label constraint; Calculate the vector similarity between the vector index of the rewritten retrieval request and the vector index under the label constraint; The structured information whose similarity calculation score exceeds a set threshold is used as the associated information of the search request.

5. The method according to claim 1, characterized in that Sorting the associated information by means of multi-feature fusion includes: Determining a feature combination of the associated information, wherein the feature combination includes a similarity calculation score, timeliness, and a quality score of a corresponding text material; Adjusting the weight of each feature in the feature combination based on the user's historical behavior and preferences; Performing feature-weighted summation on the associated information to determine a comprehensive score of the associated information; The associated information is sorted in descending order of comprehensive scores.

6. The method according to claim 1, characterized in that The multi-dimensional structured information includes: story summary, detailed introduction, character portrait and story chain.

7. The method according to claim 1, characterized in that Before extracting multi-dimensional structured information from each text material in the material library, the method includes: Get multiple text materials to be processed; Preprocessing the text material to be processed, wherein the preprocessing includes text denoising, rule segmentation and paragraph division; The preprocessed text material is stored in the material library according to the specified format.

8. A material information retrieval device, characterized in that: The device comprises: An extraction module is used to extract multi-dimensional structured information from each text material in the material library; A construction module, used to construct a label system of the material library according to the multi-dimensional structured information and each text material, wherein each structured information corresponds to a label and each text material corresponds to a label; A mapping module, used to map the search request input by the user to the tag system and determine the associated information corresponding to the hit tag; The sorting module is used to sort the related information by means of multi-feature fusion, and output the material information whose sorting position is before the set position.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 7 are implemented.