Information retrieval method and device for physical training, electronic equipment and storage medium

By combining intelligent agents and large language models to retrieve knowledge from the physical training knowledge base, the problems of information fragmentation and low retrieval efficiency in adolescent physical training are solved, and efficient and accurate knowledge acquisition is achieved.

CN120821884APending Publication Date: 2025-10-21BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510830773.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Problems of information fragmentation, low retrieval efficiency and inaccurate knowledge matching in the process of acquiring physical training knowledge for teenagers.

Method used

By using an intelligent agent to retrieve information from a pre-built physical training knowledge base and combining it with a large language model to provide responses, the accuracy and efficiency of the question answers are ensured.

Benefits of technology

It improves the accuracy and efficiency of answers to questions related to physical training, adapts to the ever-increasing knowledge demand, and supports multi-level, fine-grained knowledge retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821884A_ABST
    Figure CN120821884A_ABST
Patent Text Reader

Abstract

The invention provides an information retrieval method and device for physical training, electronic equipment and a storage medium, and relates to the technical field of data processing, in particular to the technical field of natural language processing, intelligent searching, intelligent agents and large models. The method comprises the following steps: retrieving a pre-constructed physical training knowledge base through an intelligent agent according to physical training related problems; and obtaining reply information of the physical training related questions according to the retrieval result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, specifically to the field of natural language processing, intelligent search, intelligent agents and large model technologies, and especially to an information retrieval method, device, electronic device and storage medium for physical training. Background Art

[0002] Currently, the process of acquiring knowledge about physical training for teenagers relies on manual review of paper books or electronic documents, which is time-consuming and labor-intensive, and has problems such as information fragmentation, low retrieval efficiency, and inaccurate knowledge matching. Summary of the Invention

[0003] The present disclosure provides a method, device, electronic device and storage medium for information retrieval of physical training.

[0004] According to one aspect of the present disclosure, a method for retrieving information about physical training is provided, comprising:

[0005] Receive questions related to physical training;

[0006] Searching a pre-built physical training knowledge base according to the physical training-related questions by the intelligent agent;

[0007] According to the search results, answer information of the physical training related questions is obtained.

[0008] According to another aspect of the present disclosure, there is provided a physical training information retrieval device, comprising:

[0009] A receiving module, used for receiving questions related to physical training;

[0010] A retrieval module, configured to search a pre-built physical training knowledge base based on the physical training-related questions through an intelligent agent;

[0011] The acquisition module is used to obtain answer information of the physical training related questions based on the search results.

[0012] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so as to enable the at least one processor to execute the method described in the embodiment of the first aspect.

[0013] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the embodiment of the first aspect.

[0014] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the steps of the method according to the first aspect when executed by a processor.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0017] Figure 1 is a schematic diagram of a physical training information retrieval method provided by an embodiment of the present disclosure;

[0018] Figure 2 is a schematic diagram of another physical training information retrieval method provided by an embodiment of the present disclosure;

[0019] Figure 3 is a schematic diagram of another physical training information retrieval method provided by an embodiment of the present disclosure;

[0020] Figure 4 is a schematic diagram of another physical training information retrieval method provided by an embodiment of the present disclosure;

[0021] Figure 5 This is a logic diagram of an intelligent agent operation provided by an embodiment of the present disclosure;

[0022] Figure 6 It is a structural diagram of an information retrieval device for physical training provided by an embodiment of the present disclosure;

[0023] Figure 7 is a schematic block diagram of an example electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0025] Data processing is the collection, storage, retrieval, processing, transformation and transmission of data. It runs through all areas of social production and social life. Its basic purpose is to extract and derive valuable and meaningful data for certain specific people from large amounts of disorganized and difficult-to-understand data.

[0026] Natural Language Processing (NLP) is an important research direction in the field of artificial intelligence. It integrates knowledge from multiple disciplines such as linguistics, computer science, machine learning, mathematics, and cognitive psychology. It includes two main aspects: natural language understanding and natural language generation. The research content includes multiple levels such as characters, words, phrases, sentences, paragraphs, and chapters. It is a bridge for communication between machine language and human language. It aims to enable machines to understand, interpret, and generate human language, achieve effective communication between humans and machines, and enable computers to perform tasks such as language translation, sentiment analysis, and text summarization.

[0027] Intelligent Search is a new generation of search that combines artificial intelligence. In addition to providing traditional quick search and relevance sorting functions, it also provides user role levels, automatic identification of user interests, semantic understanding of content, intelligent information filtering and push, and other functions.

[0028] An agent is an agent that can perceive the environment and take actions to achieve specific goals. It can be software, hardware, or a system with autonomy, adaptability, and interaction. The agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then performs actions to influence the environment or achieve predetermined goals.

[0029] A large model refers to a machine learning model with a large number of parameters and complex structure. It can process massive amounts of data and complete various complex tasks, such as natural language processing, computer vision, and speech recognition. Large models have the characteristics of large parameters, large training data, and large computing resources. They have the ability to solve general tasks, follow human instructions, and perform complex reasoning.

[0030] The disclosed embodiments can be applied to physical training information retrieval scenarios for people of all stages, including children, teenagers, middle-aged or elderly people; Figure 1 FIG. 1 is a schematic diagram of a method for retrieving information about physical training provided by an embodiment of the present disclosure. Figure 1 As shown, the method includes:

[0031] S101, receive questions related to physical training.

[0032] Optionally, questions related to physical training can be received from a client, and the client can be a port such as an application or a browser where users can perform input operations. In this embodiment, taking physical training for teenagers as an example, questions related to physical training can be "How do teenagers aged 10-14 arrange strength training and aerobic training every month?"

[0033] S102: The intelligent agent searches a pre-built physical training knowledge base based on questions related to physical training.

[0034] Optionally, the intelligent agent may be an intelligent model including matching and retrieval functions; in this embodiment, the intelligent agent includes at least a retrieval module, based on which physical training-related questions are retrieved from a physical training knowledge base.

[0035] In some embodiments, the intelligent agent may perform keyword extraction on the physical training-related questions to obtain a keyword set of the physical training-related questions, and search a pre-built physical training knowledge base based on the keyword set.

[0036] Optionally, the pre-built physical training knowledge base may be a database consisting of a large amount of physical training knowledge. In this embodiment, the physical training knowledge base may include the physical training knowledge of the teenagers that can be collected.

[0037] In some embodiments, keyword matching can be performed in a physical training knowledge base based on a keyword set to obtain relevant knowledge matching the keyword set from the physical training knowledge base, such as background knowledge or popular science knowledge.

[0038] It is understandable that the search results may include two results: relevant knowledge is retrieved from the pre-built physical training knowledge base, or relevant knowledge is not retrieved from the pre-built physical training knowledge base.

[0039] S103: Obtain answer information for questions related to physical training based on the search results.

[0040] Optionally, if the retrieval result is relevant knowledge retrieved from a pre-built physical training knowledge base, answer information for physical training related questions can be obtained based on the retrieved relevant knowledge. For example, question extraction can be performed based on the retrieved relevant knowledge to obtain multiple related similar questions. Corresponding answer information can be searched in the knowledge base based on similar questions, and the corresponding answer information of all similar questions can be summarized or spliced ​​to obtain answer information for physical training related questions, thereby improving the accuracy of question answers.

[0041] Optionally, if the search result shows that no relevant knowledge is found in the pre-built physical training knowledge base, the physical training related questions can be input into the big model, and the big model can reason and output the answer information of the physical training related questions. The big model can be used to answer the physical training related questions, thereby improving the recognition of the user's question intention and improving the efficiency of answering questions.

[0042] In this embodiment, physical training related questions are received, and a pre-constructed physical training knowledge base is searched through the intelligent agent based on the physical training related questions. The search function of the intelligent agent is used to determine whether there is a search result of relevant knowledge matching the physical training related questions in the physical training knowledge base. Based on whether there is relevant knowledge matching the physical training related questions in the search result, answer information of the physical training related questions is obtained. When the search result is that relevant knowledge is retrieved, similar questions are extracted based on the retrieved relevant knowledge, and corresponding answer information is determined based on the similar questions, and then the answer information of the physical training related questions is obtained by fusion, thereby ensuring the accuracy of the answer to the question. When the search result is that no relevant knowledge is retrieved, the large model is called to obtain the answer information of the physical training related questions, thereby improving the recognition of the intention of the user's question and improving the efficiency of answering the question.

[0043] Figure 2 FIG. 1 is a schematic diagram of another method for retrieving information about physical training provided by an embodiment of the present disclosure. Figure 2 As shown, the method includes:

[0044] S201, receiving questions related to physical training.

[0045] In the embodiment of the present disclosure, the implementation method of step S201 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0046] S202: The intelligent agent searches a pre-built physical training knowledge base based on questions related to physical training.

[0047] In the embodiment of the present disclosure, the implementation method of step S202 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0048] S203 : In response to the physical training knowledge base including background knowledge associated with physical training-related issues, determine a question set based on the background knowledge and the physical training issues.

[0049] In some embodiments, if the search result is that background knowledge associated with physical training related questions is retrieved from the physical training knowledge base, a question set is determined based on the background knowledge and the physical training questions.

[0050] Optionally, the questions involved in the background knowledge may be obtained, and all the questions raised in the background knowledge and the physical training question may be combined into a question set.

[0051] S204: performing information retrieval in a physical training knowledge base based on the question set to obtain answer information related to physical training.

[0052] In some embodiments, the questions in the question set can be encoded to obtain a first embedding vector of the question; a second embedding vector of the information pair in the physical training knowledge base is obtained; optionally, encoding can be performed based on an embedding model embeddingmodel, and the embedding vector obtained by encoding each question in the question set is the first embedding vector corresponding to the question; the embedding vector obtained by encoding the information pair in the physical training knowledge base is the second embedding vector, and feature matching is performed based on the embedding vector to improve matching accuracy and efficiency.

[0053] In some embodiments, the physical training knowledge base includes multiple levels of titles and content brief information associated with each title, and each title and the content brief information associated with the title form an information pair.

[0054] Furthermore, the answer information can be obtained based on the first embedding vector and the second embedding vector; optionally, for question i in the question set, the inner product operation is performed on the first embedding vector of question i and the second embedding vector of the information pair in the physical training knowledge base to obtain the inner product set i corresponding to question i, where i is an integer greater than or equal to 1.

[0055] Optionally, the maximum inner product corresponding to question i can be determined from the inner product set i, the target information pair corresponding to the maximum inner product can be determined, and the candidate answer information for question i can be determined based on the target information pair.

[0056] Optionally, the target content set associated with the target information pair can be determined, and the target content set can be determined as candidate answer information for question i; it can be understood that the target information pair is composed of a title and content introduction information associated with the title, so the content corresponding to the title and the content introduction associated with the title can be obtained as the target content set, for example, the content under the title or the content in all subtitles under the title is the target content set, and the target content set is determined to be candidate answer information for question i.

[0057] Furthermore, after obtaining candidate answer information for all questions in the question set, the candidate answer information for the questions in the question set may be fused to obtain answer information, thereby ensuring the accuracy of obtaining the answer information.

[0058] In some embodiments, the candidate answer information for questions in a question set can be sorted according to the maximum inner product corresponding to the question; for example, the candidate answer information for a question can be sorted from large to small according to the size of the inner product, and the candidate answer information can be spliced ​​according to the sorting result to obtain the answer information, making the answer more layered.

[0059] In some embodiments, the order of the questions in the question set may also be determined; candidate answer information of the questions may be spliced ​​in order to obtain answer information, thereby achieving a faster response speed.

[0060] S205 , in response to not obtaining background knowledge from the physical training knowledge base, calling the large language model, and having the large language model answer questions related to physical training based on its own knowledge base to obtain answer information.

[0061] If the search result shows that no background knowledge related to physical training-related questions is found in the physical training knowledge base, the pre-trained large language model is called, and the large language model answers the physical training-related questions based on its own knowledge base to obtain answer information, thereby more fully understanding the user's intention and improving the accuracy of the answer.

[0062] In this embodiment, after receiving questions related to physical training, the intelligent agent searches a pre-constructed physical training knowledge base based on the questions related to physical training to determine whether there is background knowledge of the questions related to physical training in the physical training knowledge base. When background knowledge exists, a question set is obtained based on the background knowledge and the questions related to physical training, and a target information pair is determined from the physical training knowledge base based on the question set. Based on the target information pair, candidate answer information is determined, and the candidate answer information is spliced ​​to obtain the final answer information, thereby improving the accuracy of answers to questions related to physical training; when background knowledge does not exist, a large language model is called to answer to obtain answer information, thereby improving the accuracy of recognizing user intentions and improving the efficiency of obtaining answer information.

[0063] Figure 3 FIG. 1 is a schematic diagram of another method for retrieving information about physical training provided by an embodiment of the present disclosure. Figure 3 As shown, the method includes:

[0064] S301, obtaining reference materials related to physical training.

[0065] In some embodiments, multimodal candidate reference materials may be collected; and the multimodal candidate reference materials may be processed based on a lightweight markup language MD to obtain reference materials.

[0066] Optionally, the multimodal candidate reference materials include but are not limited to books, electronic documents, or pictures, etc. In some embodiments, all paper materials can be scanned and converted into electronic files.

[0067] Furthermore, all electronic documents are converted into markup language Markdown (MD) documents. MD documents are documents written in plain text format and formatted using simple symbols. They have concise syntax and strong cross-platform compatibility.

[0068] Optionally, for the same electronic document, the images therein are saved, and the image reference name is used as the image file name in the image folder corresponding to the electronic document. For example, for electronic document A, the images in electronic document A are saved in image folder A corresponding to electronic document A, and each image is named as the reference name of the image in electronic document A. For example, any image reference name is: Figure 1-1 , then the image in the image folder A is named Figure 1-1 , so as to directly correspond; in this embodiment, the text related to the referenced image is also saved, for example, electronic document A includes "Specific training points refer to Figure 1-1 ", then the text "For specific training points, see Figure 1-1 ” is retained to facilitate image positioning and association.

[0069] After all multimodal candidate reference materials are converted into formats and saved accordingly, they can be used as reference materials and can also be updated and added to, making the search scope wider and more flexible.

[0070] S302, extracting titles and hierarchical relationships of titles from reference materials.

[0071] It is understandable that the reference material may include hierarchical relationships such as first-level titles, second-level titles, and n-level titles, and all titles in the reference material and the hierarchical relationships between titles are extracted.

[0072] S303: Based on the hierarchical relationship of the titles, a collection of content related to the titles is structured and stored to generate a physical training knowledge base.

[0073] In some embodiments, a title and a set of content related to the title may be obtained, where the set of content related to the title refers to the content under the title, and prompt words of a large language model may be generated based on the title and the set of content related to the title.

[0074] In some embodiments, if the content set includes the node content and leaf list associated with the title, the prompt words of the large language model are generated according to the sub-content set associated with the title, node content and leaf list; it can be understood that the elements in the leaf list refer to the sub-titles under the title level, and each sub-title will obtain a corresponding leaf list according to the sub-titles under its own level. If the content set includes the node content and leaf list associated with the title, it means that there are sub-titles of smaller levels under the level to which the title belongs. The prompt words of the large language model are generated according to the node content and leaf list associated with the title, so as to instruct the large language model to generate content introduction information that is more in line with the title based on the comprehensive prompt words.

[0075] In some embodiments, in response to the content collection only including node content associated with the title, prompt words of the large language model are generated based on the title and the node content. That is, when the current title is a title of the smallest level, prompt words of the large language model are generated based on the title and the corresponding node content.

[0076] Furthermore, a large language model is called to generate content summary information associated with the title based on the prompt words, thereby ensuring the accuracy of the generated content summary information.

[0077] Optionally, an information pair may be generated based on the title and the content brief information associated with the title; the information pair may be encoded to obtain a second embedding vector of the information pair; for example, the second embedding vector of the information pair may be obtained by encoding according to an embedding model.

[0078] Based on the hierarchical relationship of the titles, the index information of the titles is determined. In this embodiment, the index information is the path information for locating the current title. The index information is used as the key and the second embedded vector is used as the value for dictionary storage, thereby obtaining a structured storage physical training knowledge base that supports multi-level and fine-grained knowledge retrieval.

[0079] S304, receiving questions related to physical training.

[0080] In the embodiment of the present disclosure, the implementation method of step S304 can be implemented by using any of the methods in the various embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0081] S305: The intelligent agent searches a pre-built physical training knowledge base based on questions related to physical training.

[0082] In the embodiment of the present disclosure, the implementation method of step S305 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0083] S306: Obtain answer information for questions related to physical training based on the search results.

[0084] In the embodiment of the present disclosure, the implementation method of step S306 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0085] In this embodiment, reference materials are obtained by collecting multimodal candidate reference materials and processing them, and the hierarchical relationships between titles in the reference materials are extracted, so as to perform structured storage on the content sets related to the titles, and generate a dictionary with index information as the key and the second embedded vector as the value as the physical training knowledge base. The knowledge base is constructed through the structured dictionary to support multi-level and fine-grained knowledge retrieval, and new reference materials can continue to be added to the physical training knowledge base to adapt to the growing knowledge needs. Question retrieval is performed based on the richer physical training knowledge base to obtain answer information, thereby improving the accuracy of the answer.

[0086] Figure 4 FIG. 1 is a schematic diagram of another method for retrieving information about physical training provided by an embodiment of the present disclosure. Figure 4 As shown, the method includes:

[0087] S401, obtaining reference materials related to physical training.

[0088] In the embodiment of the present disclosure, the implementation method of step S401 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0089] S402, extracting titles and hierarchical relationships of titles from reference materials.

[0090] In the embodiment of the present disclosure, the implementation method of step S402 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0091] S403: Based on the hierarchical relationship of the titles, a collection of content related to the titles is structured and stored to generate a physical training knowledge base.

[0092] In the embodiment of the present disclosure, the implementation method of step S403 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0093] S404, receiving questions related to physical training.

[0094] In the embodiment of the present disclosure, the implementation method of step S404 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0095] S405: The intelligent agent searches a pre-built physical training knowledge base based on questions related to physical training.

[0096] In the embodiment of the present disclosure, the implementation method of step S405 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0097] S406 , in response to the physical training knowledge base including background knowledge associated with physical training-related issues, determine a question set based on the background knowledge and the physical training issues.

[0098] In the embodiment of the present disclosure, the implementation method of step S406 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0099] S407: Search for information in a physical training knowledge base based on the question set to obtain answer information related to physical training.

[0100] In the embodiment of the present disclosure, the implementation method of step S407 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0101] S408: In response to not obtaining background knowledge from the physical training knowledge base, the large language model is called, and the large language model responds to the physical training related questions based on its own knowledge base to obtain response information.

[0102] In the embodiment of the present disclosure, the implementation method of step S408 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.

[0103] In this embodiment, by collecting multimodal candidate reference materials and processing them to obtain reference materials, the titles and hierarchical relationships of the titles in the reference materials are extracted, thereby performing structured storage on the content set related to the titles, generating a dictionary with index information as the key and the second embedded vector as the value, as a physical training knowledge base, constructing a knowledge base through a structured dictionary, supporting multi-level and fine-grained knowledge retrieval, and being able to continue to add new reference materials to the physical training knowledge base to adapt to the growing knowledge needs, after receiving questions related to physical training, the intelligent agent performs pre-constructed search based on the questions related to physical training. A physical training knowledge base is searched to determine whether there is background knowledge of physical training-related issues in the physical training knowledge base. When background knowledge exists, a question set is obtained based on the background knowledge and physical training-related issues, and a target information pair is determined from the physical training knowledge base according to the question set. Based on the target information pair, candidate answer information is determined, and the candidate answer information is spliced ​​to obtain the final answer information, thereby improving the accuracy of answers to physical training-related questions. When background knowledge does not exist, a large language model is called to answer to obtain answer information, thereby improving the accuracy of recognizing user intentions and improving the efficiency of obtaining answer information.

[0104] Based on the above embodiments, Figure 5 This is a logic diagram of the working of an intelligent agent applicable to this embodiment. After the user asks a question, the question module searches for background knowledge based on the user question. If background knowledge exists, a question set is generated, and a retrieval module searches the question set to obtain candidate answer information. All candidate answer information is input into the answer module for splicing, and the answer information is output; if the question module does not have background knowledge based on the user question search, the answer module receives an indication that the background knowledge is empty from the question module, and answers the input user question based on the large language model, and obtains answer information for output, thereby improving the efficiency and accuracy of obtaining answer information.

[0105] Figure 6 This is a schematic diagram of the structure of a physical training information retrieval device provided by an embodiment of the present disclosure. Figure 6 As shown, the physical training information retrieval device 600 includes:

[0106] Receiving module 601, for receiving questions related to physical training;

[0107] A retrieval module 602 is configured to search a pre-built physical training knowledge base based on physical training-related questions through an intelligent agent;

[0108] The acquisition module 603 is used to obtain answer information of questions related to physical training according to the search results.

[0109] In some embodiments, the acquisition module 603 includes:

[0110] In response to the physical training knowledge base including background knowledge associated with physical training related issues, determining a set of issues based on the background knowledge and the physical training issues;

[0111] According to the question set, information is retrieved in the physical training knowledge base to obtain answer information related to physical training.

[0112] In some embodiments, the acquisition module 603 includes:

[0113] Encode the questions in the question set to obtain the first embedding vector of the question;

[0114] Obtaining a second embedding vector of an information pair in a physical training knowledge base, wherein the information pair includes a title and content brief information associated with the title;

[0115] Reply information is obtained according to the first embedding vector and the second embedding vector.

[0116] In some embodiments, the acquisition module 603 includes:

[0117] For problem i in the problem set, perform inner product operations on the first embedding vector of problem i and the second embedding vector of the information pair in the physical training knowledge base to obtain the inner product set i corresponding to problem i, where i is an integer greater than or equal to 1;

[0118] Determine the maximum inner product corresponding to problem i from the inner product set i;

[0119] Determine the target information pair corresponding to the maximum inner product;

[0120] Determine candidate answer information for question i based on target information;

[0121] The candidate answer information of the questions in the question set is integrated to obtain the answer information.

[0122] In some embodiments, the acquisition module 603 includes:

[0123] A target content set associated with the target information pair is determined, and the target content set is determined as candidate answer information for question i.

[0124] In some embodiments, the acquisition module 603 includes:

[0125] Sort the candidate answer information of the questions in the question set according to the maximum inner product corresponding to the question;

[0126] The candidate reply information is spliced ​​according to the sorting results to obtain the reply information.

[0127] In some embodiments, the acquisition module 603 includes:

[0128] Determine the order of questions in the question set;

[0129] The candidate answer information of the question is spliced ​​in order to obtain the answer information.

[0130] In some embodiments, the apparatus 600 further includes:

[0131] In response to the failure to obtain background knowledge from the physical training knowledge base, the large language model is called, and the large language model responds to the physical training-related questions based on its own knowledge base to obtain response information.

[0132] In some embodiments, the retrieval module 602 includes:

[0133] Obtain reference materials related to physical training;

[0134] Extract titles and hierarchical relationships of titles from references;

[0135] According to the hierarchical relationship of the titles, the content collection related to the titles is structured and stored to generate a physical training knowledge base.

[0136] In some embodiments, the retrieval module 602 includes:

[0137] Generate content brief information associated with the title based on the title and the content collection related to the title;

[0138] Generate information pairs based on the title and the content brief information associated with the title;

[0139] Encode the information pair to obtain a second embedding vector of the information pair;

[0140] Determine the index information of the title according to the hierarchical relationship of the title;

[0141] The index information is used as the key and the second embedding vector is used as the value for dictionary storage to obtain a physical training knowledge base.

[0142] In some embodiments, the retrieval module 602 includes:

[0143] Generate prompt words for a large language model based on the title and the collection of content related to the title;

[0144] The large language model is called to generate content summary information associated with the title based on the prompt words.

[0145] In some embodiments, the retrieval module 602 includes:

[0146] In response to the content set including node content and a leaf list associated with the title, generating prompt words of the large language model according to the sub-content set associated with the title, the node content, and the leaf list;

[0147] In response to the content set only including node contents associated with the title, prompt words of a large language model are generated according to the title and the node contents.

[0148] In some embodiments, the retrieval module 602 includes:

[0149] Collect multimodal candidate reference materials;

[0150] Based on the lightweight markup language MD, the multimodal candidate reference materials are processed to obtain reference materials.

[0151] In this embodiment, by collecting multimodal candidate reference materials and processing them to obtain reference materials, the titles and hierarchical relationships of the titles in the reference materials are extracted, thereby performing structured storage on the content set related to the titles, generating a dictionary with index information as the key and the second embedded vector as the value, as a physical training knowledge base, constructing a knowledge base through a structured dictionary, supporting multi-level and fine-grained knowledge retrieval, and being able to continue to add new reference materials to the physical training knowledge base to adapt to the growing knowledge needs, after receiving questions related to physical training, the intelligent agent performs pre-constructed search based on the questions related to physical training. A physical training knowledge base is searched to determine whether there is background knowledge of physical training-related issues in the physical training knowledge base. When background knowledge exists, a question set is obtained based on the background knowledge and physical training-related issues, and a target information pair is determined from the physical training knowledge base according to the question set. Based on the target information pair, candidate answer information is determined, and the candidate answer information is spliced ​​to obtain the final answer information, thereby improving the accuracy of answers to physical training-related questions. When background knowledge does not exist, a large language model is called to answer to obtain answer information, thereby improving the accuracy of recognizing user intentions and improving the efficiency of obtaining answer information.

[0152] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0153] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0154] Figure 7is a schematic block diagram of an example electronic device used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0155] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0156] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0157] The computing unit 701 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the information retrieval method for physical training. For example, in some embodiments, the information retrieval method for physical training can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the information retrieval method for physical training described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the information retrieval method for physical training by any other appropriate means (e.g., by means of firmware).

[0158] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0159] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0160] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0162] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0163] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0164] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0165] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for retrieving information about physical training, wherein: The method comprises: Receive questions related to physical training; Searching a pre-built physical training knowledge base according to the physical training-related questions by the intelligent agent; According to the search results, answer information of the physical training related questions is obtained.

2. The method according to claim 1, wherein The answer information of the physical training related questions is obtained according to the search results, including: In response to the physical training knowledge base including background knowledge associated with the physical training-related problem, determining a problem set based on the background knowledge and the physical training problem; According to the question set, information retrieval is performed in the physical training knowledge base to obtain answer information of the physical training related questions.

3. The method according to claim 2, wherein: The step of searching for information in the physical training knowledge base based on the question set to obtain answer information to the physical training-related questions includes: Encoding the questions in the question set to obtain a first embedding vector of the question; Obtaining a second embedding vector of an information pair in the physical training knowledge base, wherein the information pair includes a title and content brief information associated with the title; The reply information is obtained according to the first embedding vector and the second embedding vector.

4. The method according to claim 3, wherein: The acquiring the reply information according to the first embedding vector and the second embedding vector includes: For problem i in the problem set, perform inner product operations on the first embedding vector of problem i and the second embedding vector of the information pair in the physical training knowledge base to obtain an inner product set i corresponding to problem i, where i is an integer greater than or equal to 1; Determine the maximum inner product corresponding to the problem i from the inner product set i; Determine the target information pair corresponding to the maximum inner product; Determining candidate answer information for the question i based on the target information; The candidate answer information of the questions in the question set is integrated to obtain the answer information.

5. The method according to claim 4, wherein The step of determining candidate answer information for question i based on the target information pair includes: A target content set associated with the target information pair is determined, and the target content set is determined as candidate answer information for the question i.

6. The method according to claim 4, wherein: The step of fusing candidate answer information of questions in the question set to obtain the answer information includes: sorting candidate answer information for questions in the question set according to the maximum inner product corresponding to the question; The candidate reply information is spliced ​​according to the sorting result to obtain the reply information.

7. The method according to claim 4, wherein: The step of fusing candidate answer information of questions in the question set to obtain the answer information includes: determining an order of the questions in the set of questions; The candidate answer information of the question is spliced ​​in the order to obtain the answer information.

8. The method according to claim 2, wherein: The method further comprises: In response to the failure to obtain the background knowledge from the physical training knowledge base, the large language model is called, and the large language model responds to the physical training-related questions based on its own knowledge base to obtain the response information.

9. The method according to any one of claims 1 to 8, wherein The process of constructing the physical training knowledge base includes: Obtain reference materials related to physical training; extracting titles and hierarchical relationships of the titles from the reference materials; According to the hierarchical relationship of the titles, a content set related to the titles is structured and stored to generate the physical training knowledge base.

10. The method according to claim 9, wherein: The step of performing structured storage on the content related to the titles according to the hierarchical relationship of the titles to generate the physical training knowledge base includes: Generate content brief information associated with the title based on the title and a collection of content related to the title; generating an information pair according to the title and the content brief information associated with the title; Encoding the information pair to obtain a second embedding vector of the information pair; Determining index information of the title according to the hierarchical relationship of the title; Dictionary storage is performed using the index information as a key and the second embedded vector as a value to obtain the physical training knowledge base.

11. The method according to claim 9, wherein The step of generating content brief information associated with the title based on the title and a set of content associated with the title includes: Generate prompt words of a large language model according to the title and a set of content related to the title; The large language model is called, and content brief information associated with the title is generated by the large language model according to the prompt word.

12. The method according to claim 11, wherein Generating prompt words of a large language model based on the title and a set of content related to the title includes: In response to the content set including node content and a leaf list associated with a title, generating prompt words of the large language model according to a sub-content set associated with the title, the node content, and the leaf list; In response to the content set only including the node content associated with the title, prompt words of the large language model are generated according to the title and the node content.

13. The method according to claim 9, wherein: The reference materials related to physical training are obtained, including: Collect multimodal candidate reference materials; Based on a lightweight markup language MD, the multimodal candidate reference materials are processed to obtain the reference resources.

14. A physical training information retrieval device comprising: A receiving module, used for receiving questions related to physical training; A retrieval module, configured to search a pre-built physical training knowledge base based on the physical training-related questions through an intelligent agent; The acquisition module is used to obtain answer information of the physical training related questions based on the search results.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-13.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.