Retrieval semantic acquisition method, keyword extraction method, equipment and product
Through a multi-stage processing method, the intent classification, splitting and semantic expansion of natural language text are carried out, which solves the output instability problem of existing semantic retrieval systems in complex and ambiguous queries, and achieves more accurate retrieval results and a more efficient user experience.
Patent Information
- Application Number
- CN202511122423.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-19
AI Technical Summary
Existing semantic retrieval systems have poor output stability and weak fault tolerance when faced with complex and ambiguous natural language queries, making it difficult to meet user query needs.
A multi-stage processing method is adopted to classify the intent of natural language text through a large model, split it into sub-questions, generate retrieval semantics using constraint templates and semantic expansion, and finally extract keywords to improve the accuracy and efficiency of queries.
It achieves accurate splitting and expansion of complex queries, improves the system's output stability and fault tolerance, provides more accurate search results, and improves user experience.
Smart Images

Figure CN120670604A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and in particular to a search semantic acquisition method, a keyword extraction method, a device and a product. Background Art
[0002] Natural language query understanding is a key component of modern intelligent information retrieval systems. Its core task is to convert natural language questions posed by users into machine-executable retrieval semantic representations, thereby enabling efficient and accurate search. As user queries become increasingly complex and structured, traditional keyword-based retrieval methods struggle to meet user demands for accurate, rich, and contextually accurate results.
[0003] Current mainstream semantic retrieval systems generally employ a unified, large-scale model or multi-stage query processing approach. However, these systems still face several key challenges, such as a lack of understanding of intent, the complex logic inherent in a large number of query structures, and a lack of a structured decomposition mechanism, which directly impacts query conversion and result accuracy. In the face of incomplete information or ambiguous expressions, the lack of effective mechanisms for contextual completion and semantic approximation expansion limits recall and coverage, resulting in poor system output stability and fault tolerance, and ultimately unsatisfactory results. Summary of the Invention
[0004] One purpose of this application is to provide a retrieval semantic acquisition method, keyword extraction method, device and product, at least to solve the problem that after the natural language questions raised by the user, the system output has poor stability and weak fault tolerance, and the final output results do not meet the user's query requirements.
[0005] To achieve the above objectives, some embodiments of the present application provide the following aspects:
[0006] In a first aspect, the present application provides a natural language-based search semantics acquisition method, the search semantics acquisition method comprising:
[0007] Get the natural language text entered by the user;
[0008] Performing semantic analysis on the natural language text by using the first model, and classifying the intent of the natural language text;
[0009] Splitting the natural language text into a plurality of sub-questions using a second model, wherein the number of the sub-questions is related to the query complexity of the natural language text;
[0010] Setting different constraint templates according to the intent classification, wherein the constraint templates are used to constrain the final result output;
[0011] The sub-questions are output according to a constraint template format, and the output results are determined as the retrieval semantics of the natural language text.
[0012] Optionally, after the step of splitting the natural language text into multiple sub-questions using the second large model, the retrieval semantics acquisition method further includes:
[0013] Reconstructing the multiple sub-problems through the third model;
[0014] The reconstructed multiple sub-problems are output according to the constraint template.
[0015] Optionally, after performing semantic parsing on the natural language text using the first large model and performing intent classification on the natural language text, the retrieval semantic acquisition method further includes:
[0016] The natural language text is semantically expanded through the fourth model to generate expansion questions related to the natural language text. According to the semantic expansion constraint template format, the expansion questions are determined as the retrieval semantics of the natural language text.
[0017] Optionally, the natural language text is split into a plurality of sub-questions using the second largest model, wherein the number of the split sub-questions is related to the query complexity of the natural language text, specifically including:
[0018] Splitting the natural language text into multiple independent sub-problems through a second large model;
[0019] Optionally, the natural language text is split into multiple interdependent sub-problems through a second large model.
[0020] Optionally, the natural language text is semantically parsed by the first large model, and the natural language text is classified into intent types, wherein the classification types include factual type, comparative type, relational type, recommendation type, and statistical type.
[0021] Optionally, different constraint templates are set according to the intent classification, and the constraint templates are used to constrain the final result output, specifically including:
[0022] When the natural language text classification category is comparative, relational, recommendation, or statistical, the number of sub-questions shall not exceed 5;
[0023] When the natural language text classification category is factual, the sub-question is 0, and the output result is an empty set.
[0024] In a second aspect, the present application further provides a method for extracting search semantic keywords based on natural language, the method comprising:
[0025] Obtain the sub-question output according to any one of the above-mentioned retrieval semantics acquisition methods;
[0026] The fifth model is used to extract keywords for each sub-question and output the keywords.
[0027] Optionally, after the step of extracting keywords for each of the sub-questions using the fifth model and outputting the keywords, the method further includes:
[0028] The keyword is rewritten using the sixth model, and the rewritten keyword is output.
[0029] In a third aspect, the present application further provides an electronic device, comprising:
[0030] one or more processors; and
[0031] A memory storing computer program instructions, wherein the computer program instructions, when executed, cause the processor to perform the steps of any one of the methods described above.
[0032] In a fourth aspect, the present application also provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.
[0033] Compared with the related art, in the solution provided by the embodiment of the present application, questions posed to natural language texts undergo a structured, multi-stage processing method, and the process of constructing retrieval semantics based on user input involves multiple steps such as intent analysis, query reconstruction, logical splitting, and semantic expansion. Through these steps, the system can more accurately understand user needs, eliminate ambiguity and ambiguity in queries, split complex queries, and expand the semantic scope, thereby providing more accurate and comprehensive retrieval results. Automatic conversion from original natural language text questions to high-quality executable retrieval sub-questions is achieved. This process not only enhances the user's search experience, but also provides strong support for the efficient operation of the system, and has strong practical value and industry promotion prospects. This method can be widely used in various scenarios such as question-and-answer systems, intelligent search, intelligent customer service, and knowledge base calls. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0035] Figure 1 A flowchart of a natural language-based search semantics acquisition method provided by an exemplary embodiment of the present disclosure;
[0036] Figure 2 A flowchart of a method for extracting search semantic keywords based on natural language is provided as an exemplary embodiment of the present disclosure;
[0037] Figure 3 An exemplary structural diagram of the electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0038] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0039] Figure 1 A natural language-based search semantics acquisition method is provided as an exemplary embodiment of the present disclosure. The search semantics acquisition method includes:
[0040] S101, obtaining natural language text input by a user;
[0041] S102: Perform semantic analysis on the natural language text using the first model, and perform intent classification on the natural language text;
[0042] Specifically, the first large model performs semantic parsing on the user's natural language query text to identify the intended category. This large model can be a general-purpose large language model (LLM) or a pre-trained large language model, used to accurately classify the intent of natural language text. For example, when a user enters "Beijing's weather," their intent is to find specific information; whereas, when a user asks, "Which is hotter today, Beijing or Shanghai?" their intent is to compare the attributes of two entities.
[0043] S103. Split the natural language text into multiple sub-questions using the second largest model, where the number of sub-questions is related to the query complexity of the natural language text;
[0044] Specifically, after clarifying the user's intent, the next step is logical decomposition, breaking complex queries into multiple simpler sub-questions to support multi-step queries. For example, when a user enters "What are the positions of Zhang San and Li Si?", the system can split the query into two sub-questions: "What is Zhang San's position?" and "What is Li Si's position?" This decomposition not only simplifies the query complexity but also improves query accuracy and efficiency. For more complex queries, such as "Who is the founder of Zhao Si's company?", the system can proceed in two steps: first querying "What companies does Zhao Si own?" and then, based on the results of the first step, further querying the founders of these companies. Through this multi-step query approach, the system can gradually narrow the query scope and ultimately provide an accurate answer. The second-largest model and the first-largest model can refer to the same large model or different large models.
[0045] S104, setting different constraint templates according to the intent classification, wherein the constraint templates are used to constrain the final result output;
[0046] Specifically, a constraint template can be created through a prompt-based mechanism, and the constraint template is used to predefine output formats and rules.
[0047] For example, each sub-question should be expressed in the simplest language possible. Keywords in the original question should not be replaced with synonyms if they are also mentioned in the sub-questions. Sub-questions should be placed in the "sub_questions" list, and the returned format should be a dictionary: {"sub_questions": ["XXX","XXX"...]}, etc.
[0048] S105 . Output the sub-questions according to the constraint template format, and determine the output results as the retrieval semantics of the natural language text.
[0049] Specifically, the final output is determined according to the above constraint template, and the output result is determined as the retrieval semantics of the natural language text.
[0050] For example: Question: Is urban development balanced in the western United States?
[0051] Answer: {"sub_questions": ["What areas does the western United States include?", "What is the urban development situation in the western United States?", "Is urban development in the western United States balanced?"]}
[0052] Question: What are the best books for financial novices to learn the basics of saving and investing?
[0053] Answer: {"sub_questions": ["What are the best books for beginners to learn about saving?", "What are the best books for beginners to learn the basics of investing?"]}
[0054] Question: What is the relationship between funds and government bonds?
[0055] Answer: {"sub_questions": ["What is a fund?", "What is a government bond?", "What is the relationship between funds and government bonds?"]}
[0056] In this embodiment, the process of constructing retrieval semantics based on user input involves multiple steps such as intent analysis, logical splitting, and structured output by processing the user's natural language text questions in a structured, multi-stage manner. It can more accurately understand user needs, eliminate ambiguity in queries, and split complex queries, thereby providing more accurate retrieval results. It realizes the automatic conversion from original natural language text questions to high-quality executable retrieval sub-questions. This process not only improves the user's search experience, but also provides strong support for the efficient operation of subsequent systems. It has strong practicality, and the multi-stage process has error tolerance and fallback mechanisms, supports progressive processing of queries, has good explainability and debugging capabilities, and is suitable for various scenarios such as open domain question answering, vertical search, and government and enterprise knowledge base calls.
[0057] In one embodiment, after the step of splitting the natural language text into multiple sub-questions using the second large model, the retrieval semantics acquisition method further includes:
[0058] Reconstructing the multiple sub-problems through the third model;
[0059] The reconstructed multiple sub-problems are output according to the constraint template.
[0060] The third largest model, the first largest model, and the second largest model may refer to the same large model or different large models. It should be noted that if the third largest model, the first largest model, and the second largest model refer to the same large model, then the large model is a pre-trained model and can be fine-tuned based on the training samples.
[0061] Specifically, after the splitting, the multiple sub-problems are reconstructed through the third model. This step mainly solves the ambiguity, ambiguity and incompleteness that may exist after the logical splitting. Ambiguity is usually reflected in the description of time or entity. For example, when the user inputs "What was the sales volume of XXX new energy vehicles last year, and what was the audience profile?", the system can convert it into "What was the sales volume of XXX new energy vehicles last year" and "What is the audience profile?" through search term splitting. However, the latter loses the original description subject "XXX new energy vehicles", so it needs to be reconstructed into "What is the audience profile of XXX new energy vehicles?".
[0062] In this embodiment, by reconstructing the split sub-problems, the ambiguity, ambiguity and incompleteness that may exist in the sub-problems after logical splitting can be resolved. By introducing an intent-driven query decomposition and reconstruction mechanism, the system can efficiently process complex queries, parallel sentences, nested relationships and other language expression structures, thereby improving the machine parsability of natural language queries.
[0063] In one embodiment, after performing semantic parsing on the natural language text using the first large model and performing intent classification on the natural language text, the retrieval semantic acquisition method further includes:
[0064] The natural language text is semantically expanded through the fourth model to generate expansion questions related to the natural language text. According to the semantic expansion constraint template format, the expansion questions are determined as the retrieval semantics of the natural language text.
[0065] Among them, the fourth largest model, the third largest model, the first largest model, and the second largest model may refer to the same large model or different large models.
[0066] Specifically, semantic expansion, that is, expanding the semantic scope of user queries through synonyms and antonyms, in order to improve the coverage of search results. For example, when a user enters "legal cases involving a collision between two cars", the system can rewrite it into "legal cases involving traffic accidents" through synonym expansion, thereby expanding the search scope and ensuring that relevant results are not missed. In addition, the system can also design a multi-level fallback strategy to gradually relax the query conditions. For example, if the initial query fails to return enough results, the system can gradually relax the query conditions, such as expanding from "legal cases involving a collision between two cars" to "legal cases involving traffic accidents" and then to "legal cases involving vehicle accidents", thereby ensuring that users can obtain enough relevant information.
[0067] Expanded questions can be generated from multiple perspectives, including temporal, spatial, and logical, to generate a series of relevant and meaningful questions. Prompt words can be used to further constrain expanded questions. For example, when expanding, consider the diversity and logical coherence of the questions; ensure that the questions cover a wide range of areas; and maintain relevance to the original query.
[0068] For example:
[0069] When diverging from a spatial perspective, it can be expanded to a larger geographical area or a smaller regional area.
[0070] When diverging from a time perspective, you can look back in history or look forward to the future.
[0071] When you branch out from a logical perspective, you can explore other areas or topics related to the original inquiry.
[0072] When diverging from a comparative perspective, comparative analysis can be performed with other similar entities.
[0073] For example:
[0074] Original query: What is the average GDP growth rate for Denmark?
[0075] Divergent questions: {"extension": ["What is the level of GDP development in Europe?", "What has been Denmark's GDP growth trend over the past 20 years?", "What is Denmark's main economic pillar?"]}
[0076] Original query: What are the applications of artificial intelligence in the medical field?
[0077] Divergent questions: {"extension": ["What are the specific applications of artificial intelligence in disease diagnosis?", "What is the role of artificial intelligence in drug development?", "What is the accuracy of artificial intelligence in medical image analysis?"]}
[0078] Original query: What are the effects of global warming on Arctic ecosystems?
[0079] Divergent questions: {"extension": ["How much impact does global warming have on Arctic glacier melting?", "How do species in Arctic ecosystems adapt to climate change?", "What is the impact of global warming on Arctic marine biodiversity?"]}
[0080] In this embodiment, semantic expansion is performed on the natural language text to ensure that the user can obtain sufficient relevant information, thereby improving the coverage and hit rate of the search results.
[0081] In one embodiment, the natural language text is split into multiple sub-questions using the second largest model, and the number of the sub-questions is related to the query complexity of the natural language text, specifically including:
[0082] Splitting the natural language text into multiple independent sub-problems through a second large model;
[0083] In one embodiment, the natural language text is split into multiple interdependent sub-questions by a second large model.
[0084] Specifically, the second model is used to split the natural language text into multiple sub-problems, which can be roughly divided into:
[0085] First complex query: The main query is a complex query that can be split into multiple independent sub-queries. Each sub-query can be processed separately, and the results are finally integrated to answer the main query.
[0086] For example: Question: What are the best books for financial novices to learn the basics of saving and investing?
[0087] Answer: {"sub_questions": ["What are the best books for financial novices to learn about saving?", "What are the best books for financial novices to learn the basics of investing"]}.
[0088] Secondary complex query: The main query is complex, but needs to be split into multiple interdependent subqueries. These subqueries have dependencies and need to be processed in a specific order to ensure accurate results.
[0089] For example: Question: Is urban development balanced in the western United States?
[0090] Answer: {"sub_questions": ["What areas does the western United States include?", "What is the urban development situation in the western United States?", "Is urban development in the western United States balanced?"]}.
[0091] In this embodiment, the second largest model breaks down natural language text into multiple sub-questions, determines the complexity of each question, and then employs different splitting methods for different questions. This not only simplifies query complexity but also improves query accuracy and efficiency. For more complex queries, this multi-step query approach allows the system to gradually narrow the query scope and ultimately provide accurate answers.
[0092] In one embodiment, the natural language text is semantically parsed by the first large model, and the natural language text is classified into intent categories, wherein the classification types include factual, comparative, relational, recommendation, and statistical.
[0093] Specifically, intent categories include but are not limited to: factual, comparative, relational, recommendation, and statistical. Output intent categories guide subsequent structural processing strategies. For example, when a user enters "Beijing's weather," their intent is to find specific information, which is factual. When a user asks, "Which is hotter today, Beijing or Shanghai?", their intent is to compare the attributes of two entities, which is comparative. And when a user asks, "How much has XX Company's profit increased in the past three years?" their intent is to calculate the attributes of the entity, which is statistical.
[0094] In one embodiment, different constraint templates are set according to the intent classification, and the constraint templates are used to constrain the final result output, specifically including:
[0095] When the natural language text classification category is comparative, relational, recommendation, or statistical, the number of sub-questions shall not exceed 5;
[0096] Specifically, the number of sub-problems that can be decomposed can be limited to less than 5 through the constraint template.
[0097] For example: Question: Is urban development balanced across the western United States? (Relationship)
[0098] Answer: {"sub_questions": ["What areas does the western United States include?", "What is the urban development situation in the western United States?", "Is urban development in the western United States balanced?"]}
[0099] Question: What are the best books for financial management beginners to learn the basics of saving and investing? (Recommended)
[0100] Answer: {"sub_questions": ["What are the best books for beginners to learn about saving?", "What are the best books for beginners to learn the basics of investing?"]}
[0101] Question: What is the relationship between funds and government bonds? (Relationship)
[0102] Answer: {"sub_questions": ["What is a fund?", "What is a government bond?", "What is the relationship between funds and government bonds?"]}
[0103] Question: How much has the profit of XX Company increased in the past three years (statistics)?
[0104] Answer: {"sub_questions": ["What is the net profit of XX Company in 2022?", "What is the net profit of XX Company in 2023?", "What is the net profit of XX Company in 2024?"]}
[0105] When the natural language text classification category is factual, the sub-question is 0, and the output result is an empty set.
[0106] Specifically, factual questions are often answered directly after searching or by searching for answers. These queries typically don't involve subqueries and are relatively straightforward to process. Therefore, when constraining templates, factual questions can be directly output as an empty set without splitting them.
[0107] For example, ask: What does Jingwei filling the sea mean?
[0108] Answer: {"sub_questions": []}
[0109] In the above embodiment, by classifying natural language text, different processing can be performed for different types of questions. For complex queries, parallel sentences, nested relationships and other language expression structures, related sub-question outputs are added, and for simple questions, corresponding sub-question outputs are reduced, which can improve processing efficiency.
[0110] Figure 2 A natural language-based search semantic keyword extraction method is provided as an exemplary embodiment of the present disclosure. The search semantic keyword extraction method includes:
[0111] S201, obtaining a sub-question output according to any one of the above-mentioned retrieval semantics obtaining methods;
[0112] S202. Extract keywords for each of the sub-questions using the fifth model and output the keywords.
[0113] Specifically, key entities are extracted from the sentence, including names of people, companies, places, events, concepts, and other terms, as well as terms describing the entity relationships within the sentence. The fifth, fourth, third, second, and first largest models can refer to the same or different large models.
[0114] Question: What are the top ten buzzwords related to “2025”?
[0115] Answer: {"keywords": ["2025", "hot words"]}
[0116] Question: Has the market value of XX company exceeded that of XXX company?
[0117] Answer: {"keywords": ["XX Company", "Market Value", "XXX Company"]}
[0118] In this embodiment, after obtaining the retrieval semantics, keywords can be further extracted based on the retrieval semantics, and the keywords can be applied to various scenarios such as question-answering systems, intelligent search, intelligent customer service, knowledge base calls, etc., thereby improving the efficiency of retrieval and understanding the user's true intentions.
[0119] In one embodiment, after extracting keywords from each sub-question using the fifth model and outputting the keywords, the method further includes:
[0120] The keyword is rewritten using the sixth model, and the rewritten keyword is output.
[0121] Among them, the sixth largest model, the fifth largest model, the fourth largest model, the third largest model, the second largest model, and the first largest model may refer to the same large model, or different large models, or may be pre-trained models for specific scenarios.
[0122] Specifically, after extracting the keywords, the keywords are rewritten through the sixth model. The rewriting methods include but are not limited to synonym replacement based on the key entity vocabulary in the keyword content, disambiguation rewriting synonym replacement, etc. And the synonym replacement can be further limited by the constraint template. For example: if the sub-question is emphasized by quotation marks or proper nouns, no replacement is performed, and the replacement synonym is required not to change the semantics of the original question. For example: "profit" can be replaced with "income". Or when performing disambiguation rewriting, only the ambiguous words are expanded or replaced, and the complete vocabulary after replacement is required not to change the original semantics. For example: in some sentences, "ant" can be disambiguated and rewritten as "Ant Financial Services Company". If a word does not require synonym replacement or disambiguation rewriting, the original vocabulary is retained in "keywords_rewrite".
[0123] Here are a few examples:
[0124] Question: Is CICC affiliated to BOC International?
[0125] Answer: {"keywords": ["CICC", "BOC International"]}
[0126] {"keywords_rewrite": ["CICC", "BOC International"]}
[0127] Question: What are the top ten buzzwords related to “2025”?
[0128] Answer: {"keywords": ["2025", "hot words"]}
[0129] {"keywords_rewrite": ["2025", "Keywords"]}
[0130] Question: Has the market value of XX company exceeded that of XXX company?
[0131] Answer: {"keywords": ["XX Company", "Market Value", "XXX Company"]}
[0132] {"keywords_rewrite": ["XX Company", "Market Value", "XXX Company"]}
[0133] The following is an overall example to illustrate keyword extraction and keyword rewriting.
[0134] For example: What do I need to prepare to withdraw my deposit in advance? Can I withdraw it without an ID card?
[0135] Output:
[0136] [{'session_id': '111', 'ori_query': 'What do I need to prepare for withdrawing my deposit in advance? Can I withdraw it without an ID card?', 'sub_query': 'What documents do I need to prepare for withdrawing my deposit in advance?', 'keywords': ['Premature withdrawal', 'Deposit', 'Documents'], 'keywords_rewrite': ['Premature withdrawal', 'Deposit', 'Documents']}, {'session_id': '111', 'ori_query': 'What do I need to prepare for withdrawing my deposit in advance? Can I withdraw it without an ID card?', 'sub_query': 'Can I withdraw my deposit in advance without an ID card?', 'keywords': ['ID card', 'Premature withdrawal', 'Deposit'], 'keywords_rewrite': []}]
[0137] In this embodiment, the extracted keywords are further rewritten, and ambiguous words can be expanded or replaced. The rewritten keywords can not only eliminate the ambiguity of the original words, but also match user needs more accurately, thereby significantly improving the accuracy and relevance of information retrieval, avoiding invalid results caused by ambiguity, and achieving more efficient information transmission.
[0138] In addition, some embodiments of the present application further provide an electronic device. The electronic device may be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device may also be various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0139] The electronic device includes: one or more processors; and a memory storing computer program instructions, wherein the computer program instructions, when executed, enable the processor to perform the steps of the method provided in any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, with each device providing some of the necessary operations. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0140] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means, with the bus connection being used as an example in the figure.
[0141] Input device 1103 can receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. Examples include a touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, and other input devices. Output device 1104 may include a display device, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors). The display device may include, but is not limited to, a liquid crystal display, a light emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0142] To provide user interaction, the electronic device may be a computer. The computer includes a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse) through which the user can provide input to the computer. Other types of devices may also be used to provide user interaction; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback), and input from the user may be received in any form (e.g., voice input or tactile input).
[0143] In the embodiments of the present application, a computer program / instruction is stored on a computer-readable medium. When executed by a processor, the computer program / instruction implements the steps of the method provided in any one or more of the above embodiments. The computer-readable medium may be included in the electronic device described in the above embodiments, or it may exist independently and not be incorporated into the device. The computer-readable medium carries one or more computer-readable instructions.
[0144] The memory 1102 can be used as a non-transitory computer-readable storage medium to store non-transitory software programs, non-transitory computer executable programs, and modules. The processor 1101 executes the non-transitory software programs, instructions, and modules stored in the memory 1102 to execute various functional applications and data processing of the server, thereby implementing the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.
[0145] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely located relative to the processor 1101, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0146] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. Computer-readable media may be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.
[0147] Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technology, compact discs, digital versatile discs or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0148] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network or a wide area network, or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0149] In the above embodiments, all or part of the steps or functions of the present invention may be implemented using software, hardware, firmware, or any combination thereof. For example, implementation may be achieved using a dedicated integrated circuit, a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application may be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like. In addition, some steps or functions of the present application may be implemented using hardware, for example, as a circuit that cooperates with a processor to perform the various steps or functions.
[0150] The computer program product provided in the embodiments of the present application includes one or more computer programs / instructions that, when executed by a processor, fully or partially produce the processes or functions described in accordance with the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0151] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-specific system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0152] The scope of this application is defined by the appended claims rather than the foregoing description and is therefore intended to encompass within this application all changes that come within the meaning and range of equivalents of the claims. Any reference signs in the claims should not be construed as limiting the claims to which they relate. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim may also be implemented by one unit or device through software or hardware. Words such as "first" and "second" are only used to distinguish the description and do not indicate any particular order, nor should they be understood as indicating or implying relative importance.
[0153] The above descriptions are merely specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art may easily propose variations or substitutions within the technical scope disclosed in the present application, and such variations or substitutions shall be encompassed within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims, and the above descriptions shall be regarded as exemplary and non-limiting.
Claims
1. A natural language-based retrieval semantics acquisition method, characterized in that: The retrieval semantics acquisition method includes: Get the natural language text entered by the user; Performing semantic analysis on the natural language text by using the first model, and classifying the intent of the natural language text; Splitting the natural language text into a plurality of sub-questions using a second model, wherein the number of the sub-questions is related to the query complexity of the natural language text; Setting different constraint templates according to the intent classification, wherein the constraint templates are used to constrain the final result output; The sub-questions are output according to a constraint template format, and the output results are determined as the retrieval semantics of the natural language text.
2. The search semantics acquisition method according to claim 1, characterized in that: After the step of splitting the natural language text into multiple sub-questions using the second largest model, the retrieval semantics acquisition method further includes: Reconstructing the multiple sub-problems through the third model; The reconstructed multiple sub-problems are output according to the constraint template.
3. The search semantics acquisition method according to claim 1, characterized in that: After the steps of semantically parsing the natural language text using the first large model and classifying the natural language text by intent, the retrieval semantic acquisition method further includes: The natural language text is semantically expanded through the fourth model to generate expansion questions related to the natural language text. According to the semantic expansion constraint template format, the expansion questions are determined as the retrieval semantics of the natural language text.
4. The search semantics acquisition method according to claim 1, characterized in that: The second model is used to split the natural language text into multiple sub-questions. The number of sub-questions is related to the query complexity of the natural language text, specifically including: Splitting the natural language text into multiple independent sub-problems through a second large model; and / or, The natural language text is split into multiple interdependent sub-problems through the second large model.
5. The search semantics acquisition method according to claim 1, characterized in that: The first large model is used to perform semantic analysis on the natural language text and to classify the intent of the natural language text, wherein the classification types include factual, comparative, relational, recommendation, and statistical.
6. The search semantics acquisition method according to claim 5, characterized in that: The different constraint templates are set according to the intent classification, and the constraint templates are used to constrain the final result output, specifically including: When the natural language text classification category is comparative, relational, recommendation, or statistical, the number of sub-questions shall not exceed 5; When the natural language text classification category is factual, the sub-question is 0, and the output result is an empty set.
7. A method for extracting search semantic keywords based on natural language, characterized in that: The search semantic keyword extraction method includes: Obtaining the sub-question output according to the retrieval semantics acquisition method according to any one of claims 1 to 6; The fifth model is used to extract keywords for each sub-question and output the keywords.
8. The search semantic keyword extraction method according to claim 7, characterized in that: After extracting the keywords of each sub-question through the fifth model and outputting the keywords, the method further includes: The keyword is rewritten using the sixth model, and the rewritten keyword is output.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method according to any one of claims 1 to 8.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.