Service corpus generation method and apparatus, and device and computer-readable storage medium
By obtaining the corpus feature vectors of the natural query corpus, determining the standard query corpus and generating business corpus, the problem of low accuracy in the recognition of natural corpus by voice assistants is solved, and a more efficient comprehension ability and recognition accuracy of the voice interaction system is achieved.
Patent Information
- Application Number
- PCT/CN2024/084730
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-03-29
- Publication Date
- 2025-07-03
AI Technical Summary
The existing voice assistants have low accuracy in identifying natural corpus intentions, resulting in insufficient understanding of complex information by voice interaction systems.
By obtaining the corpus feature vectors of the natural query corpus, determining the standard query corpus that meets the target business, and generating business corpus based on the corpus similarity values, including similarity comparison, business category matching and assembly of context corpus information, optimize the preset corpus set to improve recognition accuracy.
It improves the ability of voice interaction systems to understand complex information, enhances the accuracy of recognition of natural corpus intentions, and adapts to real-time matching of various complex application scenarios.
Smart Images

Figure CN2024084730_03072025_PF_FP_ABST
Abstract
Description
Business corpus generation method, device, equipment and computer-readable storage medium
[0001] This application claims priority to Chinese patent application No. 202311804167.9 filed on December 26, 2023, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of natural language processing technology, and in particular to a method, device, electronic device, and computer-readable storage medium for generating business corpus. Background Art
[0003] With the improvement of computer computing power, the field of artificial intelligence has also developed rapidly. In daily life, more and more voice assistants have appeared in various industries, making people's lives and work more convenient.
[0004] Existing voice assistants mainly support voice and text interaction, and are generally used to execute instructions and complete tasks according to instructions. The voice assistant model contains a set instruction language and the execution action corresponding to the instruction. Therefore, users can only communicate with the voice assistant through fixed execution instructions. When the natural language input by the user does not have a corresponding execution instruction, the voice interaction system has poor understanding of complex information and low semantic understanding processing, resulting in poor recognition of natural language and technical problems such as incorrect execution command response. Therefore, the existing voice assistants have low accuracy in recognizing the intent of natural language.
[0005] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Technical issues
[0006] The main purpose of this application is to provide a business corpus generation method, device, electronic device and computer-readable storage medium, aiming to solve the technical problem of low accuracy of existing voice assistants in recognizing the intent of natural corpus. Technical Solutions
[0007] To achieve the above objectives, the present application provides a method for generating business corpus, the method comprising:
[0008] Obtain natural query corpus for target business input;
[0009] Determining a corresponding standard query corpus based on the corpus feature vector of the natural query corpus, wherein the standard query corpus refers to a natural query corpus that meets the query standard of the target business;
[0010] The business corpus of the target business is generated according to the corpus similarity value of the standard query corpus.
[0011] In one embodiment, the step of determining the corresponding standard query corpus based on the corpus feature vector of the natural query corpus includes:
[0012] Comparing the similarity between the corpus feature vector and the target feature vector in a preset corpus to determine a corpus similarity value;
[0013] According to the corpus similarity value, a corresponding standard query corpus is determined in the preset corpus set.
[0014] In one embodiment, the step of determining the corresponding standard query corpus based on the corpus feature vector of the natural query corpus includes:
[0015] Obtaining the business category corresponding to the natural query corpus;
[0016] According to the business category, determining at least one standard query corpus under the business category in the preset corpus set;
[0017] According to the corpus feature vector, a standard query corpus corresponding to the natural query corpus is detected.
[0018] In one embodiment, the step of generating the business corpus of the target business based on the corpus similarity value of the standard query corpus includes:
[0019] When the corpus similarity value is greater than or equal to a preset value, determining a preset number of standard query corpora in the preset corpus set;
[0020] The business corpus is generated according to the standard query corpus.
[0021] In one embodiment, the step of generating the business corpus of the target business based on the corpus similarity value of the standard query corpus includes:
[0022] When the corpus similarity value is greater than or equal to a preset value, requesting the preceding corpus information and / or the following corpus information of the natural query corpus;
[0023] The business corpus is generated by assembling the natural query corpus, the preceding corpus information and / or the following corpus information.
[0024] In one embodiment, after the step of generating the business corpus of the target business based on the corpus similarity value of the standard query corpus, the business corpus generation method further includes:
[0025] Determining a matching ratio of the business corpus in each of the standard query corpora;
[0026] Comparing the matching ratio with a preset ratio;
[0027] When the matching proportion is less than the preset proportion, the preset corpus set is supplemented with corpus according to the natural query corpus.
[0028] In one embodiment, after the step of generating the business corpus of the target business based on the corpus similarity value of the standard query corpus, the business corpus generation method includes:
[0029] Obtaining the business content corresponding to the business corpus;
[0030] The business content is adjusted according to the adjustment operation performed on the business content.
[0031] In addition, to achieve the above-mentioned purpose, the present application also provides a business corpus generation device, the business corpus generation device comprising:
[0032] The acquisition module is used to obtain natural query corpus input for the target business;
[0033] a determination module, configured to determine a corresponding standard query corpus based on a corpus feature vector of the natural query corpus, wherein the standard query corpus refers to a natural query corpus that meets the query standard of the target business;
[0034] A generating module is used to generate the business corpus of the target business according to the corpus similarity value of the standard query corpus.
[0035] In one embodiment, the determining module is further configured to:
[0036] Comparing the similarity between the corpus feature vector and the target feature vector in a preset corpus to determine a corpus similarity value;
[0037] According to the corpus similarity value, the corresponding standard query corpus is determined in the preset corpus set. In one embodiment, the determination module is further configured to:
[0038] Obtaining the business category corresponding to the natural query corpus;
[0039] According to the business category, determining at least one standard query corpus under the business category in the preset corpus set;
[0040] According to the corpus feature vector, the standard query corpus corresponding to the natural query corpus is detected. In one embodiment, the generation module is further configured to:
[0041] When the corpus similarity value is greater than or equal to a preset value, determining a preset number of standard query corpora in the preset corpus set;
[0042] The business corpus is generated according to the standard query corpus.
[0043] In one embodiment, the generating module is further configured to:
[0044] When the corpus similarity value is greater than or equal to a preset value, requesting the preceding corpus information and / or the following corpus information of the natural query corpus;
[0045] The business corpus is generated by assembling the natural query corpus, the preceding corpus information and / or the following corpus information.
[0046] In one embodiment, the business corpus generating device further includes:
[0047] Determining a matching ratio of the business corpus in each of the standard query corpora;
[0048] Comparing the matching ratio with a preset ratio;
[0049] When the matching proportion is less than the preset proportion, the preset corpus set is supplemented with corpus according to the natural query corpus.
[0050] In one embodiment, the business corpus generating device further includes:
[0051] Obtaining the business content corresponding to the business corpus;
[0052] The business content is adjusted according to the adjustment operation performed on the business content.
[0053] The present application also provides an electronic device, which includes: a memory, a processor, and a business corpus generation program stored in the memory and executable on the processor, wherein the business corpus generation program is configured to implement the steps of the above-mentioned business corpus generation method.
[0054] The present application also provides a computer-readable storage medium, on which is stored a program for implementing the business corpus generation method. The program for implementing the business corpus generation method is executed by a processor to implement the steps of the business corpus generation method as described above.
[0055] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned business corpus generation method when executed by a processor. Beneficial effects
[0056] The present application obtains natural query corpus input for the target business, detects the matching results between the natural query corpus and the preset corpus set based on the corpus feature vector of the natural query corpus, and generates corresponding business corpus in the preset corpus set based on the matching results. In this way, the natural corpus information can be identified and the corresponding business information can be matched according to the corpus feature vector of the natural corpus. It can adapt to various complex application scenarios and realize real-time target business matching, thereby improving the voice interaction system's ability to understand complex information and overcoming the technical problem of low recognition efficiency of natural corpus resulting in poor effect of the voice interaction system. Therefore, the recognition accuracy of the natural corpus intention is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0058] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0059] FIG1 is a flow chart of a method for generating business corpus according to an embodiment of the present invention;
[0060] FIG2 is a flowchart of a corpus generation method according to a first embodiment of the present application;
[0061] FIG3 is a schematic diagram of corpus similarity values provided by the business corpus generation method according to the first embodiment of the present application;
[0062] FIG4 is a matching flow chart provided by a method for generating business corpus according to an embodiment of the present application;
[0063] FIG5 is a flow chart of a method for generating business corpus according to a second embodiment of the present application;
[0064] FIG6 is a matching sequence diagram provided by the business corpus generation method according to the second embodiment of the present application;
[0065] FIG7 is a schematic diagram of the module structure of a device for generating business corpus according to a third embodiment of the present application;
[0066] FIG8 is a schematic diagram of the device structure of the hardware operating environment involved in the business corpus generation method in Example 4 of the present application.
[0067] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. Modes for Carrying Out the Invention
[0068] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0069] Example 1
[0070] This application proposes a business corpus generation method according to a first embodiment, referring to FIG1 , which includes:
[0071] Step S10, obtaining a natural query corpus input for the target business;
[0072] Step S20: determining a corresponding standard query corpus based on the corpus feature vector of the natural query corpus, wherein the standard query corpus refers to a natural query corpus that meets the query standard of the target business;
[0073] Step S30: generating the business corpus of the target business according to the corpus similarity value of the standard query corpus.
[0074] In this embodiment, it should be noted that the natural query corpus is used to represent the natural language information input by the user, which can specifically be voice information, text information or image information. The natural query corpus contains the user's input intention. The standard query corpus is used to represent the natural corpus that has been pre-trained and stored in the voice interaction system. The corpus feature vector is used to represent vector values such as word vectors and sentence vectors of the natural query corpus. The vector value can specifically be an embedding vector, where embedding means mapping high-dimensional data (such as text, pictures, music and video, etc.) to a low-dimensional space, and the embedding vector is usually a vector composed of real numbers, which represents the input high-dimensional data with points in a continuous numerical space. At the same time, embedding can be learned based on the appearance pattern of words in the language context, thereby representing the semantics of words or sentences.
[0075] In addition, it should be noted that the corpus similarity value is used to indicate the degree of similarity between the intentions to be expressed by the natural query corpus and the standard query corpus. The business corpus represents the corpus information related to the target business and can be recognized by the machine. Specifically, the larger the corpus similarity value, the more similar the natural query corpus and the standard query corpus are. At the same time, since the standard query corpus corresponds to instruction information and intent content, the business corpus corresponding to the natural query corpus can be determined based on the corpus similarity value.
[0076] For example, to help understand the technical concept or technical principle of the present application, please refer to Figure 2. Figure 2 provides a corpus generation flow chart. When the user inputs natural language information through voice, automatic speech recognition technology (Automatic Speech Recognition, ASR) or other speech recognition technology is used to convert the human voice into text, and then the text information is pre-processed by pre-judgment and error correction and rewriting. The large language model related to natural language understanding (NLU) technology understands the text information and makes intent judgments. Various recommendation strategies and algorithms are used to generate the feedback information that the user wants. The feedback information can be generated through the speech platform composed of natural language generation (NLG) to generate corresponding reply speech, and then passed to the human-computer interaction interface, or it can directly generate execution instructions based on the feedback information for human-computer interaction.
[0077] As an example, steps S10 to S30 include: taking the natural corpus information input by the user as natural query corpus and identifying it, determining the existence of a standard query corpus that has a certain similarity with the natural query corpus by obtaining the corpus feature vector of the natural corpus, and determining the target business of the natural query corpus based on the business corpus corresponding to the standard query corpus.
[0078] The present application obtains natural query corpus, wherein the natural query corpus is the natural language information input by the user, and then compares the natural query corpus with each standard query corpus in the preset corpus set for similarity, determines the corpus similarity value, and matches the business corpus of the natural query corpus based on the corpus similarity value. In this way, the natural corpus information can be identified and matched with the corresponding intent information, thereby improving the voice interaction system's ability to understand complex information and overcoming the technical problem of low recognition efficiency of natural corpus resulting in poor effect of the voice interaction system. Therefore, the recognition accuracy of the natural corpus intent is high.
[0079] The step of determining the corresponding standard query corpus based on the corpus feature vector of the natural query corpus includes:
[0080] Step A10, comparing the corpus feature vector with a target feature vector in a preset corpus set to determine a corpus similarity value;
[0081] Step A20: determining a corresponding standard query corpus in the preset corpus set according to the corpus similarity value.
[0082] In this embodiment, it should be noted that the preset corpus set is used to represent a corpus set composed of standard query corpora, which may specifically include corpus text, instructions corresponding to the corpus text, embedding values corresponding to the corpus text, and business identifiers corresponding to the corpus text (i.e., the business type to which the corpus belongs). The target feature vector is used to represent vector values such as word vectors or sentence vectors of standard query sentences. The corpus similarity value is used to represent the similarity value between the corpus feature vector and the target feature vector, which may specifically be the cosine similarity between the corpus feature vector and the target feature vector, wherein the value range is between -1 and 1. Specifically, when the corpus is similar, the target feature vector is the cosine similarity between the corpus feature vector and the target feature vector. When the similarity value is close to 1, it means that the directions of the corpus feature vector and the target feature vector are very close, and the angle is close to 0 degrees, that is, the natural query corpus is very similar to the standard query corpus. When the corpus similarity value is close to -1, it means that the directions of the corpus feature vector and the target feature vector are opposite, and the angle is close to 180 degrees, that is, the natural query corpus is very dissimilar to the standard query corpus. When the corpus similarity value is close to 0, it means that the directions of the corpus feature vector and the target feature vector are approximately orthogonal, and the angle is close to 90 degrees, that is, the similarity between the natural query corpus and the standard query corpus is low or non-existent.
[0083] In one possible implementation, the corpus similarity value can be obtained by calculating the cosine similarity between the corpus feature vector and the target feature vector. Cosine similarity is a measurement method for comparing the similarity between two vectors. The similarity between the two vectors is evaluated by calculating the cosine value of the angle between the two vectors. Cosine similarity is often used to calculate the similarity between texts in the fields of natural language processing (NLP) and information retrieval.
[0084] For example, to help understand the technical concept or technical principle of the present application, please refer to Figure 3, which provides a schematic diagram of corpus similarity values. In a possible implementation, it is assumed that there are two standard query corpora in the preset corpus set, where standard query corpus 1 is: Help me find customers, the corresponding instruction is filterCustomer, and the generated embedding value mapped to the vector space is vector d1; standard query corpus 2 is: Help me follow up customers, the corresponding instruction is followUp, and the generated embedding value mapped to the vector space is vector d2; the natural query corpus input by the user is: Help me find the customers followed up yesterday, and the generated embedding value mapped to the vector space is vector q. To identify the similarity between the natural query corpus and standard query corpus 1 and standard query corpus 2, we can make a judgment by calculating the cosine value of the angle between the vector q corresponding to the natural query corpus and the vector corresponding to the standard query corpus. Since the smaller the angle between the two vectors, the larger the cosine value, and the higher the cosine similarity, as shown in Figure 3, the angle between vector q and vector d1 is α, and the angle between vector q and vector d2 is θ. Since α<θ, then cosα>cosθ, that is, vector d1 is closer to vector q (the cosine similarity between d1 and q is higher), so standard query corpus 1 is more similar to the description entered by the user.
[0085] As an example, steps A10 to A20 include: calculating the corpus similarity value between the corpus feature vector and the target feature vector in the preset corpus set so that the standard query statement can be matched with the natural query statement, thereby determining the standard query corpus corresponding to the natural query corpus.
[0086] The step of determining the corresponding standard query corpus based on the corpus feature vector of the natural query corpus includes:
[0087] Step B10, obtaining the business category corresponding to the natural query corpus;
[0088] Step B20, determining at least one standard query corpus under the business category in the preset corpus set according to the business category;
[0089] Step B30: detecting the standard query corpus corresponding to the natural query corpus based on the corpus feature vector.
[0090] In this embodiment, it should be noted that the business category is used to indicate the type of business corresponding to the natural query corpus, which may specifically include business types such as those for different industries or different departments in the same industry. Before generating the preset corpus set, the standard query corpus will also be classified according to the business category. Therefore, the business category of the natural query corpus can be determined first, and the corresponding standard query corpus can be directly matched under the business category. In addition, the embedding values corresponding to different corpora are different. Therefore, it is necessary to compare the similarity of the corpus feature vector and each target feature vector to determine the business corpus of the natural query corpus in the preset corpus set. For example, the corpus similarity value can be determined by calculating the cosine similarity between two vectors or the correlation coefficient between two vectors.
[0091] As an example, steps B10 to B30 include: determining the business category to which the natural query corpus belongs, selecting corpus consistent with the business category from the preset corpus set, and then calculating the corresponding corpus similarity value based on the similarity between the corpus feature vector and each target feature vector.
[0092] The step of generating the business corpus of the target business according to the corpus similarity value of the standard query corpus includes:
[0093] Step C10: when the corpus similarity value is greater than or equal to a preset value, determining a preset number of standard query corpora in the preset corpus set;
[0094] Step C20: Generate the business corpus based on the standard query corpus.
[0095] In this embodiment, it should be noted that the preset number of standard query corpora can be specifically 5, 10 or 15, etc. The standard query corpora represent a collection of standard query corpora that have a high similarity with the intent represented by the natural query corpus. After obtaining the corpus similarity value, each standard query corpus can also be sorted from large to small according to the corpus similarity value, and according to the sorting result, a preset number of standard query corpora can be selected from high to low, and the selected standard query corpora can be combined into an intent comparison set. Specifically, the corpus similarity value can be sorted from large to small, and the standard query corpora corresponding to the top ten corpus similarity values can be selected, and then the selected standard query corpora can be combined to supplement the preset large language model.
[0096] As an example, steps C10 to C20 include: obtaining a matching result, and when the matching result is a successful match, selecting one or more standard query corpora with a corpus similarity value greater than or equal to a preset value from the preset corpus set, thereby using the selected standard query corpora as the business corpus.
[0097] The step of generating the business corpus of the target business according to the corpus similarity value of the standard query corpus includes:
[0098] Step D10: when the corpus similarity value is greater than or equal to a preset value, requesting the preceding corpus information and / or the following corpus information of the natural query corpus;
[0099] Step D20 , generating the business corpus by assembling the natural query corpus, the preceding corpus information and / or the following corpus information.
[0100] In this embodiment, it should be noted that the preceding corpus information and the following corpus information are respectively used to represent the natural corpus information sent by the user before and after the current natural query corpus. When generating business corpus for the natural query corpus, it is also necessary to judge the context corpus information of the natural query corpus. If the preceding corpus information and / or the following corpus information exist, it is necessary to determine the business corpus of the natural query corpus based on the matching results of the corpus similarity values, the content of the context corpus information and the preset semantic format. If the preceding corpus information and / or the following corpus information do not exist, the business corpus of the natural query corpus can be directly matched based on the corpus similarity values.
[0101] In one possible implementation, if there is a chat context in the user's current chat, it is necessary to construct the context content according to a preset format, determine the context's corpus intention, and splice it with the user's latest natural corpus information to form a natural query corpus that requires business corpus generation, and then match the business corpus of the natural query corpus according to the preset large language model.
[0102] Exemplarily, to help understand the technical concept or technical principle of the present application, please refer to Figure 4, which provides a matching flow chart. First, by obtaining the current session context and corresponding instructions as well as the natural language content input by the user, and based on the user input content, the top 10 corpus and instructions are screened out, and OpenAI processes the corpus and instructions to parse the user input text. If the parsing is successful, the instruction corresponding to the action can be returned directly. If the parsing fails, it is determined whether the current session context is empty. If so, it continues to determine whether the user input content is related to the current session context. If so, the instruction corresponding to the current session context is returned. If the current context is empty or the user input content is not related to the current session context, it is matched according to other rules in combination with the top 10 corpus instructions, and the matched instructions are returned.
[0103] As an example, steps D10 to D20 include: detecting the contextual corpus information and / or the following corpus information of the natural query corpus. If the contextual corpus information exists, it is necessary to determine whether the identified corpus is related to the contextual corpus information, and combine the corpus similarity value, the contextual corpus information and / or the following corpus information to match the business corpus of the natural query corpus. If the contextual corpus information does not exist, it is necessary to match the business corpus of the natural query corpus based on the corpus similarity value.
[0104] Example 2
[0105] Based on the first embodiment of the present application, in another embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereafter. On this basis, please refer to FIG5 , after the step of generating the business corpus of the target business based on the corpus similarity value of the standard query corpus, the business corpus generation method includes:
[0106] Wherein, after the step of generating the business corpus of the target business according to the corpus similarity value of the standard query corpus, the business corpus generation method includes:
[0107] Step E10, determining the matching ratio of the business corpus in each of the standard query corpora;
[0108] Step E20, comparing the matching ratio with a preset ratio;
[0109] Step E30: When the matching ratio is less than the preset ratio, the preset corpus set is supplemented according to the natural query corpus.
[0110] In this embodiment, it should be noted that the matching ratio is used to indicate the ratio of the business corpus obtained from the standard query corpus that is consistent with the preset intent, and the preset ratio is used to indicate the expected ratio of the business corpus that is consistent with the preset intent in the generated business corpus. Specifically, after obtaining the Top 10 standard query corpora, the matching ratio of the preset intent in the business corpus can also be solved. If the matching ratio is greater than or equal to the preset ratio, it is determined that the accuracy of the matching of the Top 10 corpora is good, and the corpus can be selectively supplemented; if the matching ratio is less than the preset ratio, it is determined that the preset corpus set does not have sufficiently rich corpus to support the target instruction matching of the current user input text, and it is necessary to supplement the corpus with natural language similar to the user input this time.
[0111] In addition, it should be noted that the accuracy of business corpus generation can also be observed periodically through human intervention to determine whether the natural query corpus is consistent with the business corpus. Specifically, by regularly querying the table information generated by the natural query corpus and the business corpus results, it is possible to observe whether the user input content and the matched business corpus are reasonable. When the match is unreasonable, the preset corpus set is supplemented with corpus based on the natural corpus information entered by the user.
[0112] As an example, in one possible implementation, the user enters the natural query "Help me open Zhang San's details page," and the corresponding preset intent is "filterCustomer." The preset intent accounts for 60% of the Top 10 corpus instructions. Assuming that among the Top 10 corpus instructions, 7 are "followUp" and 3 are "filterCustomer," then the preset intent matches 30% of the Top 10 corpus instructions. This is less than the preset percentage, necessitating the addition of the preset corpus set to describe the corpus type of the natural query.
[0113] As another example, the natural language information input by the user and the generated business language results will be stored as a table. The table will be regularly obtained and sent to the person in charge by email. For example, the instruction matching status of all consultants yesterday will be sent at 10 o'clock every morning. The person in charge will observe whether the consultant's expression and the matched instructions are reasonable. If there is any unreasonable situation, it is necessary to use the instruction testing tool to test the natural language information expressed this time to further optimize and supplement the corpus.
[0114] As an example, steps E10 to E30 include: obtaining a preset intent for a natural query corpus, determining a matching ratio between the preset intent and the business corpus corresponding to each standard query corpus in the intent comparison set, comparing the matching ratio with the preset ratio, and when the matching ratio is less than the preset ratio, adding the input natural query corpus to the preset corpus set, and when the matching ratio is greater than or equal to the preset ratio, determining that the corpus matching is successful.
[0115] Wherein, after the step of generating the business corpus of the target business according to the corpus similarity value of the standard query corpus, the business corpus generation method includes:
[0116] Step F10, obtaining the business content corresponding to the business corpus;
[0117] Step F20 : adjusting the service content according to the adjustment operation performed on the service content.
[0118] In this embodiment, it should be noted that the adjustment operation is used to represent the operation performed in the human-computer interaction interface, which can be the user selecting and moving business content in a text-based or graphical user interface through a mouse or touch. The business content represents the instructions corresponding to the business corpus generation and the results generated by the instructions. Specifically, after the generated natural language results are returned to the user, the natural language results can also be selected, keywords highlighted and edited according to the position of the user's cursor.
[0119] In a possible implementation, the human-computer interaction interface is a GUI (Graphical User Interface), and the natural language results are stored in a text box in the GUI. When a keyword is present in the text box, a new element needs to be created, the text data in the text box is copied to the element, and then the highlighted text is wrapped with a specific tag, and Z is set. Axis position, place the element in the lower layer of the text box, and set the text style of the element to make it the same size as the text box, and set the font color to transparent. Finally, for the label of the highlighted element, set the font color to black and set the highlight color to achieve the highlighting of the keyword. In addition, the data content of the text box is detected in real time. When there is a keyword in the text box, it is first necessary to determine the position relationship between the current cursor and the text box, and then detect the position change of the cursor in the document of the entire page. When the cursor position changes, it is determined whether the current text selection contains the keyword based on the cursor position and the selected text. If the cursor has not selected text, it is necessary to determine whether the cursor position is within the keyword. If so, the entire keyword is selected; if the cursor has selected text, check whether the selected text contains the incompletely selected keyword part. If so, update the selected range of the cursor to include the beginning and end of the keyword, and compare the starting position of the cursor with the starting position of the highlighted word, thereby calculating the relative position of the cursor and the highlighted word to determine whether the highlighted word is in a partially selected state.
[0120] As an example, steps F10 to F20 include: outputting the natural language results corresponding to the business corpus, and then implementing text operations such as selecting, editing, and modifying the natural language results by obtaining adjustment operations performed on the business content.
[0121] For example, in order to help understand the technical concept or technical principle of the present application, please refer to Figure 6, which provides a matching sequence diagram. Specifically, when applied in a sales service scenario, when generating business corpus for the natural query corpus input by the user, it needs to go through the mobile sales interactive interface, iFlytek speech-to-text model, mobile sales service instruction generation module, openai text-to-quantity module, and MySQL vector and instruction query module, etc. The user is a sales consultant and sends a voice message on the interactive interface. iFlytek converts the voice message into text message, and then the text message is sent to the instruction service module, which obtains the result based on the text message parsed by openai. The business identifier of the vector information and text information is used to obtain the corresponding instructions in MySQL. The cosine similarity match is performed on the text vector and the vector of the queried corpus. The top 10 are taken as the result set. The top 10 corpus is used to construct a user and system chat example according to the specified format. The chat example constructed above is used to supplement the example in the prompt rule to determine whether there is dialogue context information. If so, the chat dialogue between the user and the system is constructed according to the format, and the user's latest question is spliced to encapsulate the final user chat content. The improved prompt rule and the latest user chat context text are judged and matched according to the prompt rule, and the instruction information is returned.
[0122] This embodiment detects the recognition results and, when the business corpus is inconsistent with the preset intent, supplements the preset corpus set based on the natural query corpus and optimizes the parameter model to improve the accuracy and matching efficiency of business corpus generation. In addition, by obtaining the user's touch operations on the mobile device, adjustments to the business content are achieved, such as quick selection of text or highlighting of the selected text, to facilitate user recognition. This can provide text editing functions such as copying, cutting and pasting, and is compatible with various mobile browsers, thereby meeting the user's needs for quick text editing.
[0123] Example 3
[0124] The present application also provides a business corpus generation device, referring to FIG. 7 , which includes:
[0125] Acquisition module 101, used to acquire natural query corpus input for target business;
[0126] A determination module 102 is configured to determine a corresponding standard query corpus based on the corpus feature vector of the natural query corpus, wherein the standard query corpus refers to a natural query corpus that meets the query standard of the target business;
[0127] The generating module 103 is configured to generate the business corpus of the target business according to the corpus similarity value of the standard query corpus.
[0128] In one embodiment, the determining module 102 is further configured to:
[0129] Comparing the similarity between the corpus feature vector and the target feature vector in a preset corpus to determine a corpus similarity value;
[0130] According to the corpus similarity value, the corresponding standard query corpus is determined in the preset corpus set. In one embodiment, the determination module 102 is further configured to:
[0131] Obtaining the business category corresponding to the natural query corpus;
[0132] According to the business category, determining at least one standard query corpus under the business category in the preset corpus set;
[0133] According to the corpus feature vector, a standard query corpus corresponding to the natural query corpus is detected.
[0134] In one embodiment, the generating module 103 is further configured to:
[0135] When the corpus similarity value is greater than or equal to a preset value, determining a preset number of standard query corpora in the preset corpus set;
[0136] The business corpus is generated according to the standard query corpus.
[0137] In one embodiment, the generating module 103 is further configured to:
[0138] When the corpus similarity value is greater than or equal to a preset value, requesting the preceding corpus information and / or the following corpus information of the natural query corpus;
[0139] The business corpus is generated by assembling the natural query corpus, the preceding corpus information and / or the following corpus information.
[0140] In one embodiment, the business corpus generating device further includes:
[0141] Determining a matching ratio of the business corpus in each of the standard query corpora;
[0142] Comparing the matching ratio with a preset ratio;
[0143] When the matching proportion is less than the preset proportion, the preset corpus set is supplemented with corpus according to the natural query corpus.
[0144] In one embodiment, the business corpus generating device further includes:
[0145] Obtaining the business content corresponding to the business corpus;
[0146] The business content is adjusted according to the adjustment operation performed on the business content.
[0147] The business corpus generation device provided in this application, which adopts the business corpus generation method of the above-mentioned embodiment 1 or embodiment 2, can solve the technical problem of low accuracy in recognizing the intent of natural corpus in existing voice assistants. Compared with the existing technology, the beneficial effects of the business corpus generation device provided in the embodiment of this application are the same as the beneficial effects of the business corpus generation method provided in the above-mentioned embodiment, and the other technical features of the business corpus generation device are the same as the features disclosed in the method of the above-mentioned embodiment, and are not further described here.
[0148] Example 4
[0149] An embodiment of the present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the business corpus generation method in the above-mentioned embodiment one.
[0150] Reference is now made to FIG8 , which illustrates a schematic diagram of the structure of an electronic device suitable for implementing an embodiment of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG8 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.
[0151] As shown in Figure 8 , an electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the electronic device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems may be connected to I / O interface 1006: input devices 1007, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008, such as a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 1003, such as a magnetic tape or hard disk; and communication devices 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.
[0152] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0153] The electronic device provided by this application, using the business corpus generation method in the above-mentioned embodiment, can solve the technical problem of low accuracy in recognizing the intent of natural language corpus in existing voice assistants. Compared with the prior art, the beneficial effects of the electronic device provided by the embodiment of this application are the same as the beneficial effects of the business corpus generation method provided by the above-mentioned embodiment, and the other technical features of the electronic device are the same as those disclosed in the method of the previous embodiment, which will not be repeated here.
[0154] It should be understood that various parts of the present disclosure can be implemented with hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0155] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0156] Example 5
[0157] An embodiment of the present application provides a computer-readable storage medium having computer-readable program instructions stored thereon, and the computer-readable program instructions are used to execute the business corpus generation method in the above-mentioned embodiment 1.
[0158] The computer-readable storage medium provided in the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0159] The computer-readable storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0160] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device: obtains a natural query corpus input for a target business; determines a corresponding standard query corpus based on a corpus feature vector of the natural query corpus, wherein the standard query corpus refers to a natural query corpus that meets the query standard of the target business; and generates a business corpus for the target business based on a corpus similarity value of the standard query corpus.
[0161] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0162] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0163] The modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0164] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions for executing the above-mentioned business corpus generation method, and can solve the technical problem of low accuracy of existing voice assistants in recognizing the intent of natural language corpus. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the embodiments of this application are the same as the beneficial effects of the business corpus generation method provided in the above-mentioned embodiment 1 or embodiment 2, and will not be repeated here.
[0165] Example 6
[0166] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned business corpus generation method.
[0167] The computer program product provided in this application can solve the technical problem of low accuracy in recognizing the intent of natural language content in existing voice assistants. Compared with the existing technology, the beneficial effects of the computer program product provided in the embodiments of this application are the same as the beneficial effects of the business language content generation method provided in the first or second embodiments above, and will not be repeated here.
[0168] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0169] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0170] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0171] The above are merely optional embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for generating business corpus, wherein, The method for generating service corpus includes: Obtaining natural query corpus input for a target service; Determining corresponding standard query corpus according to the corpus feature vector of the natural query corpus, where the standard query corpus refers to the natural query corpus that conforms to the query standard of the target service; Generating service corpus of the target service according to the corpus similarity value of the standard query corpus.
2. The service corpus generation method according to claim 1, wherein, The step of determining corresponding standard query corpus according to the corpus feature vector of the natural query corpus includes: Comparing the corpus feature vector with the target feature vector in a preset corpus set to determine the corpus similarity value; Determining corresponding standard query corpus in the preset corpus set according to the corpus similarity value.
3. The service corpus generation method according to claim 2, wherein, The step of determining corresponding standard query corpus according to the corpus feature vector of the natural query corpus includes: Obtaining the service category corresponding to the natural query corpus; Determining at least one standard query corpus under the service category in the preset corpus set according to the service category; Detecting the standard query corpus corresponding to the natural query corpus according to the corpus feature vector.
4. The service corpus generation method according to claim 2, wherein, The step of generating service corpus of the target service according to the corpus similarity value of the standard query corpus includes: When the corpus similarity value is greater than or equal to a preset value, determining a preset number of standard query corpus in the preset corpus set; Generating the service corpus according to the standard query corpus.
5. The service corpus generation method according to claim 1, wherein, The step of generating service corpus of the target service according to the corpus similarity value of the standard query corpus includes: When the corpus similarity value is greater than or equal to a preset value, requesting the upper context corpus information and / or lower context corpus information of the natural query corpus; Generating the service corpus by assembling the natural query corpus, the upper context corpus information and / or the lower context corpus information.
6. The method for generating service corpus according to claim 4, wherein, After the step of generating service corpus of the target service according to the corpus similarity value of the standard query corpus, the method for generating service corpus further includes: Determining the matching proportion of the service corpus in each standard query corpus; Comparing the matching proportion with a preset proportion; When the matching proportion is less than the preset proportion, supplementing the preset corpus set according to the natural query corpus.
7. The service corpus generation method according to claim 1, wherein, After the step of generating service corpus of the target service according to the corpus similarity value of the standard query corpus, the method for generating service corpus includes: Obtaining the service content corresponding to the service corpus; Adjusting the service content according to the adjustment operation performed on the service content.
8. A business corpus generation device, wherein, The service corpus generation device includes: An obtaining module, configured to obtain natural query corpus input for a target service; A determining module, configured to determine corresponding standard query corpus according to the corpus feature vector of the natural query corpus, where the standard query corpus refers to the natural query corpus that conforms to the query standard of the target service; A generating module, configured to generate service corpus of the target service according to the corpus similarity value of the standard query corpus.
9. An electronic device, wherein, The electronic device includes: a memory, a processor, and a service corpus generation program stored on the memory and executable on the processor, the service corpus generation program being configured to implement the steps of the service corpus generation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, wherein, A service corpus generation program is stored on the computer-readable storage medium, and when the program is executed by a processor, the steps of the service corpus generation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Approximate template matching for natural language queries
CN109478189A
Voice interaction method and device, terminal equipment, readable storage medium and vehicle
CN116312532A
Voice intention recognition method and device, electronic equipment and storage medium
CN116798417A
Service corpus generation method, device and equipment and computer readable storage medium
CN117473069A
In-vehicle circumstantial speech recognition
US20090164216A1