Knowledge base data processing method and server for vehicle-mounted dialogue system

By semantically driven conflict detection and automated cleaning in the knowledge base of the vehicle dialogue system, the problem of knowledge redundancy and command conflict in the knowledge base is solved, and the system performance and user experience are improved.

CN120146168APending Publication Date: 2025-06-13GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510231825.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Due to the diverse sources of knowledge information and untimely updates in the knowledge base of the vehicle dialogue system, the knowledge information may have content conflicts and other problems, which affects the system's semantic understanding ability and user experience.

Method used

By obtaining knowledge points and similar expression sets in the knowledge data, we determine the associated similar expression sets from the knowledge base, conduct semantic-driven conflict detection and automated database cleaning to eliminate redundancy and conflict information.

Benefits of technology

It effectively solves the problem of knowledge redundancy and command conflicts in the knowledge base of the on-board dialogue system, and improves system performance and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146168A_ABST
    Figure CN120146168A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge base data processing method for a vehicle-mounted dialogue system, a server and a computer readable storage medium. The method comprises the following steps: acquiring first knowledge data, wherein the first knowledge data comprises a first knowledge point and a first similar expression set corresponding to the first knowledge point; then, according to the first similar expression set, a first associated similar expression set is determined from the knowledge base, and each first similar expression in the first similar expression set at most corresponds to one first associated similar expression in the first associated similar expression set; and finally, cleaning the knowledge base according to the first similar expression set and the first associated similar expression set. Thus, through semantic-driven knowledge data conflict detection and database automatic cleaning, the problem of knowledge redundancy and instruction conflict in the knowledge base of the vehicle-mounted dialogue system is solved, the performance of the vehicle-mounted dialogue system is improved, and the user experience is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of database data processing, and particularly to a knowledge base data processing method, a server, and a computer-readable storage medium for an in-vehicle dialogue system. Background Art

[0002] In the related art, a pre-configured associated knowledge base is used to provide data information for an in-vehicle dialogue system to enhance the intention understanding ability of the in-vehicle dialogue system. However, due to the acquisition of knowledge information in the knowledge base from different information sources and the non-updated knowledge information in a timely manner, etc., there may be problems such as content conflicts in the knowledge information in the knowledge base, thus affecting the semantic understanding ability of the in-vehicle dialogue system and the user experience. Summary of the Invention

[0003] The present application provides a knowledge base data processing method, a server, and a computer-readable storage medium for an in-vehicle dialogue system.

[0004] An embodiment of the present application provides a knowledge base data processing method for an in-vehicle dialogue system, the method comprising:

[0005] Obtaining first knowledge data, the first knowledge data including a first knowledge point and a first set of similar expressions corresponding to the first knowledge point;

[0006] Determining a first set of associated similar expressions from the knowledge base according to the first set of similar expressions, wherein each first similar expression in the first set of similar expressions corresponds to at most one first associated similar expression in the first set of associated similar expressions;

[0007] Performing a cleaning process on the knowledge base according to the first set of similar expressions and the first set of associated similar expressions.

[0008] In this way, the server obtains first knowledge data, which includes a first knowledge point and a first set of similar expressions corresponding to the first knowledge point. Then, the server determines a first set of associated similar expressions from the knowledge base according to the first set of similar expressions, wherein each first similar expression in the first set of similar expressions corresponds to at most one first associated similar expression in the first set of associated similar expressions. Finally, the server performs a cleaning process on the knowledge base according to the first set of similar expressions and the first set of associated similar expressions. In this way, by detecting conflicts in knowledge data based on semantic driving and automatically cleaning the database, the problems of knowledge redundancy and instruction conflicts in the knowledge base of the in-vehicle dialogue system are solved, the performance of the in-vehicle dialogue system is improved, and the user experience is enhanced.

[0009] In some embodiments, the determining a first set of associated similar expressions from the knowledge base according to the first set of similar expressions includes:

[0010] For each of the first similar expressions, determine a second associated similar expression set from the knowledge base, where the second associated similar expression set includes a plurality of first associated similar expression subsets, and each of the first associated similar expression subsets corresponds to one of the first similar expressions;

[0011] Determine at most one first associated similar expression according to each of the first similar expressions and the corresponding first associated similar expression subset;

[0012] Determine the first associated similar expression set according to the first associated similar expression.

[0013] In this way, the server determines a second associated similar expression set from the knowledge base according to each first similar expression. The second associated similar expression set includes a plurality of first associated similar expression subsets, and each first associated similar expression subset corresponds to one first similar expression. Then, the server determines at most one first associated similar expression according to each first similar expression and the corresponding first associated similar expression subset. Finally, the server determines the first associated similar expression set according to the first associated similar expression. In this way, by matching the first similar expression with the corresponding first associated similar expression subset, the expression with the highest similarity can be accurately determined, thereby improving the accuracy of the matching.

[0014] In some embodiments, the determining, from the knowledge base, a second associated similar expression set according to each of the first similar expressions includes:

[0015] Based on a first preset algorithm, determine a third associated similar expression set from the knowledge base according to each of the first similar expressions, where the third associated similar expression set includes a plurality of second associated similar expression subsets, and each of the second associated similar expression subsets corresponds to one of the first similar expressions;

[0016] Deduplicate each of the second associated similar expression subsets according to the first similar expression to determine a plurality of the first associated similar expression subsets;

[0017] Determine the second associated similar expression set according to the first associated similar expression subset.

[0018] Thus, based on the first preset algorithm, the server determines a third associated similar expression set from the knowledge base according to each of the first similar expressions. The third associated similar expression set includes multiple second associated similar expression subsets, and each of the second associated similar expression subsets corresponds to one of the first similar expressions. Then, the server performs a deduplication process on each of the second associated similar expression subsets according to the first similar expressions to determine multiple first associated similar expression subsets. Finally, the server determines the second associated similar expression set according to the first associated similar expression subsets. In this way, through the deduplication process, the similar associated expressions of the second associated similar expression subsets and the first similar expressions under the same knowledge data can be eliminated, effectively reducing redundant information and improving the efficiency of knowledge base cleaning.

[0019] In some embodiments, the determining of at most one first associated similar expression according to each of the first similar expressions and the corresponding first associated similar expression subset includes:

[0020] Based on the second preset algorithm, according to the first target similar expression and the first target associated similar expression subset corresponding to the first target similar expression, a target associated similar expression table is determined, where the first similar expression includes the first target similar expression, and the first associated similar expression subset includes the first target associated similar expression subset;

[0021] Based on a preset similarity threshold, according to the target associated similar expression table, at most one first target associated similar expression corresponding to the first target similar expression is determined, where the first associated similar expression includes the first target associated similar expression.

[0022] Thus, based on the second preset algorithm, according to the first target similar expression and the first target associated similar expression subset corresponding to the first target similar expression, a target associated similar expression table is determined, where the first similar expression includes the first target similar expression, and the first associated similar expression subset includes the first target associated similar expression subset. Then, based on the preset similarity threshold, the server determines at most one first target associated similar expression corresponding to the first target similar expression according to the target associated similar expression table, where the first associated similar expression includes the first target associated similar expression. In this way, through the target associated similar expression table and the preset similarity threshold, the expression most relevant to the first target similar expression can be accurately identified, thereby improving the accuracy of knowledge base cleaning. Moreover, the preset similarity threshold can avoid misjudging expressions with relatively low similarity as the most relevant expressions, reducing the possibility of misjudgment.

[0023] In some embodiments, the determining of the target associated similar expression table based on the second preset algorithm according to the first target similar expression and the subset of the first target associated similar expressions corresponding to the target first similar expression includes:

[0024] Perform encoding processing on the first target similar expression to determine a first encoding;

[0025] Perform encoding processing on each second similar expression in the subset of the first target associated similar expressions respectively to determine a plurality of second encodings;

[0026] Based on a preset similarity algorithm, calculate the similarity between the first encoding and each of the second encodings respectively to obtain a plurality of similarity values;

[0027] Based on a preset sorting method, perform sorting processing on the plurality of similarity values to obtain a sorting result of the plurality of similarity values;

[0028] According to the sorting result and the subset of the first target associated similar expressions, determine a sub-table of the target associated similar expressions corresponding to the first target similar expression;

[0029] According to the sub-table of the target associated similar expressions, determine the associated similar expression table.

[0030] In this way, perform encoding processing on the first target similar expression to determine a first encoding. Then, the server performs encoding processing on each second similar expression in the subset of the first target associated similar expressions respectively to determine a plurality of second encodings. Subsequently, based on a preset similarity algorithm, the server calculates the similarity between the first encoding and each of the second encodings respectively to obtain a plurality of similarity values. Then, based on a preset sorting method, the server performs sorting processing on the plurality of similarity values to obtain a sorting result of the plurality of similarity values. The server also determines a sub-table of the target associated similar expressions corresponding to the first target similar expression according to the sorting result and the subset of the first target associated similar expressions. Finally, the server determines the associated similar expression table according to the sub-table of the target associated similar expressions. In this way, through encoding processing and a preset similarity algorithm, the similarity between similar questions can be accurately calculated, thereby improving the accuracy of knowledge base cleaning.

[0031] In some embodiments, the cleaning process of the knowledge base according to the first set of similar expressions and the first set of associated similar expressions includes:

[0032] In the case where the first set of associated similar expressions includes the first associated similar expression, determine the knowledge data to be cleaned according to the first set of similar expressions and the first set of associated similar expressions;

[0033] Perform cleaning processing on the knowledge base according to the knowledge data to be cleaned.

[0034] Thus, in the case where the first associated similar expression set includes the first associated similar expressions, the knowledge data to be cleaned is determined according to the first similar expression set and the first associated similar expression set. Then, the server cleans the knowledge base according to the knowledge data to be cleaned. In this way, by cleaning the knowledge base, the conflicting information existing in the knowledge base can be eliminated, the consistency and accuracy of the knowledge base can be ensured, thereby improving the overall quality of the knowledge base and enhancing the user experience.

[0035] In some embodiments, the determining the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set includes:

[0036] Based on the second knowledge points corresponding to each of the first associated similar expressions, a third knowledge data set is determined, and the third knowledge data set includes at least one second knowledge data that matches the first knowledge;

[0037] The knowledge data to be cleaned is determined according to the first knowledge data and the third knowledge data set.

[0038] Thus, based on the second knowledge points corresponding to each first associated similar expression, the server determines a third knowledge data set, and the third knowledge data set includes at least one second knowledge data that matches the first knowledge. Then, the server determines the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set. In this way, by analyzing the association between similar knowledge points, it is possible to accurately determine whether there are conflicts or repetitions in the knowledge data, improving the cleaning accuracy.

[0039] In some embodiments, the determining the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set includes:

[0040] According to the ratio of the first similarity quantity and the second similarity quantity, the matching degree between the second knowledge data and the first knowledge data is determined, where the first similarity quantity is the quantity of the first associated similar expressions in the currently matched second knowledge data, and the second similarity quantity is the quantity of the first similar expressions in the first similar expression set;

[0041] Based on a preset matching degree threshold, the knowledge data to be cleaned is determined according to the matching degree and the third knowledge data set.

[0042] In this way, based on the ratio of the first similar quantity to the second similar quantity, the server determines the matching degree between the second knowledge data and the first knowledge data, where the first similar quantity is the number of first associated similar expressions in the second knowledge data currently being matched, and the second similar quantity is the number of first similar expressions in the first similar expression set. Then, based on a preset matching degree threshold, the server determines the knowledge data to be cleaned according to the matching degree and the third knowledge data set. In this way, by calculating the matching degree and the preset matching degree threshold, the knowledge points conflicting with or redundant to the first knowledge data can be accurately identified, thereby improving the accuracy of knowledge base cleaning. Moreover, the preset matching degree threshold can avoid misjudging the knowledge points with a low matching degree as conflicting or redundant information, reducing the possibility of misjudgment.

[0043] In some embodiments, the cleaning process of the knowledge base according to the first similar expression set and the first associated similar expression set includes:

[0044] In the case that the first associated similar expression set does not include the first associated similar expression, it is confirmed that the first knowledge data does not need to be cleaned.

[0045] In this way, in the case that the first associated similar expression set does not include the first associated similar expression, the server confirms that the first knowledge data does not need to be cleaned. In this way, by judging whether the first associated similar expression set contains the first associated similar expression, it is possible to avoid misjudging the first knowledge data without conflicting or redundant information as the knowledge data to be cleaned.

[0046] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above method is implemented.

[0047] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0048] The additional aspects and advantages of the embodiments of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the embodiments of the present application. Description of the Drawings

[0049] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0050] Figure 1 is one of the flow diagrams of the knowledge base data processing method of some embodiments of the present application;

[0051] Figure 2 It is the second flow diagram of the knowledge base data processing method according to some embodiments of the present application;

[0052] Figure 3 It is the third flow diagram of the knowledge base data processing method according to some embodiments of the present application;

[0053] Figure 4 It is the fourth flow diagram of the knowledge base data processing method according to some embodiments of the present application;

[0054] Figure 5 It is the fifth flow diagram of the knowledge base data processing method according to some embodiments of the present application;

[0055] Figure 6 It is the sixth flow diagram of the knowledge base data processing method according to some embodiments of the present application;

[0056] Figure 7 It is the seventh flow diagram of the knowledge base data processing method according to some embodiments of the present application;

[0057] Figure 8 It is the eighth flow diagram of the knowledge base data processing method according to some embodiments of the present application;

[0058] Figure 9 It is the ninth flow diagram of the knowledge base data processing method according to some embodiments of the present application. Detailed Embodiments

[0059] The following details the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application and should not be construed as a limitation to the embodiments of the present application.

[0060] Currently, in-vehicle dialogue systems can provide a natural and smooth voice interaction experience. Users can control vehicle function points through voice interaction, such as adjusting the air conditioner, navigation, and playing music, etc., thus freeing their hands and being more focused on driving. To ensure the functional reliability and user experience of in-vehicle dialogue systems, it is usually necessary to build a knowledge base including various aspects such as vehicle information, traffic regulations information, road conditions information, and entertainment information, etc., to provide data information related to the user's voice requests for the in-vehicle dialogue system, so as to better understand the user's questions and provide accurate answers.

[0061] However, since the construction process of the knowledge base often involves obtaining knowledge information from multiple different information sources, such as official documents, third-party data interfaces, and Internet information, etc., the knowledge information obtained from different information sources may have different expression forms, update frequencies, and accuracies. In addition, the knowledge information itself may also change over time, such as the update of traffic regulations and the real-time change of road condition information, etc. The above factors may all lead to problems such as content conflicts, outdated information, and redundant information in the knowledge base.

[0062] When a user asks a question to the in-vehicle dialogue system, the system will retrieve relevant information from the knowledge base according to the user's voice request and give an answer. If there are content conflicts in the knowledge base, such as inconsistent answers to the same question, then the answer given by the system may be self-contradictory, thus affecting the user experience. In addition, if there is outdated information in the knowledge base, then the answer given by the system may not reflect the current actual situation, such as inaccurate road condition information, resulting in the user being unable to obtain correct navigation suggestions. Redundant information will cause the system's answer to be long and repetitive, affecting the dialogue efficiency. Therefore, these problems existing in the knowledge base will directly affect the semantic understanding ability of the in-vehicle dialogue system, and further affect the user experience.

[0063] Based on the above problems, please refer to Figure 1 , an embodiment of the present application provides a method for processing knowledge base data for an in-vehicle dialogue system, the method includes:

[0064] 01: Obtain first knowledge data, the first knowledge data includes a first knowledge point and a first set of similar expressions corresponding to the first knowledge point;

[0065] 02: Determine a first associated set of similar expressions from the knowledge base according to the first set of similar expressions;

[0066] 03: Clean the knowledge base according to the first set of similar expressions and the first associated set of similar expressions.

[0067] An embodiment of the present application also provides a server, including a memory and a processor. The method for processing knowledge base data for an in-vehicle dialogue system according to the embodiment of the present application can be implemented by the server according to the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain first knowledge data, the first knowledge data includes a first knowledge point and a first set of similar expressions corresponding to the first knowledge point. And determine a first associated set of similar expressions from the knowledge base according to the first set of similar expressions. And clean the knowledge base according to the first set of similar expressions and the first associated set of similar expressions.

[0068] The embodiments of the present application also provide a data processing device. The method for processing knowledge base data for an in-vehicle dialogue system according to the embodiments of the present application can be implemented by the data processing device according to the embodiments of the present application. Specifically, the data processing device includes an acquisition module, a determination module, and a data cleaning module. The acquisition module is used to acquire first knowledge data, and the first knowledge data includes a first knowledge point and a first set of similar expressions corresponding to the first knowledge point. The determination module is used to determine a first set of associated similar expressions from the knowledge base according to the first set of similar expressions. The data cleaning module is used to perform a cleaning process on the knowledge base according to the first set of similar expressions and the first set of associated similar expressions.

[0069] Specifically, an in-vehicle dialogue system, also known as an in-vehicle voice interaction system, refers to a system integrated in a vehicle that enables voice interaction between a person and the vehicle through technologies such as speech recognition, speech synthesis, and natural language processing, and can improve driving safety, driving convenience, provide personalized services, and enhance the driving experience.

[0070] The knowledge base of the in-vehicle dialogue system refers to a database stored in the in-vehicle dialogue system that is used to support the system in understanding the user's intention and providing relevant answers. The knowledge base includes various driving-related information, such as vehicle information, road condition information, traffic regulations, and entertainment information. In some embodiments, the knowledge base of the in-vehicle dialogue system can be obtained through manual construction or automatic construction. Manual construction means collecting and organizing information manually and entering it into the knowledge base. However, this method requires a large amount of manpower and material resources and is difficult to ensure timely update of information. Automatic construction means obtaining information from the Internet through relevant technologies and using natural language processing technologies to extract and structure the information and store it in the knowledge base. However, this method needs to solve problems such as information redundancy and information conflict. The method for processing knowledge base data for an in-vehicle dialogue system provided by the embodiments of the present application is aimed at the knowledge base of the automatically constructed in-vehicle dialogue system. By performing a data cleaning process on the automatically constructed knowledge base, the problems of knowledge redundancy and instruction conflict in the knowledge base of the in-vehicle dialogue system are solved, the performance of the in-vehicle dialogue system is improved, and the user experience is enhanced.

[0071] In some embodiments, the knowledge base consists of knowledge points, answers corresponding to the knowledge points, and similar questions. A knowledge point is the basic unit of the knowledge base, which refers to the voice requests that users may make, such as "How to set the navigation?", "What's the weather like today?", etc. Each knowledge point contains multiple similar questions, which are various expressions that users may use to enhance the generalization ability of the dialogue. A similar question refers to other knowledge points with similar semantics to the knowledge point, such as "How to set the GPS?" and "What's the weather forecast today?", etc. An answer is the response or operation executed by the in-vehicle dialogue system to the user's voice request, such as "Open the navigation application, click on the destination, and enter the destination address." and "It's sunny today, and the temperature is 20 degrees.", etc.

[0072] The first knowledge data refers to the knowledge data being detected in the current knowledge base. Through certain processing, it is determined whether the first knowledge data is repeated or in information conflict with other knowledge data. Please refer to Table 1, which shows part of the data in the knowledge base. Among them, the knowledge data includes knowledge points, similar questions corresponding to the knowledge points, and answers. A similar question is a similar expression, and a set of similar questions is a set of similar expressions. If the first knowledge data being detected currently is the knowledge data where knowledge point A is located, then the first knowledge point is knowledge point A, and the first set of similar expressions is similar questions A1 - 7.

[0073] Table 1

[0074]

[0075] The first associated similar expression set refers to the intermediate result of cross-knowledge point detection. It is a set of potential conflict entries that are screened by semantic similarity, have a high semantic similarity with the first similar expressions of the first knowledge point, and cross knowledge points, providing a basis for further knowledge base cleaning operations. That is, the set of first associated similar expressions that finally match each of the first similar expressions in the first similar expression set. It should be noted that each first similar expression in the first similar expression set corresponds to at most one first associated similar expression in the first associated similar expression set. That is, one similar expression has at most one associated similar expression. Please refer to Table 1 again. If the first acquaintance data is the knowledge data where the knowledge point is A, after some processing, the first associated similar expression corresponding to the similar question A1 is the similar question D1, the first associated similar expression corresponding to the similar question A2 is the similar question E2, the first associated similar expression corresponding to the similar question A3 is the similar question B3, the first associated similar expression corresponding to the similar question A4 is the similar question C4, the first associated similar expression corresponding to the similar question A5 is the similar question B5, the first associated similar expression corresponding to the similar question A6 is the similar question C6, and there is no first associated similar expression corresponding to the similar question A7. Then the first associated similar expression set is the set of the similar questions D1, E2, B3, C4, B5, and C6.

[0076] Obtain the first knowledge data from the knowledge base, including the first knowledge point and the first similar expression set corresponding to this knowledge point.

[0077] Next, according to the first similar expression set, determine the first associated similar expression set from the knowledge base. Each first similar expression corresponds to at most one first associated similar expression.

[0078] Finally, based on the first similar expression set and the first associated similar expression set, perform cleaning processing on the knowledge base to eliminate or merge conflicting information, ensuring the consistency and accuracy of the knowledge base. In this way, by cleaning the conflicting information, it is ensured that the information corresponding to each knowledge point in the knowledge base is consistent, avoiding contradictory answers.

[0079] The following uses an example to illustrate the knowledge base data processing method for an in-vehicle dialogue system provided by the embodiments of the present application. Referring to Table 1 again, the first knowledge data obtained is the knowledge data D where the knowledge point D is located, the first knowledge point is the knowledge point D, and the first set of similar expressions is the similar questions D1-3. Then, according to the obtained set of similar questions corresponding to the knowledge point D, it is determined that the first associated set of similar expressions is the set of similar questions B1 and similar questions B2. Then, according to the first set of similar expressions and the first associated set of similar expressions, the knowledge base is cleaned to eliminate or merge conflicting information, ensuring the consistency and accuracy of the knowledge base. Subsequently, after the processing of the knowledge point A and the corresponding set of similar questions and answers is completed, the other knowledge data in the database is processed until all the knowledge data in the database is completed.

[0080] In summary, in the knowledge base data processing method and server for an in-vehicle dialogue system provided by the embodiments of the present application, the server obtains the first knowledge data, and the first knowledge data includes the first knowledge point and the first set of similar expressions corresponding to the first knowledge point. Then, the server determines the first associated set of similar expressions from the knowledge base according to the first set of similar expressions, where each first similar expression in the first set of similar expressions corresponds to at most one first associated similar expression in the first associated set of similar expressions. Finally, the server cleans the knowledge base according to the first set of similar expressions and the first associated set of similar expressions. In this way, through the conflict detection of knowledge data driven by semantics and the automatic cleaning of the database, the problems of knowledge redundancy and instruction conflict in the knowledge base of the in-vehicle dialogue system are solved, the performance of the in-vehicle dialogue system is improved, and the user experience is enhanced.

[0081] Please refer to Figure 2 , in some embodiments, step 02 (determining the first associated set of similar expressions from the knowledge base according to the first set of similar expressions) includes:

[0082] 021: Determine the second associated set of similar expressions from the knowledge base according to each first similar expression;

[0083] 022: Determine at most one first associated similar expression according to each first similar expression and the corresponding subset of first associated similar expressions;

[0084] 023: Determine the first associated set of similar expressions according to the first associated similar expressions.

[0085] In some embodiments, the determining module is further configured to determine the second associated set of similar expressions from the knowledge base according to each first similar expression. And determine the second associated set of similar expressions from the knowledge base according to each first similar expression. And determine the first associated set of similar expressions according to the first associated similar expressions.

[0086] In some embodiments, the processor is further configured to determine a second associated similar expression set from the knowledge base according to each first similar expression. And determine a second associated similar expression set from the knowledge base according to each first similar expression. And determine a first associated similar expression set according to the first associated similar expressions.

[0087] Specifically, the second associated similar expression set refers to the set of similar questions that are closest to the similar questions of all other knowledge points under each knowledge data, that is, a similar question can retrieve a set of associated similar expressions corresponding to it (i.e., the first associated similar expression subset). Gathering the multiple sets of associated similar expressions corresponding to multiple similar questions together is the second associated similar expression set. Continuing with the above example, please refer to Table 1 again. If the first knowledge data is the knowledge data D where the knowledge point D is located, the first associated similar expression subset D-1 corresponding to the similar question D1 is the similar questions B1, E1, A1, and C1. The first associated similar expression subset D-2 corresponding to the similar question D2 is the similar questions B2, E2, C2, and A2. The first associated similar expression subset D-3 corresponding to the similar question D3 is the similar questions A3, B3, E3, and C3. In this way, the second associated similar expression set corresponding to the knowledge data D includes the first associated similar expression subset D-1, the first associated similar expression subset D-2, and the first associated similar expression subset D-3.

[0088] The first associated similar expression refers to the similar expression that is most similar to the first similar expression determined from multiple associated similar expressions in the first associated similar expression subset based on certain rules. Please refer to Table 1 again. Continuing with the above example, the first associated similar expression corresponding to the similar question D1 determined from the first associated similar expression subset D-1 is the similar question B1. The first associated similar expression corresponding to the similar question D3 determined from the first associated similar expression subset D-3 is the similar question A3. Due to rule restrictions, there is no corresponding first associated similar expression for the similar question D3.

[0089] Please refer to Table 1 again. Continuing with the above example, according to all the first associated similar expressions, the first associated similar expression set is determined to be the set of the similar questions B1 and B2.

[0090] In this way, the server determines a second associated similar expression set from the knowledge base according to each first similar expression. The second associated similar expression set includes multiple first associated similar expression subsets, and each first associated similar expression subset corresponds to a first similar expression. Then, the server determines at most one first associated similar expression according to each first similar expression and the corresponding first associated similar expression subset. Finally, the server determines a first associated similar expression set according to the first associated similar expressions. In this way, by matching the first similar expressions with the corresponding first associated similar expression subsets, the expression with the highest similarity can be accurately determined, thereby improving the accuracy of the matching.

[0091] Please refer to Figure 3 , in some embodiments, step 021 (determining a second associated similar expression set from the knowledge base according to each first similar expression) includes:

[0092] 0211: Based on a first preset algorithm, determine a third associated similar expression set from the knowledge base according to each first similar expression;

[0093] 0212: Remove duplicates from each second associated similar expression subset according to the first similar expression to determine multiple first associated similar expression subsets;

[0094] 0213: Determine a second associated similar expression set according to the first associated similar expression subsets.

[0095] In some embodiments, the determining module is further configured to determine a third associated similar expression set from the knowledge base according to each first similar expression based on a first preset algorithm. And remove duplicates from each second associated similar expression subset according to the first similar expression to determine multiple first associated similar expression subsets. And determine a second associated similar expression set according to the first associated similar expression subsets.

[0096] In some embodiments, the processor is further configured to determine a third associated similar expression set from the knowledge base according to each first similar expression based on a first preset algorithm. And remove duplicates from each second associated similar expression subset according to the first similar expression to determine multiple first associated similar expression subsets. And determine a second associated similar expression set according to the first associated similar expression subsets.

[0097] Specifically, the first preset algorithm refers to a text retrieval algorithm based on probability theory, which can be used to estimate the relevance between a document and a query text. That is, according to factors such as the frequency of occurrence of query terms in the document and the document length, a relevance score between the document and the query text is calculated. The higher the score, the more relevant the document is to the query text. In some embodiments, the first preset algorithm includes the BM25 algorithm, the TF-IDF algorithm, and the vector space model. Among them, the BM25 algorithm can retrieve the most relevant documents from a large number of documents according to the query terms input by the user. The TF-IDF algorithm is a text representation method based on term frequency-inverse document frequency, which can be used to represent the relevance between a document and a query. The vector space model is a text representation method that represents a document as a vector, which can be used to calculate the similarity between a document and a query.

[0098] The third set of associated similar expressions refers to the set of the most similar questions among the similar questions of each knowledge data and all knowledge points in the knowledge base. That is, a similar question can retrieve a set of associated similar expressions corresponding to it (i.e., the second subset of associated similar expressions). Gathering the multiple sets of associated similar expressions corresponding to multiple similar questions together is the third set of associated similar expressions. That is to say, the difference between the third set of associated similar expressions and the second subset of associated similar expressions is that the third set of associated similar expressions compares the first similar expression to be detected with all similar questions in the knowledge base, and it may also include the similar questions in the knowledge data where the similar question to be detected is located. While the second subset of associated similar expressions is the set of the most similar questions to the similar questions in other knowledge data, excluding the similar questions in the knowledge data where the similar question to be detected is located. Continuing with the above example, please refer to Table 1 again. If the first knowledge data is the knowledge data D where the knowledge point D is located, the second subset of associated similar expressions DD-1 corresponding to the similar question D1 is the similar questions B1, D3, E1, A1, and C1. The second subset of associated similar expressions DD-2 corresponding to the similar question D2 is the similar questions B2, D1, E2, C2, and A2. The second subset of associated similar expressions DD-3 corresponding to the similar question D3 is the similar questions D1, A3, B3, E3, and C3. In this way, the third set of associated similar expressions corresponding to the knowledge data D includes the second subset of associated similar expressions DD-1, the second subset of associated similar expressions DD-2, and the second subset of associated similar expressions DD-3.

[0099] The duplicate removal process refers to eliminating the similar expressions in each second associated similar expression subset that are under the same knowledge data as the similar expression to be detected. Continuing with the above example, please refer to Table 1 again. Eliminate the similar question D3 in the second associated similar expression subset DD-1 to obtain similar questions B1, E1, A1, and C1, which is the first associated similar expression subset D-1. Eliminate the similar question D1 in the second associated similar expression subset DD-2 to obtain similar questions B2, E2, C2, and A2, which is the first associated similar expression subset D-2. Eliminate the similar question D1 in the second associated similar expression subset DD-3 to obtain similar questions A3, B3, E3, and C3, which is the first associated similar expression subset D-3.

[0100] Continuing with the above example, please refer to Table 1 again. Based on the first preset algorithm, according to the first similar expression D1, determine from the knowledge base that the third associated similar expression set includes the second associated similar expression subset DD-1, the second associated similar expression subset DD-2, and the second associated similar expression subset DD-3. Among them, the second associated similar expression subset DD-1 includes similar questions B1, D3, E1, A1, and C1, the second associated similar expression subset DD-2 includes similar questions B2, D1, E2, C2, and A2, and the second associated similar expression subset DD-3 includes similar questions D1, A3, B3, E3, and C3.

[0101] Next, the server performs duplicate removal processing on each second associated similar expression subset according to the first similar expression to determine multiple first associated similar expression subsets. That is, eliminate the similar question D3 in the second associated similar expression subset DD-1 to obtain similar questions B1, E1, A1, and C1, which is the first associated similar expression subset D-1. Eliminate the similar question D1 in the second associated similar expression subset DD-2 to obtain similar questions B2, E2, C2, and A2, which is the first associated similar expression subset D-2. Eliminate the similar question D1 in the second associated similar expression subset DD-3 to obtain similar questions A3, B3, E3, and C3, which is the first associated similar expression subset D-3.

[0102] Finally, determine the second associated similar expression set according to the first associated similar expression subset. The second associated similar expression set includes the first associated similar expression subset D-1, the first associated similar expression subset D-2, and the first associated similar expression subset D-3.

[0103] In this way, based on the first preset algorithm, the server determines a third associated similar expression set from the knowledge base according to each first similar expression. The third associated similar expression set includes multiple second associated similar expression subsets, and each second associated similar expression subset corresponds to a first similar expression. Then, the server performs deduplication processing on each second associated similar expression subset according to the first similar expression to determine multiple first associated similar expression subsets. Finally, the server determines the second associated similar expression set according to the first associated similar expression subsets. In this way, through the deduplication processing, the similar associated expressions of the second associated similar expression subsets and the first similar expression under the same knowledge data can be eliminated, effectively reducing redundant information and improving the efficiency of knowledge base cleaning.

[0104] Please refer to Figure 4 , in some embodiments, step 022 (determining at most one first associated similar expression according to each first similar expression and the corresponding first associated similar expression subset) includes:

[0105] 0221: Based on the second preset algorithm, according to the first target similar expression and the first target associated similar expression subset corresponding to the first target similar expression, determine the target associated similar expression table;

[0106] 0222: Based on the preset similarity threshold, according to the target associated similar expression table, determine at most one first target associated similar expression corresponding to the first target similar expression.

[0107] In some embodiments, the determining module is further configured to determine the target associated similar expression table based on the second preset algorithm according to the first target similar expression and the first target associated similar expression subset corresponding to the first target similar expression. The determining module is further configured to determine at most one first target associated similar expression corresponding to the first target similar expression according to the target associated similar expression table based on the preset similarity threshold.

[0108] In some embodiments, the processor is further configured to determine the target associated similar expression table based on the second preset algorithm according to the first target similar expression and the first target associated similar expression subset corresponding to the first target similar expression. And determine at most one first target associated similar expression corresponding to the first target similar expression according to the target associated similar expression table based on the preset similarity threshold.

[0109] Specifically, the second preset algorithm refers to an algorithm that determines the similarity degree between texts by calculating the similarity between text vectors, such as the BERT vectorization algorithm. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model that can effectively capture the semantic information of texts and generate high-quality text vector representations through learning a large amount of text data.

[0110] The first target similar expression refers to the similar expression currently being processed among all the first similar expressions. The first target associated similar expression subset refers to the first associated similar expression subset corresponding to the first target similar expression among all the first associated similar expression subsets.

[0111] The target associated similar expression table includes the most similar expression and the corresponding similarity value in the first target associated similar expression subset to the first target similar expression. Continuing with the above example, please refer to Table 1 again. Based on the second preset algorithm, according to the first target similar expression and the first target associated similar expression subset corresponding to the target first similar expression, the determined target associated similar expression table is "Similar question B1: 0.95; Similar question B2: 0.93; Similar question A3: 0.87".

[0112] The preset similarity threshold refers to the set similarity score threshold used to determine whether the similarity is high enough. In the implementation manner of this application, 0.90 is used as the preset similarity threshold.

[0113] Based on the second preset algorithm, according to the first target similar expression and the first target associated similar expression subset corresponding to the first target similar expression, determine the target associated similar expression table, where the first similar expression includes the first target similar expression, and the first associated similar expression subset includes the first target associated similar expression subset. Continuing with the above example, please refer to Table 1 again. When the first target similar expression is the first similar expression D1, that is, when currently processing the first similar expression D1, calculate the similarity value of each similar expression in the first target associated similar expression subset D-1 corresponding to the first similar expression D1 respectively. The similarity value of similar question B1 is 0.95, the similarity value of similar question E1 is 0.91, the similarity value of similar question A1 is 0.83, and the similarity value of similar question C1 is 0.79, so as to determine the target associated similar expression table as "Similar question B1: 0.95".

[0114] Next, based on a preset similarity threshold, the server determines at most one first target-associated similar expression corresponding to the first target similar expression according to the target-associated similar expression table, where the first associated similar expression includes the first target-associated similar expression. Continuing with the above example, please refer to Table 1 again. Since 0.95 is greater than 0.90, the first target-associated similar expression corresponding to the first similar expression D1 is the similar question B1.

[0115] Similarly, when the first target similar expression is the first similar expression D2, it can be determined that the target-associated similar expression table is "similar question B2: 0.93". Subsequently, since 0.93 is greater than 0.90, the first target-associated similar expression corresponding to the first similar expression D2 is the similar question B2. This will not be elaborated here.

[0116] When the first target similar expression is the first similar expression D3, the similarity values of each similar expression in the first target-associated similar expression subset D-1 corresponding to the first similar expression D1 are calculated respectively. The similarity value of the similar question A3 is 0.87, the similarity value of the similar question B3 is 0.82, the similarity value of the similar question E3 is 0.81, and the similarity value of the similar question C is 0.79. Thus, it is determined that the target-associated similar expression table is "similar question A3: 0.87". Subsequently, since 0.87 is less than 0.90, there is no first target-associated similar expression corresponding to the first similar expression D2.

[0117] In this way, based on the second preset algorithm, according to the first target similar expression and the first target-associated similar expression subset corresponding to the first target similar expression, the target-associated similar expression table is determined, where the first similar expression includes the first target similar expression, and the first associated similar expression subset includes the first target-associated similar expression subset. Next, based on the preset similarity threshold, the server determines at most one first target-associated similar expression corresponding to the first target similar expression according to the target-associated similar expression table, where the first associated similar expression includes the first target-associated similar expression. In this way, through the target-associated similar expression table and the preset similarity threshold, the expression most relevant to the first target similar expression can be accurately identified, thereby improving the accuracy of knowledge base cleaning. Moreover, the preset similarity threshold can avoid misjudging expressions with relatively low similarity as the most relevant expressions, reducing the possibility of misjudgment.

[0118] Please refer to Figure 5 , in some embodiments, step 0221 (based on the second preset algorithm, according to the first target similar expression and the first target-associated similar expression subset corresponding to the first target similar expression, determine the target-associated similar expression table) includes:

[0119] 02211: Perform encoding processing on the first target similar expression to determine the first code;

[0120] 02212: Code each second similar expression in the first target-associated similar expression subset respectively to determine a plurality of second encodings;

[0121] 02213: Based on a preset similarity algorithm, calculate the similarity between the first encoding and each second encoding respectively to obtain a plurality of similarity values;

[0122] 02214: Based on a preset sorting method, perform a sorting process on the plurality of similarity values to obtain a sorting result of the plurality of similarity values;

[0123] 02215: According to the sorting result and the first target-associated similar expression subset, determine a target-associated similar expression sub-table corresponding to the first target similar expression;

[0124] 02216: Determine an associated similar expression table according to the target-associated similar expression sub-table.

[0125] In some embodiments, the determining module is further configured to perform an encoding process on the first target similar expression to determine a first encoding. And perform an encoding process on each second similar expression in the first target-associated similar expression subset respectively to determine a plurality of second encodings. And based on a preset similarity algorithm, calculate the similarity between the first encoding and each second encoding respectively to obtain a plurality of similarity values. The determining module is further configured to perform a sorting process on the plurality of similarity values based on a preset sorting method to obtain a sorting result of the plurality of similarity values. And according to the sorting result and the first target-associated similar expression subset, determine a target-associated similar expression sub-table corresponding to the first target similar expression. And determine an associated similar expression table according to the target-associated similar expression sub-table.

[0126] In some embodiments, the processor is further configured to perform an encoding process on the first target similar expression to determine a first encoding. And perform an encoding process on each second similar expression in the first target-associated similar expression subset respectively to determine a plurality of second encodings. And based on a preset similarity algorithm, calculate the similarity between the first encoding and each second encoding respectively to obtain a plurality of similarity values. The processor is further configured to perform a sorting process on the plurality of similarity values based on a preset sorting method to obtain a sorting result of the plurality of similarity values. And according to the sorting result and the first target-associated similar expression subset, determine a target-associated similar expression sub-table corresponding to the first target similar expression. And determine an associated similar expression table according to the target-associated similar expression sub-table.

[0127] Specifically, the first encoding refers to a vector encoding obtained by performing an encoding process on the first target similar expression based on a second preset algorithm. Taking the BERT vectorization algorithm as an example, the first encoding Vec1 = Bert_Embedding(Similar question D1).

[0128] The second encoding refers to the vector encoding obtained by encoding the second similar expressions in the first target associated similar expression subset based on a second preset algorithm. The second encoding Vec2 = Bert_Embedding(Similar question B1).

[0129] The preset similarity algorithm refers to an algorithm such as the cosine similarity algorithm that can calculate the similarity value between texts, and is used to further calculate the similarity between similar questions to obtain a more accurate similarity value for determining whether there are conflicts or duplicates. Taking the cosine similarity algorithm as an example, the calculation formula for the similarity value is as follows:

[0130]

[0131] Among them, Vec1 i is the first target similar expression being processed currently, and Vec2 i is the second similar expression being processed currently in the first target associated similar expression subset.

[0132] The preset sorting method refers to sorting the similar questions according to the similarity value. The higher the similarity value, the higher the similarity between the similar expressions, and the greater the potential for conflicts or duplicates.

[0133] Encode the first target similar expression to determine the first encoding. Continuing with the above example, encode the similar question D1 to obtain the first encoding X1.

[0134] Next, the server encodes each second similar expression in the first target associated similar expression subset separately to determine multiple second encodings. Continuing with the above example, encode the similar questions B1, E1, A1, and C1 separately to obtain the second encoding Y1, the second encoding Y2, the second encoding Y3, and the second encoding Y4.

[0135] Subsequently, based on the preset similarity algorithm, the server calculates the similarity between the first encoding and each second encoding to obtain multiple similarity values. Continuing with the above example, calculate the similarity between the first encoding and each second encoding, and the similarity value of the similar question B1 is 0.95, the similarity value of the similar question E1 is 0.91, the similarity value of the similar question A1 is 0.83, and the similarity value of the similar question C1 is 0.79.

[0136] Then, based on the preset sorting method, the server sorts the multiple similarity values to obtain the sorting result of the multiple similarity values. Continuing with the above example, the sorting result is that the similarity value of the similar question B1 is 0.95, the similarity value of the similar question E1 is 0.91, the similarity value of the similar question A1 is 0.83, and the similarity value of the similar question C1 is 0.79.

[0137] The server also determines a target associated similar expression sub-table corresponding to the first target similar expression based on the sorting result and the first target associated similar expression subset. Continuing with the above example, based on the sorting result and the first target associated similar expression subset, the target associated similar expression sub-table corresponding to the similar question D1 is determined as "Similar question B1: 0.95". Similarly, the target associated similar expression sub-table corresponding to the similar question D2 can be obtained as "Similar question B2: 0.9", and the target associated similar expression sub-table corresponding to the similar question D3 is "Similar question A3: 0.87".

[0138] Finally, the server determines the associated similar expression table based on the target associated similar expression sub-table. Continuing with the above example, the associated similar expression table of the first knowledge data D is determined as "Similar question B1: 0.95; Similar question B2: 0.9; Similar question A3: 0.87".

[0139] In this way, the first target similar expression is encoded to determine the first code. Then, the server encodes each second similar expression in the first target associated similar expression subset respectively to determine a plurality of second codes. Subsequently, based on a preset similarity algorithm, the server calculates the similarity between the first code and each second code respectively to obtain a plurality of similarity values. Then, based on a preset sorting method, the server sorts the plurality of similarity values to obtain a sorting result of the plurality of similarity values. The server also determines a target associated similar expression sub-table corresponding to the first target similar expression based on the sorting result and the first target associated similar expression subset. Finally, the server determines the associated similar expression table according to the target associated similar expression sub-table. In this way, through the encoding process and the preset similarity algorithm, the similarity between similar questions can be accurately calculated, thereby improving the accuracy of knowledge base cleaning.

[0140] Please refer to Figure 6 , in some embodiments, step 03 (cleaning the knowledge base according to the first similar expression set and the first associated similar expression set) includes:

[0141] 031: When the first associated similar expression set includes the first associated similar expression, determine the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set;

[0142] 032: Clean the knowledge base according to the knowledge data to be cleaned.

[0143] In some embodiments, the determination module is further configured to, when the first associated similar expression set includes the first associated similar expression, determine the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set, and clean the knowledge base according to the knowledge data to be cleaned.

[0144] In some embodiments, the processor is further configured to, when the first associated similar expression set includes the first associated similar expression, determine the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set. And clean the knowledge base according to the knowledge data to be cleaned.

[0145] Specifically, the knowledge data to be cleaned refers to the knowledge data that has been screened and is considered to have potential conflicts or duplicates.

[0146] The cleaning process refers to further analyzing and judging the knowledge data to be cleaned, identifying and resolving the conflicts existing in the knowledge base, and ensuring the consistency and accuracy of the knowledge base. In some embodiments, a preset rule base is used to automatically identify and process the conflicts in the knowledge base without manual intervention. These rules are customized according to the actual application scenarios and domain knowledge.

[0147] First, check whether the first associated similar expression set contains the first associated similar expression. If it contains, determine the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set.

[0148] Next, perform a cleaning process on the knowledge data to be cleaned, such as deleting duplicate similar questions, updating outdated or inaccurate information, and optimizing the correspondence between questions and answers. Through the cleaning process, redundant information can be removed, outdated information can be updated, and the correspondence between questions and answers can be optimized, thereby improving the overall quality of the knowledge base.

[0149] Thus, when the first associated similar expression set includes the first associated similar expression, determine the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set. Then, the server cleans the knowledge base according to the knowledge data to be cleaned. In this way, by cleaning the knowledge base, the conflicting information existing in the knowledge base can be eliminated, the consistency and accuracy of the knowledge base can be ensured, thereby improving the overall quality of the knowledge base and enhancing the user experience.

[0150] Please refer to Figure 7 , in some embodiments, step 031 (determine the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set) includes:

[0151] 0311: Determine the third knowledge data set based on the second knowledge points corresponding to each first associated similar expression;

[0152] 0312: Determine the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set.

[0153] In some embodiments, the determining module is further configured to determine a third knowledge data set based on the second knowledge points corresponding to each first associated similar expression. And determine the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set.

[0154] In some embodiments, the processor is further configured to determine a third knowledge data set based on the second knowledge points corresponding to each first associated similar expression. And determine the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set.

[0155] Specifically, the second knowledge point refers to the knowledge point corresponding to the first associated similar expression. Please refer to Table 1 again. Continuing with the above example, the first associated similar expressions corresponding to the first knowledge data D include similar question B1 and similar question B2, then the second knowledge point is knowledge point B.

[0156] The second knowledge data refers to the knowledge data that has been processed to obtain knowledge data that may potentially conflict with or duplicate the first knowledge data. Continuing with the above example, the second knowledge data corresponding to the first knowledge data D is knowledge data B.

[0157] The third knowledge data set refers to a set composed of the second knowledge data. Continuing with the above example, the third knowledge data set includes knowledge data B.

[0158] Based on the second knowledge points corresponding to each first associated similar expression, the server determines the third knowledge data set, and the third knowledge data set includes at least one second knowledge data that matches the first knowledge. Continuing with the above example, the first associated similar expressions corresponding to the first knowledge data D include similar question B1 and similar question B2, then the second knowledge point is knowledge point B. According to knowledge point B, the second knowledge data corresponding to the first knowledge data D is determined to be knowledge data B.

[0159] Next, the server determines the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set.

[0160] In this way, based on the second knowledge points corresponding to each first associated similar expression, the server determines the third knowledge data set, and the third knowledge data set includes at least one second knowledge data that matches the first knowledge. Next, the server determines the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set. In this way, by analyzing the association between similar knowledge points, it is possible to accurately determine whether there are conflicts or duplications in the knowledge data, improving the cleaning accuracy.

[0161] Please refer to Figure 8 , in some embodiments, step 0312 (determine the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set) includes:

[0162] 03121: Determine the matching degree between the second knowledge data and the first knowledge data according to the ratio of the first similarity quantity and the second similarity quantity;

[0163] 03122: Based on a preset matching degree threshold, determine the knowledge data to be cleaned according to the matching degree and the third knowledge data set.

[0164] In some embodiments, the determining module is further configured to determine the matching degree between the second knowledge data and the first knowledge data according to the ratio of the first similarity quantity and the second similarity quantity. And based on a preset matching degree threshold, determine the knowledge data to be cleaned according to the matching degree and the third knowledge data set.

[0165] In some embodiments, the processor is further configured to determine the matching degree between the second knowledge data and the first knowledge data according to the ratio of the first similarity quantity and the second similarity quantity. And based on a preset matching degree threshold, determine the knowledge data to be cleaned according to the matching degree and the third knowledge data set.

[0166] Specifically, the first similarity quantity is the quantity of the first associated similar expressions in the second knowledge data currently being matched. Continuing the above example, the second knowledge data currently being matched is knowledge data B, and the first associated similar expressions in knowledge data B related to knowledge data D include similar question B1 and similar question B2, so the first similarity quantity is 2.

[0167] The second similarity quantity is the quantity of the first similar expressions in the first similar expression set. Continuing the above example, referring to Table 1 again, the first knowledge data currently being processed is knowledge data D, and the corresponding first similar expression set of knowledge data D is the similar question set D, including similar question D-1, similar question D-2, and similar question D-3, so the second similarity quantity is 3.

[0168] The matching degree refers to the similarity degree between two knowledge data. Continuing the above example, the matching degree between knowledge data D and knowledge data B is 2 / 3 = 0.67.

[0169] The preset matching degree threshold refers to a preset value used to determine whether the similarity between two knowledge data is high enough to determine whether there is a conflict between them. When the matching degree between two knowledge data is greater than the preset matching degree threshold, it indicates that there may be a conflict between these two knowledge data and further analysis and processing are required. When the matching degree between two knowledge data is less than the preset matching degree threshold, it indicates that there may be no conflict between these two knowledge data and they can continue to be retained in the knowledge base. The preset matching degree threshold needs to be adjusted according to the specific application scenario and the characteristics of the knowledge base. Here, taking the preset matching degree threshold as 0.6 as an example, the embodiments of the present application are described.

[0170] Continuing with the above example, the matching degree between knowledge data D and knowledge data B is 2 / 3 = 0.67, and 0.67 is greater than 0.6. Therefore, knowledge data D and knowledge data B are taken as the data to be cleaned.

[0171] It should be noted that in some embodiments, if there is more than one second knowledge data, that is, there is more than one second knowledge point corresponding to the first associated similar expression, it is necessary to calculate the matching degree between each second knowledge data and the first knowledge data, and then compare it with the preset matching degree threshold to determine the data to be cleaned.

[0172] In this way, according to the ratio of the first similar quantity to the second similar quantity, the server determines the matching degree between the second knowledge data and the first knowledge data, where the first similar quantity is the quantity of the first associated similar expressions in the currently matched second knowledge data, and the second similar quantity is the quantity of the first similar expressions in the first similar expression set. Then, based on the preset matching degree threshold, the server determines the knowledge data to be cleaned according to the matching degree and the third knowledge data set. In this way, by calculating the matching degree and the preset matching degree threshold, the knowledge points that conflict with or are redundant to the first knowledge data can be accurately identified, thereby improving the accuracy of knowledge base cleaning. Moreover, the preset matching degree threshold can avoid misjudging knowledge points with a low matching degree as conflicting or redundant information, reducing the possibility of misjudgment.

[0173] Please refer to Figure 9 , in some embodiments, step 03 (cleaning the knowledge base according to the first similar expression set and the first associated similar expression set) includes:

[0174] 033: When the first associated similar expression set does not include the first associated similar expression, confirm that the first knowledge data does not need to be cleaned.

[0175] In some embodiments, the confirmation module is further configured to confirm that the first knowledge data does not need to be cleaned when the first associated similar expression set does not include the first associated similar expression.

[0176] In some embodiments, the processor is further configured to confirm that the first knowledge data does not need to be cleaned when the first associated similar expression set does not include the first associated similar expression.

[0177] Specifically, if the first associated similar expression set does not include the first associated similar expression, it is confirmed that the first knowledge data does not need to be cleaned. By judging whether the first associated similar expression set contains the first associated similar expression, it is possible to avoid misjudging the first knowledge data without conflicting or redundant information as the knowledge data to be cleaned.

[0178] Thus, in the case that the first associated similar expression set does not include the first associated similar expression, the server confirms that the first knowledge data does not need to be cleaned. In this way, by determining whether the first associated similar expression set contains the first associated similar expression, it is possible to avoid misjudging the first knowledge data without conflicting or redundant information as the knowledge data to be cleaned.

[0179] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the knowledge base data processing method for an in-vehicle dialogue system as described above are implemented.

[0180] It can be understood that the computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), and software distribution media, etc.

[0181] In the description of this specification, the descriptions referring to terms such as "specifically", "furthermore", "specially", "understandably", etc. mean that the specific features, structures, materials, or characteristics described in combination with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0182] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of executable request code including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a manner that is not shown or discussed in sequence, including in a substantially simultaneous manner or in a reverse order according to the involved functions, which should be understood by those skilled in the art of the embodiments of the present application.

[0183] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A knowledge base data processing method for a vehicle-mounted dialogue system, characterized in that: The method comprises: Acquire first knowledge data, the first knowledge data comprising a first knowledge point and a first similar expression set corresponding to the first knowledge point; Determine a first associated similar expression set from the knowledge base according to the first similar expression set, wherein each first similar expression in the first similar expression set corresponds to at most one first associated similar expression in the first associated similar expression set; The knowledge base is cleaned according to the first similar expression set and the first associated similar expression set.

2. The method according to claim 1, characterized in that The determining, based on the first similar expression set, from the knowledge base, a first associated similar expression set comprises: According to each of the first similar expressions, determining a second associated similar expression set from the knowledge base, wherein the second associated similar expression set includes a plurality of first associated similar expression subsets, and each of the first associated similar expression subsets corresponds to one of the first similar expressions; Determine at most one first associated similar expression according to each of the first similar expressions and the corresponding first associated similar expression subset; The first associated similar expression set is determined according to the first associated similar expressions.

3. The method according to claim 2, characterized in that The step of determining a second set of associated similar expressions from the knowledge base according to each of the first similar expressions comprises: Based on a first preset algorithm, according to each of the first similar expressions, determine a third associated similar expression set from the knowledge base, the third associated similar expression set including a plurality of second associated similar expression subsets, each of the second associated similar expression subsets corresponding to one of the first similar expressions; According to the first similar expression, performing deduplication processing on each of the second associated similar expression subsets to determine a plurality of the first associated similar expression subsets; The second associated similar expression set is determined according to the first associated similar expression subset.

4. The method according to claim 2, characterized in that: The determining at most one first associated similar expression according to each of the first similar expressions and the corresponding first associated similar expression subset comprises: Based on a second preset algorithm, determining a target-associated similarity expression table according to a first target similarity expression and a first target-associated similarity expression subset corresponding to the first target similarity expression, wherein the first similarity expression includes the first target similarity expression, and the first associated similarity expression subset includes the first target-associated similarity expression subset; Based on a preset similarity threshold, at most one first target-associated similar expression corresponding to the first target-associated similar expression is determined according to the target-associated similar expression table, wherein the first associated similar expression includes the first target-associated similar expression.

5. The method according to claim 4, characterized in that The determining of the target-associated similarity statement table based on the second preset algorithm and the first target-associated similarity statement subset corresponding to the target first similarity statement includes: Encoding the first target similar expression to determine the first code; Performing encoding processing on each second similar expression in the first target-related similar expression subset to determine a plurality of second codes; Based on a preset similarity algorithm, respectively calculating the similarity between the first code and each of the second codes to obtain a plurality of similarity values; Based on a preset sorting method, the multiple similarity values ​​are sorted to obtain a sorting result of the multiple similarity values; Determine a target-associated similar expression sub-table corresponding to the first target-associated similar expression according to the ranking result and the first target-associated similar expression subset; The associated similarity statement table is determined according to the target associated similarity statement sub-table.

6. The method according to claim 1, characterized in that The step of cleaning the knowledge base according to the first similar expression set and the first associated similar expression set includes: In a case where the first associated similar expression set includes the first associated similar expression, determining the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set; The knowledge base is cleaned according to the knowledge data to be cleaned.

7. The method according to claim 6, characterized in that The step of determining the knowledge data to be cleaned according to the first similar expression set and the first associated similar expression set includes: Determine a third knowledge data set based on the second knowledge point corresponding to each of the first associated similar expressions, wherein the third knowledge data set includes at least one second knowledge data matching the first knowledge; The knowledge data to be cleaned is determined according to the first knowledge data and the third knowledge data set.

8. The method according to claim 7, characterized in that The step of determining the knowledge data to be cleaned according to the first knowledge data and the third knowledge data set includes: Determine the matching degree between the second knowledge data and the first knowledge data according to the ratio of the first similarity number to the second similarity number, wherein the first similarity number is the number of the first associated similar expressions in the second knowledge data currently being matched, and the second similarity number is the number of the first similar expressions in the first similar expression set; Based on a preset matching degree threshold, the knowledge data to be cleaned is determined according to the matching degree and the third knowledge data set.

9. The method according to claim 1, characterized in that: The step of cleaning the knowledge base according to the first similar expression set and the first associated similar expression set includes: When the first associated similar expression set does not include the first associated similar expression, it is confirmed that the first knowledge data does not need to be cleaned.

10. A server, characterized in that: The server includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the method according to any one of claims 1 to 9 is implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.