Multi-domain knowledge automatic extraction and analysis platform and method

By collecting user intention information, calculating user similarity and problem scores, and using logistic regression algorithm to filter answers, the problem of ignoring user historical problems in the existing technology is solved, and the accuracy and relevance of answers are improved.

CN119940499APending Publication Date: 2025-05-06SHENZHEN LIANGYI INTERACTIVE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510013906.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the system outputs answers based on the questions asked by the user, ignoring the user's historical questions, resulting in the answers being answered that may not conform to the user's true intentions.

Method used

By collecting user intention information, calculating user similarity and problem scores, using logistic regression algorithm to calculate the conformity coefficients, and filtering out answers that meet the actual needs of users.

Benefits of technology

This makes the filtered answer more in line with the actual needs of users and improves the accuracy and relevance of the answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940499A_ABST
    Figure CN119940499A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-domain knowledge automatic extraction and analysis platform and method, and particularly relates to the technical field of knowledge automatic extraction and analysis. According to problems input by a user, user intention information is collected, and user similarity is obtained through historical knowledge domain searching frequency calculation; calculating question scores according to the inquiry question interval time and the inquiry question similarity, performing weight assignment on each question score according to an optimal sequence graph method, and determining a weight value of each question; a plurality of user questions are integrated, so that the answering is more accurate; calculating the similarity between each question and a historical question according to the user intention information, and carrying out weighted summation on a plurality of questions asked by the user to obtain the comprehensive similarity of the plurality of questions; calculating a coincidence coefficient by using a logistic regression algorithm according to the user similarity and the comprehensive similarity of the plurality of problems; and screening out answers according to the coincidence coefficient and a prediction threshold value, and outputting the answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic knowledge extraction and analysis, and more specifically, to a multi-domain knowledge automatic extraction and analysis platform and method. Background Art

[0002] The background technology of multi-domain knowledge automatic extraction and analysis method and platform is an important research direction in the current field of artificial intelligence and big data. It involves how to automatically identify, extract, organize and analyze valuable knowledge from massive data, and effectively use this knowledge in multiple fields or multiple application scenarios. This technology has been widely used in natural language processing, knowledge graph, intelligent question answering, information retrieval and other fields.

[0003] In the prior art, the system outputs answers based on the questions asked by the user, ignoring the user's previous questions, resulting in the answer not being able to meet the user's true intention. Summary of the invention

[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a multi-domain knowledge automatic extraction and analysis platform and method to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] The following steps are involved:

[0007] Step S1: Collect user intention information;

[0008] Step S2: Calculate user similarity using cosine similarity based on the frequency of user historical searches for different knowledge fields; Calculate question scores based on the time interval between questions asked and the similarity of questions asked, and assign weights to each question score using the priority graph method to determine the weight value of each question; Calculate the similarity of several questions asked by the user based on user intent information and questions asked by historical users, and perform weighted summation on several questions asked by the user to obtain the comprehensive similarity of several questions asked by the user;

[0009] Step S3: Calculate the matching coefficient using a logistic regression algorithm based on the user similarity and the comprehensive similarity of several questions;

[0010] Step S4: Filter out answers based on the matching coefficient and prediction threshold.

[0011] In a preferred embodiment, in step S1, the user intention information includes question keywords, question types, and predicate verbs.

[0012] In a preferred embodiment, in step S2, the frequency of users' historical searches for different knowledge fields is determined, and the user similarity is calculated using the cosine similarity formula.

[0013] In a preferred embodiment, in step S2, the user asks several questions, and the time interval between the questions asked now and the questions asked in the past is calculated by calculating the time difference between the questions asked now and the questions asked in the past. The user intention information is converted into a low-dimensional vector representation using a word embedding model to measure the similarity of the two question vectors; the similarity of the user questions is calculated by cosine similarity; and the similarity of the questions asked now and the questions asked in the past are calculated respectively to obtain the similarity of the questions asked now.

[0014] In a preferred embodiment, in step S2, the time interval between questions asked and the similarity of questions asked are determined and normalized; the score of each question is calculated based on the normalized time interval between questions asked and the similarity of questions asked, and the score of each question is weighted according to the priority graph method to determine the weight value of each question, and the similarity between each question and historical user questions is calculated based on user intent information, and the weighted sum is performed to obtain the comprehensive similarity of several questions.

[0015] In a preferred embodiment, in step S4, the compliance coefficient is determined, the compliance coefficient is compared with a system preset threshold, historical answers whose compliance coefficient is greater than the system threshold are screened out, and the historical answers are sorted and output.

[0016] In a preferred embodiment, it includes a user intention information collection and analysis module, a calculation module, a judgment module, an output module, and a storage module;

[0017] The user intention information collection module is used to obtain user intention information and record the time when the user asks a question, and transmit the user intention information and the time when the user asks a question to the calculation module;

[0018] The calculation module is used to calculate user similarity based on user historical searches; calculate question scores based on the time interval between questions asked and the similarity of questions asked, and assign weights to each question score according to the priority graph method to determine the weight value of each question; calculate the similarity between each question and historical questions based on user intention information, and perform weighted summation on several questions asked by the user to obtain the comprehensive similarity of several questions; transmit the user similarity and the comprehensive similarity of several questions to the judgment module;

[0019] The judgment module is used to calculate the matching coefficient according to the user similarity and the comprehensive similarity of several questions, and compare the matching coefficient with the system threshold, screen out the historical answer information with the matching coefficient greater than the system threshold, and transmit the historical answer information to the output module;

[0020] The output module is used to output the answer;

[0021] The storage module is used to store historical answer information and user historical search information; and can transmit the historical answer information to the judgment module and transmit the user historical search information to the calculation module.

[0022] Technical effects and advantages of the present invention:

[0023] The present invention collects user intention information through questions input by users, obtains user similarity through frequency calculation of historical search knowledge fields, calculates question scores according to the interval time between questions asked and the similarity of questions asked, and assigns weights to the scores of each question according to the priority graph method to determine the weight value of each question; calculates the similarity between each question and historical questions according to the user intention information, performs weighted summation on several questions asked by the user to obtain the comprehensive similarity of several questions; uses a logistic regression algorithm to deduce the conformity coefficient according to the user similarity and the comprehensive similarity of several questions; screens out answers according to the conformity coefficient and the prediction threshold, and outputs the answers.

[0024] The comprehensive similarity of several questions asked by the user is calculated by weighted summation, so that the screened answers are more in line with the actual needs of the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to facilitate understanding by those skilled in the art, the present invention is further described below in conjunction with the accompanying drawings;

[0026] Figure 1 It is a flowchart of the method for automatically extracting and analyzing multi-domain knowledge of the present invention;

[0027] Figure 2 It is a structural diagram of the multi-domain knowledge automatic extraction and analysis platform of the present invention. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] The present invention collects user intention information through questions input by users, obtains user similarity through frequency calculation of historical search knowledge fields, calculates question scores according to the interval time between questions asked and the similarity of questions asked, and assigns weights to the scores of each question according to the priority graph method to determine the weight value of each question; calculates the similarity between each question and historical questions according to the user intention information, performs weighted summation on several questions asked by the user to obtain the comprehensive similarity of several questions; uses a logistic regression algorithm to deduce the conformity coefficient according to the user similarity and the comprehensive similarity of several questions; screens out answers according to the conformity coefficient and the prediction threshold, and outputs the answers.

[0030] Example 1

[0031] Multi-domain knowledge automatic extraction and analysis methods, such as Figure 1 As shown, the following steps are included:

[0032] Step S1: Collect user intention information;

[0033] Step S2: Calculate user similarity using cosine similarity based on the frequency of user historical searches for different knowledge fields; Calculate question scores based on the time interval between questions asked and the similarity of questions asked, and assign weights to each question score using the priority graph method to determine the weight value of each question; Calculate the similarity of several questions asked by the user based on user intent information and questions asked by historical users, and perform weighted summation on several questions asked by the user to obtain the comprehensive similarity of several questions asked by the user;

[0034] Step S3: Calculate the matching coefficient using a logistic regression algorithm based on the user similarity and the comprehensive similarity of several questions;

[0035] Step S4: Filter out answers based on the matching coefficient and prediction threshold.

[0036] The specific steps are as follows:

[0037] In step S1, after the user enters the question to be answered in the dialog box, in the process of identifying the sentence structure, the subject, predicate, object and other components are determined by analyzing the relationship between the various components in the sentence. Common technologies include:

[0038] Dependency Parsing: Dependency parsing identifies how words in a sentence depend on other words and thus identifies the subject. The subject is usually a direct dependent of the verb, and in most cases, the verb points to a noun or noun phrase as the subject.

[0039] Phrase structure syntactic analysis: Phrase structure syntactic analysis identifies the subject through the phrase hierarchy in the sentence. The subject is usually a noun phrase (NP).

[0040] First, you need to determine the subject word of the sentence. The subject word of the sentence is usually the core object discussed in the question. It is the most important noun or noun phrase in the question. By analyzing the subject part of the sentence, you can determine this subject word.

[0041] The second step is to identify the type of question. There are usually different types of questions, and each type of question focuses on different points, for example:

[0042] Who / What: Ask about people, things, identities, etc.

[0043] Why: Asking for reasons or motivations.

[0044] How: Ask about the process, method or approach.

[0045] Where: Asking about a place or location.

[0046] When: Asking about the time.

[0047] Identifying the question type can help you focus on specific key points. For example, "who" usually concerns people, "how" focuses on methods or steps, and "where" is related to locations.

[0048] The third step is to identify keywords and predicate verbs. Keywords: nouns, verbs, adjectives, etc. in a sentence are the key components of a question. For example, "color" is a key noun in "What color is an apple?" and "president" is a key noun in "Who is the president of the United States?" Predicate verbs: Predicate verbs help identify the core action or requirement of the question. For example, in "What do you like to eat?", the verb "like" determines that the subject of the question is "food," and "what" is an interrogative pronoun, indicating that a specific answer is being sought.

[0049] In step S2, the user's historical search frequency for different knowledge fields is determined.

[0050] Calculate user similarity:

[0051] Construct the frequency vector of user historical search in different knowledge fields: u1=(r u1,1 ,r u2,2 ,....r. u1,,n )u2=(r u2,1 ,r u2,2 ,....,r u2,n );

[0052] where r u1,i and r u2,i Respectively represent the search frequency of user u1 and user u2 on the same knowledge field.

[0053] Calculate the cosine similarity between two users:

[0054]

[0055] Where I is the set of knowledge areas that both users have searched.

[0056] Determine similarity:

[0057] If the calculated cosine similarity value is close to 1, it means that the interests of the two users are very similar;

[0058] If the cosine similarity value is close to 0, it means that the interests of the two users are not similar;

[0059] If the cosine similarity value is close to -1, it means that the interests of the two users are completely opposite.

[0060] In practical applications, we can set a threshold to determine whether two users are similar, depending on the specific situation. If the cosine similarity is greater than the threshold, the two users are considered very similar; if the cosine similarity is less than the threshold and is not much different from the threshold, the two users are considered to have a certain degree of similarity; if the cosine similarity is much less than the threshold, the two users are considered to have little similarity.

[0061] Several questions are raised by the user. This embodiment takes five questions A, B, C, D, and E as examples, where question E is the last question raised and question A is the first question raised; the time interval between asking question A is the difference between the time when question E is asked and the time when question A is asked. Similarly, the time interval between asking question B is the difference between the time when question E is asked and the time when question B is asked; the time interval between asking question C is the difference between the time when question E is asked and the time when question C is asked; the time interval between asking question D is the difference between the time when question E is asked and the time when question D is asked; the time interval between asking question E is 0; the longer the time interval between asking questions, the lower the importance of the question, and vice versa.

[0062] Use word embedding models (such as Word2Vec, GloVe, FastText, etc.) to convert user intent information into low-dimensional vector representations to measure the similarity of two question vectors. These models are trained with a large number of corpora and can capture the semantic relationship between words. The similarity of user questions can be calculated by cosine similarity, as described above; calculate the similarity between question E and questions E, D, C, B, and A respectively. The higher the similarity of the question, the more important the question is, and vice versa, the less important the question is.

[0063] Linear normalization is used to map the question interval and question similarity to the interval [0,1], the formula is: In the formula, Normalized_Score is the normalized value.

[0064] After normalizing the time interval between questions and the similarity of questions, they are marked as a and b respectively, and the formula for calculating the score of each question is: Tem(i) = ba. In the formula, Tem(i) is the score of each question, and i represents the sequence number of each question;

[0065] After calculating the score of each question, the scores of each question are compared.

[0066] Obviously, the shorter the time interval between questions / the higher the similarity of questions asked, the higher the question score will be, and the more attention it needs to receive, that is, the more important it is, and the greater the weight it should take in question consideration.

[0067] The score of each question is calculated based on the normalized interval between questions and the similarity of the questions asked, and the score of each question is weighted according to the priority graph method to determine the weight value of each question.

[0068] Specifically, the weights of the scores of each question are assigned according to the priority diagram method as shown in Table 1 below:

[0069]

[0070]

[0071] Table 1

[0072] It should be noted that in Table 1, questions A, B, C, D, and E correspond one by one to the five user questions in this example. After sorting the questions according to the size of the question scores, they correspond to the questions in the order of E, D, C, B, and A from large to small.

[0073] According to the user intention information, the similarities Ta, Tb, Tc, Td, and Te of questions A, B, C, D, and E with other historical user questions are calculated respectively. The calculation method is the same as the similarity calculation between A, B, C, D, and E mentioned above; the comprehensive similarity of several questions is weighted: T = a*Ta+b*Tb+c*Tc+d*Td+e*Te; where T represents the comprehensive similarity of several questions; a, b, c, d, and e represent the weights of questions A, B, C, D, and E respectively.

[0074] In step S3, the user similarity and the comprehensive similarity of several questions are determined, and the user similarity and the comprehensive similarity of several questions are combined to calculate the matching coefficient according to the logistic regression algorithm: P = 1-e -(T1×D1+T2×D2) ; P represents the matching coefficient, D1 represents the user similarity, where the greater the user similarity, the greater the matching coefficient; otherwise, the smaller the matching coefficient; D2 represents the comprehensive similarity of several questions, where the greater the comprehensive similarity of several questions, the smaller the matching coefficient; T1, T2 are the regression coefficients of each variable, and T1, T2 are both greater than zero.

[0075] In step S4, the compliance coefficient is determined, the compliance coefficient is compared with the system preset threshold, the historical answers with the compliance coefficient greater than the system threshold are screened out, the historical answers are extracted from the multi-domain knowledge, and the answers are sorted and output.

[0076] Example 2

[0077] The present invention provides Figure 2 The multi-domain knowledge automatic extraction and analysis platform shown includes a user intention information collection and analysis module, a calculation module, a judgment module, an output module, and a storage module;

[0078] The user intention information collection module is used to obtain user intention information and record the time when the user asks a question, and transmit the user intention information and the time when the user asks a question to the calculation module;

[0079] The calculation module is used to calculate user similarity based on user historical searches; calculate question scores based on the time interval between questions asked and the similarity of questions asked, and assign weights to each question score according to the priority graph method to determine the weight value of each question; calculate the similarity between each question and historical questions based on user intention information, and perform weighted summation on several questions asked by the user to obtain the comprehensive similarity of several questions; transmit the user similarity and the comprehensive similarity of several questions to the judgment module;

[0080] The judgment module is used to calculate the matching coefficient according to the user similarity and the comprehensive similarity of several questions, and compare the matching coefficient with the system threshold, screen out the historical answer information with the matching coefficient greater than the system threshold, and transmit the historical answer information to the output module;

[0081] The output module is used to output the answer;

[0082] The storage module is used to store historical answer information and user historical search information; and can transmit the historical answer information to the judgment module and transmit the user historical search information to the calculation module.

[0083] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0084] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0085] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0086] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0087] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A multi-domain knowledge automatic extraction and analysis method, characterized in that: The following steps are involved: Step S1: Collect user intention information; Step S2: Calculate user similarity using cosine similarity based on the frequency of user historical searches for different knowledge fields; Calculate question scores based on the time interval between questions asked and the similarity of questions asked, and assign weights to each question score using the priority graph method to determine the weight value of each question; Calculate the similarity of several questions asked by the user based on user intent information and questions asked by historical users, and perform weighted summation on several questions asked by the user to obtain the comprehensive similarity of several questions asked by the user; Step S3: Calculate the matching coefficient using a logistic regression algorithm based on the user similarity and the comprehensive similarity of several questions; Step S4: Filter out answers based on the matching coefficient and prediction threshold.

2. The multi-domain knowledge automatic extraction and analysis method according to claim 1 is characterized in that: In step S1 , the user intention information includes question keywords, question types, and predicate verbs.

3. The multi-domain knowledge automatic extraction and analysis method according to claim 1 is characterized in that: In step S2, the frequency of users' historical searches for different knowledge fields is determined, and the user similarity is calculated using the cosine similarity formula.

4. The multi-domain knowledge automatic extraction and analysis method according to claim 1 is characterized in that: In step S2, the user asks several questions, and the time interval between the questions asked now and the questions asked in the past is calculated by calculating the time difference between them. The user intention information is converted into a low-dimensional vector representation using a word embedding model to measure the similarity of the two question vectors. The similarity of the user questions is calculated by cosine similarity. The similarity of the questions asked now and the questions asked in the past are calculated to obtain the similarity of the questions asked.

5. The multi-domain knowledge automatic extraction and analysis method according to claim 4 is characterized in that: In step S2, the time interval between questions and the similarity between questions are determined and normalized; the score of each question is calculated based on the normalized time interval between questions and the similarity between questions, and the score of each question is weighted according to the priority graph method to determine the weight value of each question, and the similarity between each question and historical user questions is calculated based on the user intent information, and the weighted sum is performed to obtain the comprehensive similarity of several questions.

6. The multi-domain knowledge automatic extraction and analysis method according to claim 1 is characterized in that: In step S3, the user similarity and the comprehensive similarity of several questions are determined, and the matching coefficient is calculated according to the logistic regression algorithm: P = 1-e -(T1×D1+T2×D2) ; Among them, P represents the matching coefficient, D1 represents the user similarity, D2 represents the comprehensive similarity of several questions, and T1 and T2 are the regression coefficients of each variable.

7. The multi-domain knowledge automatic extraction and analysis method according to claim 1 is characterized by: In step S4, the matching coefficient is determined, the matching coefficient is compared with a system preset threshold, historical answers whose matching coefficient is greater than the system threshold are screened out, and the historical answers are sorted and output.

8. A multi-domain knowledge automatic extraction and analysis platform, used to implement the multi-domain knowledge automatic extraction and analysis method according to any one of claims 1 to 7, characterized in that: It includes a user intention information collection and analysis module, a calculation module, a judgment module, an output module, and a storage module; The user intention information collection module is used to obtain user intention information and record the time when the user asks a question, and transmit the user intention information and the time when the user asks a question to the calculation module; The calculation module is used to calculate user similarity based on user historical searches; Calculate the question score based on the time interval between questions asked and the similarity of the questions asked, and assign weights to the scores of each question according to the priority graph method to determine the weight value of each question; Calculate the similarity between each question and the historical questions based on the user's intention information, perform weighted summation on the questions asked by the user to obtain the comprehensive similarity of the questions; transmit the user similarity and the comprehensive similarity of the questions to the judgment module; The judgment module is used to calculate the matching coefficient according to the user similarity and the comprehensive similarity of several questions, and compare the matching coefficient with the system threshold, screen out the historical answer information with the matching coefficient greater than the system threshold, and transmit the historical answer information to the output module; The output module is used to output the answer; The storage module is used to store historical answer information and user historical search information; and can transmit the historical answer information to the judgment module and transmit the user historical search information to the calculation module.