Method and system for case retrieval and computer readable storage medium

By calculating similarity between problem keywords and case entities, and sorting based on field similarity, the method enhances case retrieval accuracy and efficiency in enterprise libraries.

CN120316129APending Publication Date: 2025-07-15HENAN MUYUAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510251764.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In traditional case search methods, the correlation between case content and search problems is low, resulting in insufficient search accuracy.

Method used

By extracting the similarity between the problem keywords in the search question and the case entity for recall, sorting it in combination with the similarity of the case fields, using user information for personalized sorting, improving search accuracy and efficiency.

Benefits of technology

It realizes rapid and accurate search of cases that meet the search problems, meets the personalized needs of different users, and improves the effectiveness of case search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316129A_ABST
    Figure CN120316129A_ABST
Patent Text Reader

Abstract

The invention discloses a case retrieval method and system and a computer readable storage medium. A case comprises a plurality of case entities and case fields, and the method comprises the steps that a retrieval problem and user information are received, effective information of the retrieval problem is obtained, and the effective information comprises one or more problem keywords; obtaining a first similarity between the question keyword and the case entity, and performing recall based on the first similarity to obtain a preliminary retrieval case; obtaining a second similarity between the question keyword and the case field; and based on the second similarity or the user information, sorting the preliminary retrieval cases to obtain a final retrieval result. According to the method, the requirements of different users on case retrieval can be met, and the case retrieval effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the technical field of natural language processing. More specifically, the present invention relates to a method, a system and a computer-readable storage medium for case retrieval. Background Art

[0002] An enterprise case library is a software system for storing and managing enterprise case information. An enterprise case database usually contains relevant information such as the basic information of cases, the classification of cases, the description of cases, and the solutions to cases. It can help enterprises centrally collect, manage and share internal and external cases of the enterprise, and provide strong support for enterprise decision-making.

[0003] If it is necessary to retrieve cases in the enterprise case library, the traditional case retrieval method usually only retrieves cases according to the titles of cases, which results in a low correlation between the content in the retrieved cases and the retrieval problem, thereby resulting in a low accuracy of retrieving cases.

[0004] In view of this, there is an urgent need to provide a solution for case retrieval to improve the accuracy of retrieving cases. Summary of the Invention

[0005] In order to solve at least one or more of the above-mentioned technical problems, the present invention proposes a method, a system and a computer-readable storage medium for case retrieval in multiple aspects.

[0006] In a first aspect, the present invention provides a method for case retrieval, where the cases include multiple case entities and case fields, and the method includes: receiving a retrieval problem and user information, and obtaining valid information of the retrieval problem, where the valid information includes one or more problem keywords; obtaining a first similarity between the problem keywords and the case entities, and performing recall based on the first similarity to obtain preliminary retrieved cases; obtaining a second similarity between the problem keywords and the case fields; and sorting the preliminary retrieved cases based on the second similarity or the user information to obtain a final retrieval result.

[0007] In a second aspect, the present invention provides a system for case retrieval, including a processor and a memory, where the processor is configured to execute program instructions; the memory is configured to store the program instructions, and when the program instructions are loaded and executed by the processor, the system executes the method for case retrieval provided in the first aspect.

[0008] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program for case retrieval is stored. When the computer program is executed by a processor, the method for case retrieval provided in the first aspect is implemented.

[0009] Through the method for case retrieval provided above, in the embodiments of the present invention, the first similarity between the problem keywords in the retrieval problem and the case entities is extracted. Then, based on the first similarity, a recall is performed, and cases in which the problem keywords are associated with the case entities, that is, preliminary retrieval cases, can be retrieved to retrieve cases whose content is similar to the retrieval problem. Then, the similarity between the problem keywords and the case fields in the case, that is, the second similarity, is extracted, and based on the second similarity, the preliminary retrieval cases are sorted. Cases with a higher second similarity among the preliminary retrieval cases can be ranked at the front of the retrieval results, so that users can more quickly view cases that meet the retrieval problem, thereby improving the accuracy and efficiency of case retrieval. Moreover, sorting the preliminary retrieval cases based on user information can perform personalized sorting of the preliminary retrieval cases according to the relevant information of the user, which can meet the needs of different users for retrieved cases and improve the effect of retrieved cases. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0011] Figure 1 An exemplary flowchart of the method for case retrieval according to some embodiments of the present invention is shown;

[0012] Figure 2 An exemplary flowchart of the method for obtaining the first similarity according to some embodiments of the present invention is shown;

[0013] Figure 3 An exemplary flowchart of the method for obtaining the second similarity according to some embodiments of the present invention is shown;

[0014] Figure 4 An exemplary flowchart of the method for obtaining the third similarity according to some embodiments of the present invention is shown;

[0015] Figure 5 An exemplary flowchart of the method for obtaining the fourth similarity according to some embodiments of the present invention is shown;

[0016] Figure 6 An exemplary flowchart of the method for sorting the preliminary retrieval cases according to some embodiments of the present invention is shown;

[0017] Figure 7 An exemplary flowchart of a method for sorting preliminary search cases in some other embodiments of the present invention is shown;

[0018] Figure 8 An exemplary structural block diagram of a system for case retrieval according to an embodiment of the present invention is shown. Detailed implementation manners

[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] It should be understood that when terms such as "first", "second", "third" and "fourth" are used in the claims, the description and the drawings of the present invention, they are only used to distinguish different objects and not to describe a specific order. The terms "comprising" and "including" used in the description and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0021] It should also be understood that the terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the description and claims of the present invention, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the description and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0022] As used in this specification and the claims, the term "if" can be interpreted as "when", "once", "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" according to the context.

[0023] The specific implementation manners of the present invention will be described in detail below with reference to the accompanying drawings.

[0024] Exemplary application scenarios.

[0025] In a traditional case retrieval method, by treating cases as ordinary knowledge and retrieving cases only based on the titles of cases, the relevance between the content in the retrieved cases and the retrieval problem is low, resulting in low accuracy of retrieving cases.

[0026] Exemplary application scenarios.

[0027] In view of this, embodiments of the present invention provide a case retrieval scheme. By extracting the first similarity between the problem keywords in the retrieval problem and the case entities, and then, based on the first similarity, performing a recall to retrieve cases where the problem keywords are associated with the case entities, that is, preliminary retrieval cases, to retrieve cases whose content is similar to the retrieval problem; then, extracting the similarity between the problem keywords and the case fields in the cases, that is, the second similarity, and based on the second similarity, sorting the preliminary retrieval cases, the cases corresponding to higher second similarities among the preliminary retrieval cases can be ranked at the front of the retrieval results, enabling users to more quickly view cases that meet the retrieval problem, thereby improving the accuracy and efficiency of case retrieval; moreover, sorting the preliminary retrieval cases based on user information can perform personalized sorting of the preliminary retrieval cases according to the relevant information of the user, meeting the needs of different users for retrieved cases and improving the effect of retrieved cases.

[0028] In a first aspect, the present invention provides a method for case retrieval.

[0029] Figure 1 An exemplary flowchart of the method for case retrieval according to some embodiments of the present invention is shown.

[0030] As shown in the figure, in step S100, a retrieval problem and user information are received, and the effective information of the retrieval problem is obtained, where the effective information includes one or more problem keywords; in step S200, the first similarity between the problem keywords and the case entities is obtained, and a recall is performed based on the first similarity to obtain preliminary retrieval cases; in step S300, the second similarity between the problem keywords and the case fields is obtained; in step S400, based on the second similarity or user information, the preliminary retrieval cases are sorted to obtain the final retrieval result.

[0031] In an embodiment of the present disclosure, the above "retrieval question" refers to the question statement that a user wants to obtain an answer when retrieving cases. The retrieval question can be a sentence or one or more words or phrases. For example, the retrieval question can be "What are the successful marketing cases in the Internet industry?" A case includes multiple case entities and case fields. An entity refers to an identifiable thing in the real world, and an entity can be a person, an event, an object, or an abstract concept. For example, in an enterprise database, entities may include customers, products, orders, employees, etc. Here, the entities in the enterprise database are referred to as "case entities". The "case field" here refers to the database field used to store specific case information. The case field can include at least one of the work section, location, and time in the case. A work section refers to the construction organization form divided according to specific circumstances, and it is responsible for specific construction tasks.

[0032] Through step S100, after receiving the retrieval question and user information, the effective information of the retrieval question is obtained. The effective information here refers to various information related to the case in the retrieval question. For example, location, proper nouns, time, etc. The effective information includes one or more keywords in the retrieval question. The keywords in the retrieval question can be one or more characters. For the convenience of description, here the "keywords in the retrieval question" are referred to as "question keywords".

[0033] Next, in step S200, the similarity degree between the question keyword and the case entity is determined. Here, the "similarity degree between the question keyword and the case entity" is referred to as the "first similarity". The above first similarity can be the similarity between the text corresponding to the question keyword and the text corresponding to the case entity, or the cosine similarity between the vector corresponding to the question keyword and the vector corresponding to the case entity.

[0034] Moreover, based on the first similarity, the cases stored in the enterprise case library are recalled to obtain the cases where the question keyword is similar to the case entity. The "recall" here refers to finding one or more cases in the enterprise case library. The cases obtained by recalling the cases stored in the enterprise case library based on the first similarity are referred to as "preliminary retrieval cases". The above "the question keyword is similar to the case entity" means that the question keyword and the case entity are synonyms, represent the same concept, and the cosine similarity between the vector corresponding to the question keyword and the vector corresponding to the case entity is greater than a preset first threshold.

[0035] The above first threshold can be set by oneself. Those skilled in the art can understand that the setting method and specific value of the first threshold are not limited by any means. The first threshold is a critical value used to determine whether the cosine similarity between the vector corresponding to the question keyword and the vector corresponding to the case entity is large.

[0036] Specifically, the greater the cosine similarity between the vector corresponding to the problem keyword and the vector corresponding to the case entity, the higher the degree of association between the problem keyword and the case entity. If the cosine similarity between the vector corresponding to the problem keyword and the vector corresponding to the case entity is greater than the first threshold, it indicates that the internal connection between the problem keyword and the case entity is very close; conversely, the smaller the cosine similarity between the vector corresponding to the problem keyword and the vector corresponding to the case entity, the lower the degree of association between the problem keyword and the case entity. If the cosine similarity between the vector corresponding to the problem keyword and the vector corresponding to the case entity is less than the first threshold, it indicates that the degree of internal connection between the problem keyword and the case entity is very low.

[0037] For example, if the problem keyword is "electric energy meter", search for cases in the enterprise case library where the case entity is "electric energy meter" or "electric meter". Then, the cases corresponding to the case entity "electric energy meter" or "electric meter" are the preliminary retrieval cases.

[0038] After step S100, through step S200, cases related to the retrieval problem can be searched based on the degree of association between the problem keyword and the case entity.

[0039] Then, in step S300, the similarity between the problem keyword and the case field is extracted. Here, the similarity between the problem keyword and the case field is called the "second similarity". The second similarity can be the similarity between the text corresponding to the problem keyword and the text corresponding to the case field, or the cosine similarity between the vector corresponding to the problem keyword and the vector corresponding to the case field.

[0040] In some embodiments, step S400 is executed to sort the preliminary retrieval cases based on the second similarity. Here, the result after sorting the preliminary retrieval cases is called the "final retrieval result". Through step S400, based on the similarity between the problem keyword and the case field, the cases with a higher second similarity in the preliminary retrieval cases can be ranked at the front of the retrieval result, enabling the user to more quickly view the cases that meet the retrieval problem, thereby improving the accuracy and efficiency of case retrieval.

[0041] The above-mentioned "user information" refers to the relevant information of the person conducting the case retrieval, including but not limited to the entry time, department and position, and historical retrieval problems. The above-mentioned historical retrieval problems refer to the retrieval problems used by the user before.

[0042] In some other embodiments, through step S400, the preliminary retrieved cases can also be personalized sorted based on the user information, so that the cases in the preliminary retrieved cases that are more in line with the user's personal relevant information are ranked in the front, enabling the user to more quickly view the cases that meet the user's personal relevant information, thereby improving the accuracy and efficiency of case retrieval.

[0043] As can be seen from the description of the above embodiments, through steps S100 to S400, the first similarity between the problem keywords in the retrieval problem and the case entities is extracted. Then, based on the first similarity, recall is performed to retrieve the cases where the problem keywords are associated with the case entities, that is, the preliminary retrieved cases. Then, the similarity between the problem keywords and the case fields in the cases, that is, the second similarity, is extracted, and based on the second similarity, the preliminary retrieved cases are sorted, so that the cases corresponding to the higher second similarity in the preliminary retrieved cases are ranked in the front of the retrieval results, enabling the user to more quickly view the cases that meet the retrieval problem, thereby improving the accuracy and efficiency of case retrieval. Moreover, sorting the preliminary retrieved cases based on the user information can perform personalized sorting of the preliminary retrieved cases according to the user's relevant information, which can meet the needs of different users for retrieved cases and improve the effect of retrieved cases.

[0044] In some embodiments, the case entities in the cases can be extracted through a Named Entity Recognition (NER) model. The NER model is trained in the following manner: First, the case entity samples in the case samples are labeled according to the case entity labels to obtain case-labeled samples. Here, the case samples refer to the cases used to train the NER model, and the case entities in the case samples are called "case entity samples". The "case entity label" refers to the label corresponding to the case entity sample, and each case entity sample has a case entity label. For example, if the case entity sample is cold medicine, the case entity label corresponding to this case entity sample is cold medicine.

[0045] In some embodiments of the present disclosure, the case entity samples in the case samples can be labeled by the BIO (Begin Inside Outside) annotation method. Those skilled in the art can understand that the BIO annotation method is a sequence annotation method in natural language processing. B- represents the starting position of the case entity, I represents the part inside the case entity, and O- represents that the word is not part of any entity. For example, if the case entity sample is "tumor hospital", then "tumor" is labeled as B-NP and "hospital" is labeled as I-NP.

[0046] After annotating the case entity samples in the case samples, the case entities in the case samples are extracted through the BERT (Bidirectional Encoder Representation from Transformers) model. Here, the case entities extracted from the case samples by the BERT model are called case training entities, and the BERT model is tuned with the case annotation samples and case training entities to train the BERT model. Finally, the trained BERT model is used as the NER model. Those skilled in the art can understand that the BERT model refers to a machine learning model based on the Transformer architecture.

[0047] In some embodiments, the case fields include at least one of work section, location, and time. The embodiments of the present disclosure can extract the case fields in the cases. For example, when a work section appears in a case, the work section is extracted and used as the case field of the case.

[0048] Figure 2 An exemplary flowchart of the method for obtaining the first similarity in some embodiments of the present invention is shown. It can be understood that the method for obtaining the first similarity is a specific implementation in the foregoing step S200. Therefore, the features described above in conjunction with Figure 1 can be similarly applied herein.

[0049] As shown in the figure, in step S210, the question keyword is converted into a question keyword vector, and the case entity is converted into a case entity vector; in step S220, the similarity between the question keyword vector and the case entity vector is calculated to obtain the first similarity.

[0050] Through step S210, the question keyword is converted from text into a corresponding vector. Here, the vector corresponding to the question keyword is called the "question keyword vector". Specifically, the question keyword can be converted into a corresponding word vector, that is, the question keyword vector, through an embedding layer.

[0051] Similarly, according to the specific manner of converting the question keyword into the question keyword vector, the case entity is converted into the case entity vector, which will not be elaborated here.

[0052] After step S210, step S220 is executed to calculate the cosine similarity between the problem keyword vector and the case entity vector. Here, the "cosine similarity between the problem keyword vector and the case entity vector" is referred to as the first similarity. This first similarity is used to represent the internal connection or correlation degree between the problem keyword and the case entity. If the first similarity is larger, the correlation degree between the problem keyword and the case entity is higher, and the correlation degree between the retrieval problem and the case is higher; conversely, if the first similarity is smaller, the correlation degree between the problem keyword and the case entity is lower, and the correlation degree between the retrieval problem and the case is lower.

[0053] In some other embodiments, the Jaccard Similarity between the problem keyword and the case entity can also be calculated, and the Jaccard Similarity between the problem keyword and the case entity is used as the first similarity. If the Jaccard Similarity between the problem keyword and the case entity is higher, it indicates that the similarity between the retrieval problem and the case is higher; if the Jaccard Similarity between the problem keyword and the case entity is lower, it indicates that the similarity between the retrieval problem and the case is lower.

[0054] Through steps S210 to S220, the first similarity between the problem keyword and the case entity can be determined, so as to subsequently find the cases associated with the problem in the enterprise case library according to the first similarity.

[0055] In some embodiments, the foregoing step S200 for recall based on the first similarity to obtain the preliminary retrieval cases specifically includes: in response to the first similarity being greater than a preset first threshold, taking the current case as the preliminary retrieval case.

[0056] If there is any one or more cases where the first similarity between the case entity of the case and the problem keyword is greater than the first threshold, it indicates that the internal connection between the problem keyword and the case entity is very close, that is, the correlation degree between the retrieval problem and the case is relatively high. At this time, this case is taken as the preliminary retrieval case. The number of cases in this preliminary retrieval case is zero or a positive integer. In this way, the cases with a relatively high correlation degree with the retrieval problem can be retrieved, avoiding extracting cases irrelevant to the retrieval problem, thereby improving the accuracy and effectiveness of case retrieval.

[0057] In some other embodiments, step S200 may further include: performing recall based on the first similarity through Elasticsearch to obtain the preliminary retrieval cases. Here, Elasticsearch refers to an open-source search engine that provides a distributed, multi-tenant full-text search engine, capable of achieving fast search and being able to scale to hundreds or even thousands of servers and structured or unstructured data.

[0058] Figure 3 An exemplary flowchart of a method for obtaining a second similarity in some embodiments of the present invention is shown. It can be understood that the method for obtaining the second similarity is a specific implementation in the foregoing step S300. Therefore, the features described above in conjunction with Figure 1 can be similarly applied herein.

[0059] As shown in the figure, in step S310, a third similarity between the current problem keyword and the case field is obtained; in step S320, a fourth similarity between the historical search keyword and the case field is obtained.

[0060] In the embodiments of the present disclosure, the second similarity includes the third similarity and the fourth similarity. The problem keyword includes the user's current problem keyword and historical search keyword. The "current problem keyword" refers to the keyword in the user's current retrieval problem, and the "historical search keyword" refers to the keyword in the retrieval problems used by the user previously. The "third similarity" refers to the degree of proximity between the current problem keyword and the case field, and the "fourth similarity" is the degree of proximity between the historical search keyword and the case field.

[0061] It should be noted that the third similarity can be the similarity between the text corresponding to the current problem keyword and the text corresponding to the case entity, or the cosine similarity between the vector corresponding to the current problem keyword and the vector corresponding to the case entity. The fourth similarity can be the similarity between the text corresponding to the historical search keyword and the text corresponding to the case entity, or the cosine similarity between the vector corresponding to the historical search keyword and the vector corresponding to the case entity.

[0062] Moreover, the manner of obtaining the third similarity and the manner of obtaining the fourth similarity in step S310 can be the same as the manner of obtaining the first similarity in the foregoing step S200.

[0063] Through step S310 and step S320, the relevant connection between the user's current retrieval problem and the case, and the relevant connection between the retrieval problems retrieved by the user previously and the case can be extracted, that is, the third similarity and the fourth similarity are obtained. Thus, in the subsequent process, the cases retrieved can be sorted according to the relevant connections between the retrieval problems at different times and the case respectively, so as to retrieve the common cases that more conform to the user's field and historical retrieval information, thereby improving the accuracy of retrieval.

[0064] In some embodiments, the third similarity includes a fifth similarity and a sixth similarity. Figure 4FIG. shows an exemplary flowchart of a method for obtaining a third similarity in some embodiments of the present invention. It can be understood that the method for obtaining the third similarity is a specific implementation in the foregoing step S310. Therefore, the features described in combination with the foregoing Figure 1 and Figure 3 can be similarly applied herein.

[0065] As shown in the figure, in step S311, the current problem keyword is converted into a current problem keyword vector, and the case field is converted into a case field vector; in step S312, the similarity between the current problem keyword vector and the case field vector is calculated to obtain a fifth similarity; in step S313, the text similarity between the current problem keyword and the case field is calculated to obtain a sixth similarity.

[0066] In some embodiments, the case field includes the title, content summary, analysis, etc. of the case.

[0067] The "current problem keyword vector" here refers to the vector corresponding to the current problem keyword. In step S311, the word2vec model is used to convert the current problem keyword into a current problem keyword vector, and the word2vec model is used to convert the case field into a case field vector. Those skilled in the art can understand that the word2vec model is a model in the field of natural language processing for mapping words into a vector space to obtain word vectors.

[0068] In some other embodiments, the manner of converting the current problem keyword into a current problem keyword vector in step S311 is the same as the specific manner of converting the problem keyword into a problem keyword vector in the foregoing step S210, and the specific manner of converting the case field into a case field vector in step S311 is the same as the specific manner of converting the case field into a case field vector in the foregoing step S210, which will not be elaborated herein.

[0069] Moreover, the specific manner of calculating the similarity between the current problem keyword vector and the case field vector in step S312 is the same as the specific manner of calculating the similarity between the problem keyword vector and the case entity vector in the foregoing step S220.

[0070] The above-mentioned "fifth similarity" refers to the similarity between the current problem keyword vector and the case field vector, which is used to measure the degree of association between the current problem keyword vector and the case field vector. If the fifth similarity is greater, it indicates a higher degree of association between the current problem keyword vector and the case field vector, that is, a higher degree of association between the current problem keyword and the case field; if the fifth similarity is smaller, it indicates a lower degree of association between the current problem keyword vector and the case field vector, that is, a lower degree of association between the current problem keyword and the case field.

[0071] In some embodiments, the text similarity in step S313 refers to the Jaccard Similarity. Through step S313, the Jaccard similarity between the current problem keyword and the case field is calculated, and the Jaccard similarity between the current problem keyword and the case entity is used as the sixth similarity. If the sixth similarity is greater, it indicates a higher similarity between the current retrieval problem and the case; if the sixth similarity is smaller, it indicates a lower similarity between the current retrieval problem and the case.

[0072] It should be noted that step S312 can be executed first, followed by step S313, or step S313 can be executed first, followed by step S312.

[0073] Through steps S311 to S313, the similarity between the current problem keyword vector and the case field vector can be obtained, as well as the text similarity between the current problem keyword vector and the case field vector. From the perspective of word vectors and text, the relationship between the current problem keyword and the fields such as the title, content summary, and analysis of the case in the case is extracted, so as to obtain the internal connection between the current retrieval problem and the case. Furthermore, it is convenient to sort the retrieved cases according to the internal connection between the current retrieval problem and the case later, and rank the cases that more conform to the current retrieval problem in front of the case retrieval results. In this way, users can view the cases that meet the expectations more quickly, thereby improving the accuracy and efficiency of case retrieval.

[0074] Similarly, Figure 5 FIG. shows an exemplary flowchart of a method for obtaining the fourth similarity according to some embodiments of the present invention. It can be understood that the method for obtaining the fourth similarity is a specific implementation of the foregoing step S320. Therefore, the features described in the foregoing Figure 1 and Figure 3 can be similarly applied herein. The fourth similarity includes a seventh similarity and an eighth similarity.

[0075] As shown in the figure, in step S321, the historical search keywords are converted into a historical search keyword vector; in step S322, the similarity between the historical search keyword vector and the case field vector is calculated to obtain a seventh similarity; in step S323, the text similarity between the historical search keywords and the case fields is calculated to obtain an eighth similarity.

[0076] The above-mentioned "historical search keyword vector" refers to the vector corresponding to the historical search keywords. In step S321, the word2vec model is used to convert the historical search keywords into a historical search keyword vector.

[0077] After converting the historical search keywords into a historical search keyword vector, the similarity between the historical search keyword vector and the case field vector is calculated through step S322. For the sake of description, the "seventh similarity" here refers to the similarity between the historical search keyword vector and the case field vector.

[0078] In step S323, the text similarity between the historical search keywords and the case fields is calculated. In the embodiments of the present disclosure, the text similarity in step S23 is the Jaccard similarity, and the "eighth similarity" refers to the Jaccard similarity between the historical search keywords and the case fields.

[0079] It should be noted that step S322 can be executed first, followed by step S323, or step S323 can be executed first, followed by step S322.

[0080] Through steps S321 to S323, the similarity between the historical search keyword vector and the case field vector can be obtained, and the text similarity between the historical search keyword vector and the case field vector can be obtained. The relationship between the historical search keywords and the fields such as the title, content summary, and analysis of the cases in the cases is extracted from the perspective of word vectors and the perspective of text, so as to obtain the internal connection between the retrieval problems of the user before and the cases, and then facilitate subsequent sorting of the retrieved cases according to the internal connection between the historical search keywords and the cases. In this way, the user's historical search records can be fully considered, and the cases that are more in line with the historical search problems can be ranked in front of the case retrieval results, so that the user can view the cases that meet the expectations more quickly, thereby improving the accuracy and efficiency of case retrieval.

[0081] Figure 6 An exemplary flowchart of a method for sorting preliminary retrieved cases according to some embodiments of the present invention is shown. It can be understood that the method for sorting preliminary retrieved cases is a specific implementation of the foregoing step S400, and therefore the features described above in conjunction with Figure 1 can be similarly applied herein.

[0082] As shown in the figure, in step S410, the case field score of each case field is determined according to the second similarity; in step S420, according to the weight of each field in each preliminary retrieved case, the case field scores are weighted and summed to obtain the case score corresponding to each preliminary retrieved case; in step S430, the preliminary retrieved cases are sorted based on the case scores to obtain the final retrieval result.

[0083] In some embodiments, the case fields include at least one of the title, summary, and analysis of the case. Each case field has a corresponding field weight, and the field weight is used to represent the importance of the field in the case. The value range of the field weight is greater than zero and less than 1. For example, the field weight corresponding to the title can be 0.3, the field weight corresponding to the summary can be 0.5, and the field weight corresponding to the analysis can be 0.2. Those skilled in the art can understand that the specific content of the case fields and the specific values of the field weights are not limited in any way.

[0084] In step S410, the case field score of each case field is determined according to the second similarity. Specifically, if the second similarity is greater than or equal to the second threshold, the case field score of this case field is set to 1; if the second similarity is less than the second threshold, the case field score of this case field is set to 0.

[0085] The above-mentioned "second threshold" is used to judge the similarity between the problem keyword and the case field, that is, a critical value for whether the second similarity is large. If the second similarity is greater than or equal to the second threshold, it indicates that the similarity between the problem keyword and this case field is large. On the contrary, if the second similarity is less than the second threshold, it indicates that the similarity between the problem keyword and this case field is small. Those skilled in the art can understand that the above-mentioned second threshold can be set by themselves, and the setting method and specific value of the second threshold are not limited in any way.

[0086] In the embodiments of the present disclosure, the above-mentioned "case field score" is used to measure the similarity between the case field and the problem keyword in the retrieval problem. Through step S410, the similarity between the problem keyword and the case field is converted into a fixed numerical range, and the value of the case field score is 0 or 1.

[0087] Next, in step S420, each case field score is multiplied by the corresponding case field weight, and the products obtained from each multiplication operation are added together. Here, the sum obtained is called the "case score". The case score is used to measure the degree of association between the preliminary retrieved case and the retrieval problem. If the case score is higher, it indicates that the degree of association between the preliminary retrieved case and the retrieval problem is higher; if the case score is lower, it indicates that the degree of association between the preliminary retrieved case and the retrieval problem is lower.

[0088] Then, in step S430, the preliminary retrieved cases are sorted according to the magnitudes of the case scores. Specifically, the case scores are sorted from largest to smallest, so as to arrange the preliminary retrieved cases corresponding to the case scores. For example, the case score corresponding to the preliminary retrieved case a1 is 0.85, and the case score corresponding to the preliminary retrieved case a2 is 0.6. Then, the preliminary retrieved case a1 is ranked in front of the preliminary retrieved case a2.

[0089] Through steps S410 to S430, the preliminary retrieved cases are scored according to the second similarity to quantify the degree of association between the preliminary retrieved cases and the retrieval problem. Then, the preliminary retrieved cases with higher scores are ranked in front, so that the preliminary retrieved cases with a higher degree of association with the retrieval problem can be extracted and ranked in front, enabling the user to view the cases that meet the retrieval problem more quickly, and improving the accuracy and efficiency of retrieving cases.

[0090] It should be noted that if the user retrieves data for the first time, there is no relevant data on the user's previous retrievals, such as the cases retrieved by the user previously. Therefore, case retrieval cannot be performed based on the relevant data on the user's previous retrievals. In the embodiments of the present disclosure, it is not necessary to obtain the cases retrieved by the user previously. Instead, the preliminary retrieved cases are scored according to the second similarity to quantify the degree of association between the preliminary retrieved cases and the retrieval problem. In this way, the preliminary retrieved cases with higher scores can be ranked in front, so that the preliminary retrieved cases with a higher degree of association with the retrieval problem can be extracted and ranked in front, enabling the user to view the cases that meet the retrieval problem more quickly, and improving the accuracy and efficiency of retrieving cases.

[0091] In some embodiments, the user information includes at least one of user qualifications, user positions, and user departments. Figure 7 An exemplary flowchart of a method for sorting preliminary retrieved cases according to other embodiments of the present invention is shown. It can be understood that the method for sorting preliminary retrieved cases is a specific implementation in the foregoing step S400. Therefore, the features described above in conjunction with Figure 1 can be similarly applied herein.

[0092] As shown in the figure, in step S440, each preliminary retrieved case is scored according to the user information to obtain the confidence level corresponding to each preliminary retrieved case; in step S450, the preliminary retrieved cases are sorted based on the confidence level to obtain the final retrieval result.

[0093] In step S440, the trained scoring model scores each preliminary retrieval case according to the user information. Here, the scoring model refers to a machine learning model used to score the preliminary retrieval cases according to the user information. The scoring model determines whether the user clicks on the preliminary retrieval case according to the user information, and the degree of certainty of whether the user clicks on the preliminary retrieval case is called the confidence level corresponding to the preliminary retrieval case. Specifically, the user information and the preliminary retrieval case are input into the scoring model, and the scoring model outputs the degree of certainty of whether the user clicks on the preliminary retrieval case, that is, the confidence level.

[0094] In some embodiments, the scoring model can be a Light Gradient Boosting Machine (LightGBM for short). LightGBM is a tree-based ensemble learning method that adopts the gradient boosting technique. By combining multiple weak learners into a powerful model, it will not be elaborated here.

[0095] It should be noted that the method of training the above scoring model can be any one in the art and is not subject to any limitation.

[0096] After step S440, in step S450, the confidence levels are sorted from largest to smallest, so as to sort the preliminary retrieval cases. The result after sorting is called the "final retrieval result".

[0097] Through steps S440 to S450, sorting the preliminary retrieval cases based on the user information can perform personalized sorting of the preliminary retrieval cases according to the relevant information of the user, can meet the needs of different users for retrieval cases, and improve the effect of the retrieval cases.

[0098] In summary, in the embodiment of the present invention, by extracting the first similarity between the problem keyword in the retrieval problem and the case entity, then, based on the first similarity for recall, cases in which the problem keyword is associated with the case entity can be retrieved, that is, preliminary retrieval cases, to retrieve cases whose content is similar to the retrieval problem; then, the similarity between the problem keyword and the case field in the case is extracted, that is, the second similarity, and based on the second similarity, the preliminary retrieval cases are sorted, so that the cases corresponding to the higher second similarity in the preliminary retrieval cases are ranked in the front of the retrieval results, so that the user can more quickly view the cases that meet the retrieval problem, thereby improving the accuracy and efficiency of case retrieval; moreover, sorting the preliminary retrieval cases based on the user information can perform personalized sorting of the preliminary retrieval cases according to the relevant information of the user, can meet the needs of different users for retrieval cases, and improve the effect of the retrieval cases.

[0099] In a second aspect, the present invention provides a system 800 for case retrieval.Figure 8 FIG. shows an exemplary structural block diagram of a system for case retrieval according to an embodiment of the present invention. As shown in the figure, the system 800 for case retrieval includes a processor 810 and a memory 820. The processor 810 is configured to execute program instructions; the memory 820 is configured to store the program instructions, and when the program instructions are loaded and executed by the processor, the system is caused to execute the method for case retrieval provided in the first aspect of the present invention.

[0100] In a third aspect, the present invention provides a computer-readable storage medium having stored thereon a computer program for case retrieval, and when the computer program is executed by a processor, the method for case retrieval provided in the first aspect of the present invention is implemented.

[0101] Although multiple embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many changes, variations, and alternative methods may occur to those skilled in the art without departing from the spirit and scope of the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the present invention. The appended claims are intended to define the scope of the present invention and thus cover equivalents or alternatives within the scope of these claims.

Claims

1. A method for case retrieval, characterized in that, The case includes multiple case entities and case fields, and the method includes: Receiving a retrieval question and user information, and obtaining valid information of the retrieval question, where the valid information includes one or more question keywords; Obtaining a first similarity between the question keyword and the case entity, and performing recall based on the first similarity to obtain a preliminary retrieval case; Obtaining a second similarity between the question keyword and the case field; and Sorting the preliminary retrieval cases based on the second similarity or the user information to obtain a final retrieval result.

2. The method according to claim 1, wherein The obtaining the first similarity between the question keyword and the case entity includes: Converting the question keyword into a question keyword vector, and converting the case entity into a case entity vector; Calculating the similarity between the question keyword vector and the case entity vector to obtain the first similarity.

3. The method according to claim 1, wherein The performing recall based on the first similarity to obtain a preliminary retrieval case includes: In response to the first similarity being greater than a preset first threshold, taking the current case as the preliminary retrieval case.

4. The method according to claim 1, wherein The question keyword includes the user's current question keyword and historical search keywords, the second similarity includes a third similarity and a fourth similarity, and the obtaining the second similarity between the question keyword and the case field includes: Obtaining the third similarity between the current question keyword and the case field; Obtaining the fourth similarity between the historical search keyword and the case field; where the case field includes at least one of the title, summary, and analysis of the case.

5. The method according to claim 4, characterized in that The third similarity includes a fifth similarity and a sixth similarity, and the obtaining the third similarity between the current question keyword and the case field includes: Converting the current question keyword into a current question keyword vector, and Converting the case field into a case field vector; Calculating the similarity between the current question keyword vector and the case field vector to obtain the fifth similarity; Calculating the text similarity between the current question keyword and the case field to obtain the sixth similarity.

6. The method according to claim 5, characterized in that, The fourth similarity includes a seventh similarity and an eighth similarity, and the obtaining the fourth similarity between the historical search keyword and the case field includes: Converting the historical search keyword into a historical search keyword vector; Calculating the similarity between the historical search keyword vector and the case field vector to obtain the seventh similarity; Calculating the text similarity between the historical search keyword and the case field to obtain the eighth similarity.

7. The method according to claim 4, characterized in that The case field includes at least one of the title, summary, and analysis of the case, and each case field has a corresponding field weight. Wherein, sorting the preliminary retrieval cases based on the second similarity to obtain a final retrieval result includes: Determining a case field score for each case field according to each second similarity; Perform a weighted summation of the scores of each case field according to the weight of each field in each of the preliminary retrieval cases to obtain a case score corresponding to each preliminary retrieval case; Sort the preliminary retrieval cases based on the case scores to obtain a final retrieval result.

8. The method according to claim 4, characterized in that, The user information includes at least one of user qualifications, user positions, and user departments. Among them, sorting the preliminary retrieval cases based on the user information to obtain a final retrieval result includes: Score each of the preliminary retrieval cases according to the user information to obtain a confidence level corresponding to each preliminary retrieval case; Sort the preliminary retrieval cases based on the confidence level to obtain a final retrieval result.

9. A system for case retrieval, characterized in that, It includes a processor and a memory, where The processor is configured to execute program instructions; the memory is configured to store the program instructions. When the program instructions are loaded and executed by the processor, the system executes the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, A computer program for case retrieval is stored on the storage medium. When the computer program is executed by a processor, the method for case retrieval according to any one of claims 1-8 is implemented.