Case Relevance Detection Method and Terminal Device

By constructing the structured data word vector matrix of the case and determining the center point, the traditional case recommendation system is solved inadequate accuracy under the problem of data sparseness, and more accurate case recommendations and lower computing resource requirements are achieved.

CN111966924BActive Publication Date: 2025-06-27PINGAN ZHITONG CONSULT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011021188.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-25
Publication Date
2025-06-27
Estimated Expiration
2040-09-25

AI Technical Summary

Technical Problem

Traditional case-like recommendation systems have sparse data problems when the data volume is large, making it difficult to obtain accurate and effective recommendations.

Method used

By extracting the structured data of cases that meet the preset requirements in the case library, converting them into word vectors, constructing vector matrix, determining the center point of the case collection, and calculating the correlation between the case and the case collection based on the center point, cases with high correlation are recommended.

Benefits of technology

It realizes more accurate case recommendations, avoids the problem of sparse data, does not need to set up data burial points, only daily office data can be used to recommend, and the online computing resources are required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111966924B_ABST
    Figure CN111966924B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of information processing technology, and provides a method, device, and terminal device for detecting case relevance. The method for detecting case relevance includes: extracting multiple structured data of each first case that meets preset requirements in a case library, and all the first cases form a target case set; converting each structured data of each first case into a corresponding word vector to obtain a first vector matrix corresponding to the target case set; determining the center point of the target case set in a multi-dimensional vector space based on the first vector matrix; determining a second vector matrix of each second case in the case library, and determining the relevance between each second case and the target case set based on the second vector matrix and the center point, where the second case is a case in the case library other than the target case set; and sending the link information and / or relevant files of the second cases whose relevance meets the preset requirements to a user terminal. The case relevance detection result of this application is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of information processing, and particularly relates to a method for detecting case relevance and a terminal device. Background Art

[0002] Case recommendation is a technology that recommends specific cases to users through methods such as data filtering and information retrieval, and can screen out relevant cases from a large number of cases. Traditional case recommendation usually clusters based on static data, dynamic data, or rules, and recommends similar cases according to user preferences through algorithms such as collaborative filtering. However, algorithms such as collaborative filtering have serious data sparsity problems in recommendation systems with a large amount of data. The number of user views only accounts for a very small part of the total content, and most content has no click data, resulting in very little intersection between different users. Therefore, it is difficult to obtain accurate and effective recommendations. Summary of the Invention

[0003] To overcome the problems existing in the related art, embodiments of this application provide a method for detecting case relevance and a terminal device.

[0004] This application is implemented through the following technical solutions:

[0005] In a first aspect, an embodiment of this application provides a method for detecting case relevance, including:

[0006] Extracting multiple structured data of each first case that meets the preset requirements in the case library; wherein, all the first cases constitute a target case set;

[0007] Converting each structured data of each first case into a corresponding word vector to obtain a first vector matrix corresponding to the target case set;

[0008] Determining the center point of the target case set in the multi-dimensional vector space based on the first vector matrix;

[0009] Determining a second vector matrix of each second case in the case library, and determining the relevance between each second case and the target case set based on the second vector matrix and the center point; wherein, the second case is a case in the case library other than the target case set;

[0010] Sending the link information and / or relevant files of the second cases whose relevance meets the preset requirements to the user terminal.

[0011] In a possible implementation manner of the first aspect, the extracting multiple structured data of each first case that meets the preset requirements in the case library includes:

[0012] Annotate the training text data; among them, the annotated information includes multiple types of information;

[0013] Input the annotated training text data into the target network to determine the loss functions corresponding to various types of information;

[0014] Train the target network based on the loss functions corresponding to various types of information;

[0015] Input the text of each first case into the trained target network to extract the structured data.

[0016] In a possible implementation manner of the first aspect, the target network includes multiple sub-networks, and each sub-network corresponds to one type of information;

[0017] The training of the target network based on the loss functions corresponding to various types of information includes:

[0018] Train the corresponding sub-network based on the loss functions corresponding to various types of information.

[0019] In a possible implementation manner of the first aspect, the training of the target network based on the loss functions corresponding to various types of information includes:

[0020] Calculate the proportion of the number of texts corresponding to various types of information;

[0021] Determine the total loss function according to the sum of the products of each loss function and the corresponding proportion;

[0022] Train the target network based on the total loss function.

[0023] In a possible implementation manner of the first aspect, the extraction of multiple structured data of each first case in the target case set includes:

[0024] Determine a regular expression according to the element to be extracted itself and the context information related to the element itself;

[0025] Divide the text of the first case into multiple single sentences, and determine the target single sentences that match the regular expression among the multiple single sentences;

[0026] Extract structured data from the target single sentences.

[0027] In a possible implementation manner of the first aspect, the extraction of structured data from the target single sentences includes:

[0028] Perform field division on the target single sentences; among them, each target single sentence is divided into at least one field, and each field corresponds to a field name and a field value;

[0029] For each relevant field with the same field content but different field names or field values, normalize the field names or field values;

[0030] Extract structured data from each of the fields after the normalization process.

[0031] In a possible implementation of the first aspect, the structured data is text data or a numerical value, and the converting each structured data of each of the first cases into a corresponding word vector to obtain a first vector matrix corresponding to the target case set includes:

[0032] Convert the text data in each structured data of each of the first cases into word vectors;

[0033] Use the word vectors and numerical values of each of the first cases as elements of the first vector matrix; wherein, each of the first cases corresponds to a row of elements or a column of elements.

[0034] In a possible implementation of the first aspect, the determining the center point of the target case set in the multi-dimensional vector space based on the first vector matrix includes:

[0035] Calculate the mean vector of the word vectors corresponding to the same structured data of each of the first cases; wherein, each of the mean vectors is the vector of the center point;

[0036] The determining the relevance between each of the second cases and the target case set based on the second vector matrix and the center point includes:

[0037] Determine the distance between each of the second cases and the center point according to the vector of the center point and each of the second vector matrices;

[0038] Determine the relevance between each of the second cases and the target case set based on each of the distances.

[0039] In a second aspect, an embodiment of the present application provides a case relevance detection device, including:

[0040] A structured data extraction module, configured to extract multiple structured data of each of the first cases that meet preset requirements in a case library; wherein, all of the first cases constitute a target case set;

[0041] A vector matrix generation module, configured to convert each structured data of each of the first cases into a corresponding word vector to obtain a first vector matrix corresponding to the target case set;

[0042] A center point determination module, configured to determine the center point of the target case set in the multi-dimensional vector space based on the first vector matrix;

[0043] A relevance determination module, configured to determine the second vector matrix of each second case in the case library, and determine the relevance between each second case and the target case set based on the second vector matrix and the center point; wherein, the second case is a case in the case library other than the target case set;

[0044] A sending module, configured to send the link information and / or relevant documents of the second cases whose relevance meets the preset requirements to the user terminal.

[0045] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the case relevance detection method according to any one of the first aspects is implemented.

[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the case relevance detection method according to any one of the first aspects is implemented.

[0047] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute the case relevance detection method according to any one of the first aspects described above.

[0048] It can be understood that the beneficial effects of the above second aspect to the fifth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here.

[0049] The beneficial effects of the embodiments of the present application compared with the prior art are:

[0050] Extract multiple structured data of each first case in the target case set, convert the structured data of each first case into corresponding word vectors to obtain a first vector matrix corresponding to the target case set, determine the center point of the target case set in the multi-dimensional vector space based on the first vector matrix, determine the second vector matrix of each second case in the case library, and determine the relevance between each second case and the target case set based on the second vector matrix and the center point, so as to send the link information and / or relevant documents of the second cases whose relevance meets the preset requirements to the user terminal. In the embodiment of the present application, the case text is converted into a vector, and the similar cases to be sent to the user terminal are determined through the relevance between the second case and the target case set, which is more accurate than the traditional case relevance detection result, and there is no need to set data buried points. Only the regular data in daily office work is required to determine the target case set for similar case recommendation, and the center point of the first vector matrix can be completed offline. Only the relevance between the recommended cases and the center point needs to be calculated online, and the requirement for online computing resources is relatively low.

[0051] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. Brief Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a schematic diagram of the application scenario of the case relevance detection method provided by an embodiment of the present application;

[0054] Figure 2 It is a schematic flowchart of the case relevance detection method provided by an embodiment of the present application;

[0055] Figure 3 It is a schematic flowchart of the case relevance detection method provided by an embodiment of the present application;

[0056] Figure 4 It is a schematic flowchart of the case relevance detection method provided by an embodiment of the present application;

[0057] Figure 5 It is a schematic flowchart of the case relevance detection method provided by an embodiment of the present application;

[0058] Figure 6 It is a schematic flowchart of the case relevance detection method provided by an embodiment of the present application;

[0059] Figure 7 It is a schematic flowchart of a case relevance detection method provided by an embodiment of the present application;

[0060] Figure 8 It is a schematic structural diagram of a case relevance detection device provided by an embodiment of the present application;

[0061] Figure 9 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0062] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0063] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0064] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0065] As used in the specification of the present application and the appended claims, the term "if" may be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" may be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.

[0066] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0067] References to "an embodiment" or "some embodiments" or the like described in the specification of the present application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0068] Case recommendation is a technology that recommends specific cases to users through methods such as data filtering and information retrieval, and can screen out relevant cases from a large number of cases. Traditional case recommendation usually clusters based on static data, dynamic data, or rules, and recommends similar cases according to user preferences through algorithms such as collaborative filtering. However, algorithms such as collaborative filtering have serious data sparsity problems in recommendation systems with a large amount of data. The user browsing volume only accounts for a very small part of the total content volume, and most content has no click data, resulting in very few intersections between different users. Therefore, it is very difficult to obtain accurate and effective recommendations.

[0069] Based on the above problems, the case relevance detection method in the embodiments of the present application extracts multiple structured data of each case in the target case set, converts each structured data of each case into a corresponding word vector to obtain a first vector matrix corresponding to the target case set, determines the center point of the target case set in the multi-dimensional vector space based on this first vector matrix, determines the second vector matrix of each case in the case library, and determines the relevance between each case in the case library and the target case set based on the second vector matrix and the center point. Thus, the link information and / or relevant documents of the cases whose relevance meets the preset requirements are sent to the user terminal. The embodiments of the present application convert case texts into vectors and determine the similar cases to be sent to the user terminal based on the relevance between each case in the case library and the target case set. Compared with traditional case relevance detection, the results are more accurate. Moreover, there is no need to set data points, and only the regular data in daily office work is required to determine the target case set for case recommendation. Moreover, the center point of the first vector matrix can be completed offline, and only the relevance between the recommended cases and the center point needs to be calculated online, which requires lower online computing resources.

[0070] For example, the embodiments of the present application can be applied to, such as Figure 1In the exemplary scenario shown. In this scenario, a user (such as a judge or a lawyer, etc.) can work through the terminal 10, and the server 20 can obtain the user's work data with the user's authorization and then determine the target case set, which may include the cases handled by the user or the cases selected by the user. After obtaining the target case set, the server 20 determines the similar cases in the case library with a relatively high relevance to the target case set and sends them to the terminal 10 for recommendation to the user.

[0071] Specifically, the server 20 can extract multiple structured data of each first case in the target case set, convert the structured data of each first case into corresponding word vectors to obtain the first vector matrix corresponding to the target case set, and determine the center point of the target case set in the multi-dimensional vector space based on the first vector matrix; after the user designates the case library through the terminal 10, the server 20 determines the second vector matrix of each second case in the case library, and determines the relevance of each second case to the target case set based on the second vector matrix and the above center point, and sends the link information and / or relevant files of the second cases whose relevance meets the preset requirements to the terminal 10, so as to recommend them to the user.

[0072] Among them, the process of the server 20 determining the center point of the target case set in the multi-dimensional vector space can be completed offline. After the user designates the case library through the terminal 10, the server 20 can directly determine the relevance of each second case in the case library to the target case set according to the previously determined center point, and send the link information and / or relevant files of the second cases whose relevance meets the preset requirements to the terminal 10.

[0073] The following combines Figure 1 to elaborate in detail on the case relevance detection method of the present application.

[0074] Figure 2 is a schematic flowchart of the case relevance detection method provided by an embodiment of the present application. Referring to Figure 2 the detailed description of the case relevance detection method is as follows:

[0075] In step 101, multiple structured data of each first case that meets the preset requirements in the case library are extracted; among them, all the first cases constitute the target case set.

[0076] In this step, the above target case set can be a set of cases handled by the user or a set of cases selected by the user in the case library, and the embodiments of the present application do not limit this. Specifically, the target case set may include multiple first cases, and each first case may be a case handled by the user or a case selected by the user in the case library.

[0077] The above structured data can be the basic information of a case, for example, it can include fields (such as the name of the plaintiff and / or the name of the defendant), amount (such as the amount involved in the case and / or the compensation amount), year, cause of action, the court handling the case, etc. Additionally, the above structured data can also be information extracted from unstructured texts related to the case (such as documents like the complaint, the plaintiff's evidence list, etc.).

[0078] In a possible implementation manner, the structured data of each first case can be extracted through a neural network. Specifically, referring to Figure 3 , step 101 can include the following steps:

[0079] In step 201, the training text data is annotated.

[0080] Among them, the annotated information can include multiple types of information. Specifically, the information to be annotated can be set according to the above structured data. For example, the annotated information can include name, amount, year, cause of action, the court handling the case, etc.

[0081] In step 202, the annotated training text data is input into the target network to determine the loss functions corresponding to various types of information.

[0082] Among them, for each training text data, information such as the name, amount, year, cause of action, the court handling the case, etc. in the training text data can be annotated. Therefore, each training text data can be input into the target network to determine the loss function value corresponding to each type of information.

[0083] Exemplarily, taking five types of information, namely name, amount, year, cause of action, and the court handling the case, as an example, each training text data is input into the target network. The input data undergoes forward propagation through layers such as the convolutional layer, downsampling layer, and fully connected layer of the target network to obtain output values. According to the output values and target values, the loss function values corresponding to various types of information are determined. For example, the loss function value corresponding to the name information is l1, the loss function value corresponding to the amount information is l2, the loss function value corresponding to the year information is l3, the loss function value corresponding to the cause of action information is l4, and the loss function value corresponding to the court handling the case information is l5. After obtaining the loss function values corresponding to various types of information, in step 203, the target network is trained according to the loss function values corresponding to various types of information.

[0084] In step 203, the target network is trained based on the loss functions corresponding to various types of information.

[0085] In one embodiment, the target network can include multiple sub-networks, and each sub-network corresponds to one type of information; step 203 can specifically be: training the corresponding sub-network based on the loss functions corresponding to various types of information.

[0086] Exemplarily, the loss function corresponding to the corresponding information can be determined by each sub-network. After obtaining the loss functions corresponding to various types of information, the loss function of each type of information is passed back to the corresponding sub-network, and based on this, the errors of the fully connected layer, downsampling layer, convolutional layer, etc. of the sub-network are obtained, and the parameters of the current sub-network are adjusted based on each error, thereby realizing the training of each sub-network.

[0087] In one embodiment, step 203 may include:

[0088] Calculate the proportion of the number of texts corresponding to various types of information;

[0089] Determine the total loss function according to the sum of the products of each loss function and the corresponding proportion;

[0090] Train the target network based on the total loss function.

[0091] Exemplarily, taking a training text data and b types of information as an example for illustration, where the loss function of the mth type of information is l i , 1 ≤ i ≤ b. After information annotation of a training text data, count the number of texts s1... s corresponding to various types of information b , calculate the proportion of the number of texts corresponding to various types of information Calculate each loss function l i and the corresponding proportion r i The sum of the products is Obtain the total loss function l, and then train the target network based on the total loss function l.

[0092] Specifically, the total loss function l can be passed back to the target network, and based on this, the errors of the fully connected layer, downsampling layer, convolutional layer, etc. of the target network are obtained, and the parameters of the target network are adjusted based on each error, thereby realizing the training of the target network.

[0093] In step 204, the text of each first case is input into the trained target network, and the structured data is extracted.

[0094] Among them, after training the target network with the labeled training text data, the text of each first case is input into the trained target network, and the structured data corresponding to various types of information is extracted by the trained target network.

[0095] In a possible implementation manner, it can also be through

[0096] Specifically, referring to Figure 4 , step 101 may include the following steps

[0097] In step 301, a regular expression is determined according to the text information corresponding to the structured data and the context information related to the text information.

[0098] Exemplarily, taking name, amount, year, case cause, and handling court as examples, the structured data of name can be a person's name, the structured data of amount can be a numerical value, the structured data of year can be a time, the structured data of case cause can be the cause of the case involved, and the structured data of handling court can be the court name. Therefore, the text information corresponding to the structured data of name can be a person's name, the text information corresponding to the structured data of amount can be a numerical value, the text information corresponding to the structured data of year can be a time, the text information corresponding to the structured data of case cause can be the cause of the case involved, and the text information corresponding to the structured data of handling court can be the court name.

[0099] The process of determining the regular expression according to the text information and the context information related to the text information can specifically be: generating a regular expression according to the positional relationship between the text information and the context information.

[0100] For example, for the text information "name", its context information can be "defendant", "plaintiff", "witness", etc., and these context information can be before or after the text information "name". Based on the text information "name", the context information "defendant", "plaintiff", "witness", etc., and the positional relationship between the text information and the context information, a corresponding regular expression is generated, and this regular expression matches the target single sentence related to the name in the text of the first case in step 302.

[0101] For example, for the text information "amount", its context information can be "amount involved in the case", "yuan", "compensation amount", etc., and these context information can be before or after the text information "amount". Based on the text information "amount", the context information "amount involved in the case", "yuan", "compensation amount", etc., and the positional relationship between the text information and the context information, a corresponding regular expression is generated, and this regular expression matches the target single sentence related to the amount in the text of the first case in step 302.

[0102] In step 302, the text of the first case is divided into multiple single sentences, and target single sentences that match the regular expression are determined among the multiple single sentences.

[0103] Among them, the text of the first case can be divided into multiple single sentences through semantic analysis, and target single sentences that match each regular expression are determined in each single sentence.

[0104] Exemplarily, in each single sentence, the target single sentences related to the name can be matched through the regular expression corresponding to the name, the target single sentences related to the amount can be matched through the regular expression corresponding to the amount, the target single sentences related to the case cause can be matched through the regular expression corresponding to the case cause, the target single sentences related to the year can be matched through the regular expression corresponding to the year, and the target single sentences related to the handling court can be matched through the regular expression corresponding to the handling court.

[0105] In step 303, extract the structured data from the target single sentences.

[0106] In this step, after matching the target single sentences corresponding to each text information, extract the structured data corresponding to the text information from each target single sentence.

[0107] For example, for the target single sentences related to the name, the names of people, such as Zhang San and Li Si, can be extracted from them; for the target single sentences related to the amount, the amount values, such as 5 million yuan, can be extracted from them; for the target single sentences related to the case cause, the reasons for the case involved, such as financial loan disputes, intellectual property disputes, etc., can be extracted from them; for the target single sentences related to the year, the time, such as 2019, can be extracted from them; for the target single sentences related to the handling court, the name of the court, such as the Higher People's Court of Guangdong, can be extracted from them.

[0108] In some embodiments, refer to Figure 5 , step 303 may specifically include:

[0109] In step 3031, perform field division on the target single sentences; wherein, each target single sentence is divided into at least one field, and each field corresponds to a field name and a field value.

[0110] In step 3032, for each relevant field with the same field content but different field names or field values, perform normalization processing on the field names or field values.

[0111] In step 3033, extract the structured data from each field after the normalization processing.

[0112] Specifically, after performing field division on the target single sentences, for each relevant field with the same field content but different field names or field values, it is necessary to perform normalization processing on the field names or field values, and extract the structured data from the fields after the normalization processing.

[0113] For example, for the fields of "Shanghai High Court" and "Shanghai Higher People's Court", they represent the same court, that is, the field contents are the same, but their field values are different in the database. At this time, the field values of the "Shanghai High Court" field and the "Shanghai Higher People's Court" field need to be set to the same field value, so that the structured data "Shanghai High Court" can be extracted from this field based on this field value.

[0114] In step 102, each structured data of each of the first cases is converted into a corresponding word vector to obtain a first vector matrix corresponding to the target case set.

[0115] Among them, in this embodiment, each structured data can be converted into a corresponding word vector through word2vec, and the word vectors of each first case are written into the vector matrix to generate the first vector matrix.

[0116] Exemplarily, the structured data can be text data or numerical values. Refer to Figure 6 , step 102 may specifically include:

[0117] In step 1021, the text data in each structured data of each of the first cases is converted into a word vector.

[0118] Among them, multiple types of information are extracted from each case, and the information that is text data among various types of information is converted into a word vector through word2vec, and the information of numerical data is not processed; the respective word vectors corresponding to each case can be sequentially located in the same row (or the same column) in the vector matrix, thereby constituting the vector matrix.

[0119] Exemplarily, for cases 1, 2,..., N, the extracted structured data such as case cause, judge's court, amount, etc. are shown in Table 1.

[0120] Table 1

[0121] Case subject matter Court handling the case … Amount Case 1 Financial loan dispute Shenzhen Intermediate People's Court … 3000000 Case 2 Intellectual property dispute Shenzhen Qianhai Court … 600000 … … … … … Case n Fraud Guangdong High People's Court … 10000000

[0122] The text data in each structured data in Table 1 is converted into a word vector through word2vec, as shown in Table 2. Each text data corresponds to a word vector. For example, the word vector of financial loan dispute is [x 11 , y 11 , …, z 11 , and the word vector of Shenzhen Intermediate People's Court is [x 12 , y 12 , …, z 12 .

[0123] Table 2

[0124] Case subject matter Court handling the case … Amount Case 1 <![CDATA[[x 11 ,y 11 ,…,z 11 > <![CDATA[[x 12 ,y 12 ,…,z 12 > … 3000000 Case 2 <![CDATA[[x 21 ,y 21 ,…,z 21 > <![CDATA[[x 22 ,y 22 ,…,z 22 > … 600000 … … … … … Case n <![CDATA[[x n1 ,y n1 ,…,z n1 > <![CDATA[[x n2 ,y n2 ,…,z n2 > … 10000000

[0125] In step 1022, the word vectors and values of each of the first cases are used as elements of the first vector matrix.

[0126] Wherein, each of the first cases corresponds to a row of elements or a column of elements.

[0127] Exemplarily, as shown in Table 1, the word vectors of case 1 and "3000000" can be written into the respective elements of the first row of the first vector matrix, the word vectors of case 2 and "600000" can be written into the respective elements of the second row of the first vector matrix, and the word vectors of case n and "10000000" can be written into the respective elements of the nth row of the first vector matrix.

[0128] In step 103, based on the first vector matrix, the center point of the target case set in the multi-dimensional vector space is determined.

[0129] Exemplarily, step 103 may specifically include:

[0130] Calculate the mean vector of the word vectors corresponding to the same structured data of each first case, and each of the mean vectors is the vector of the center point.

[0131] For example, extract 5 types of structured data, namely structured data 1, structured data 2, structured data 3, structured data 4, and structured data 5. Calculate the mean of the word vectors corresponding to structured data 1 in all first cases to obtain the vector mean 1 corresponding to structured data 1. Calculate the vector mean 2 corresponding to structured data 2, the vector mean 3 corresponding to structured data 3, the vector mean 4 corresponding to structured data 4, and the vector mean 5 corresponding to structured data 5 in this way. The vector mean 1, vector mean 2, vector mean 3, vector mean 4, and vector mean 5 constitute the above-mentioned mean vector matrix.

[0132] Specifically, the mean matrix can be calculated by the following formula:

[0133]

[0134] Where n is the number of first cases in the target case set, m is the number of structured data, is the vector mean corresponding to structured data m, is the feature vector of structured data m of the first case n, constitute the mean vector matrix.

[0135] In addition, weight coefficients can also be set for each structured data of each first case, and the weight coefficients of each structured data can be combined in the process of determining the center point, so that the determined center point is more accurate.

[0136] Specifically, for a structured data, a weight coefficient can be set for the structured data according to the number of times the structured data appears in each first case. For example, the mean matrix can be calculated by the following formula:

[0137]

[0138] where is the vector mean corresponding to the structured data m, is the feature vector of the structured data m of the first case n, c nm is the number of times the structured data m appears in the first case n, constitutes the mean vector matrix.

[0139] In addition, in some cases, the reference significance of cases closer to the current time may be greater, or the reference significance of cases handled by a certain handling department may be greater. Therefore, the weights of all structured data corresponding to the target cases that meet the preset time conditions or handling department conditions can be relatively increased to make the vector matrix more in line with the actual needs.

[0140] In addition, some cases can be further selected from the target case set in step 101 as the final target case set by selecting additional conditions, or the target case set can be divided into multiple sub-target case sets, and then the final target case set or sub-target case sets are processed in the above steps. For example, the target case set can be divided into multiple sub-target case sets by case field, or some cases can be selected from the target case set as the final target case set.

[0141] In step 104, determine the second vector matrix of each second case in the case library, and determine the relevance between each second case and the target case set based on the second vector matrix and the center point.

[0142] Among them, the above second case can be any case in the case library except the target case set. In this step, the relevance between each second case and the above target case set can be determined according to the distance between the second vector matrix and the above center point. For example, the Euclidean distance between the vector matrix of the second case and the center point can be calculated, and the relevance between the second case and the target case set can be determined according to the calculated Euclidean distance.

[0143] Among them, the second vector matrix of each second case can be obtained according to the methods in steps 101 and 102, which will not be elaborated in this step.

[0144] See Figure 7 , step 104 can specifically include:

[0145] In step 1041, according to the vector of the center point and each of the second vector matrices, determine the distance between each of the second cases and the center point.

[0146] For example, the vector of the center point is The vector of any one of the second cases is The Euclidean distance between the vector matrix of the second case and the center point is:

[0147]

[0148] where a i is the mean vector corresponding to structured data i, and b oi is the feature vector corresponding to structured data i in the second case o.

[0149] In step 1041, based on each of the distances, determine the relevance of each of the second cases to the target case set.

[0150] Among them, the relevance is inversely proportional to the above distance, that is, the smaller the distance of the second case, the higher the relevance to the target case set, and the larger the distance of the second case, the smaller the relevance to the target case set. Therefore, the relevance of each second case to the target case set can be determined according to the calculated distance. For example, the relevance of each second case to the target case set can be determined according to the reciprocal of the calculated distance.

[0151] In step 105, send the link information and / or relevant files of the second cases whose relevance meets the preset requirements to the user terminal.

[0152] Exemplarily, the relevance corresponding to each second case can be sorted, and a preset number of second cases can be selected in descending order of relevance, and the connection information of the preset number of second cases can be sent to the user terminal, or the relevant files of the preset number of second cases can be sent to the user terminal. Additionally, a relevance threshold can be set, and the relevance corresponding to each second case can be compared with the relevance threshold, and the link information and / or relevant files of the second cases whose relevance is greater than or equal to the relevance threshold can be sent to the user terminal.

[0153] Among them, the connection information can be the storage path information corresponding to the second case, and the user can click on the connection information on the user terminal to view the relevant files of the second case. Of course, the relevant files of the second cases whose relevance meets the preset requirements can also be directly sent to the user terminal. The user can view the relevant files after downloading them on the user terminal, or preview the relevant files and then select the corresponding relevant files for download and viewing.

[0154] The above case relevance detection method extracts multiple structured data of each first case in the target case set, converts each structured data of each first case into a corresponding word vector to obtain a first vector matrix corresponding to the target case set, determines the center point of the target case set in the multi-dimensional vector space based on the first vector matrix, determines the second vector matrix of each second case in the case library, and determines the relevance between each second case and the target case set based on the second vector matrix and the center point, so as to recommend the second cases whose relevance meets the preset requirements to the user. In the embodiment of the present application, the case text is converted into a vector, and the similar cases to be recommended to the user are determined through the relevance between the second case and the target case set. Compared with the traditional rule matching, the recommended result is more accurate, and there is no need to set data buried points. Only the conventional data in daily office work is required to determine the target case set for similar case recommendation. Moreover, the center point of the first vector matrix can be completed offline, and only the relevance between the recommended case and the center point needs to be calculated online, with relatively low requirements for online computing resources.

[0155] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0156] Corresponding to the case relevance detection method described in the above embodiments, Figure 8 The structural block diagram of the case relevance detection device provided by the embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.

[0157] See Figure 8 , the case relevance detection device in the embodiment of the present application may include a structured data extraction module 401, a vector matrix generation module 402, a center point determination module 403, a relevance determination module 404, and a sending module 405.

[0158] Among them, the structured data extraction module 401 is used to extract multiple structured data of each first case in the case library that meet the preset requirements; among them, all the first cases constitute a target case set;

[0159] The vector matrix generation module 402 is used to convert each structured data of each first case into a corresponding word vector to obtain a first vector matrix corresponding to the target case set;

[0160] The center point determination module 403 is used to determine the center point of the target case set in the multi-dimensional vector space based on the first vector matrix;

[0161] A relevance determination module 404, configured to determine a second vector matrix of each second case in the case library, and determine the relevance between each second case and the target case set based on the second vector matrix and the center point; wherein, the second case is a case in the case library other than the target case set;

[0162] A sending module 405, configured to send link information and / or relevant documents of the second cases whose relevance meets a preset requirement to a user terminal.

[0163] As a possible implementation manner, the structured data extraction module 401 may include:

[0164] An annotation unit, configured to annotate training text data; wherein, the annotated information includes multiple types of information;

[0165] A loss function determination unit, configured to input the annotated training text data into a target network to determine loss functions corresponding to various types of information;

[0166] A training unit, configured to train the target network based on the loss functions corresponding to various types of information;

[0167] A first extraction unit, configured to input the text of each first case into the trained target network to extract the structured data.

[0168] Optionally, the target network includes multiple sub-networks, and each sub-network corresponds to one type of information; specifically, the training unit may be configured to: train the corresponding sub-network based on the loss functions corresponding to various types of information.

[0169] Optionally, the training unit may be specifically configured to:

[0170] Calculate the proportion of the text quantity corresponding to each type of information;

[0171] Determine a total loss function according to the sum of the products of each loss function and the corresponding proportion;

[0172] Train the target network based on the total loss function.

[0173] As a possible implementation manner, the structured data extraction module 401 may include:

[0174] A regular expression determination unit, configured to determine a regular expression according to the text information corresponding to the structured data and the context information related to the text information;

[0175] A determination unit, configured to divide the text of the first case into multiple single sentences, and determine target single sentences that match the regular expression among the multiple single sentences;

[0176] A second extraction unit for extracting structured data from the target single sentence.

[0177] Optionally, the second extraction unit is specifically configured to:

[0178] Perform field division on the target single sentence; wherein, each target single sentence is divided into at least one field, and each field corresponds to a field name and a field value;

[0179] For each relevant field with the same field content but different field names or field values, perform normalization processing on the field name or field value;

[0180] Extract structured data from each of the fields after the normalization processing.

[0181] Optionally, the structured data is text data or a numerical value, and the vector matrix generation module 402 is specifically configured to:

[0182] Convert the text data in the structured data of each of the first cases into word vectors;

[0183] Use the word vectors and numerical values of each of the first cases as elements of the first vector matrix; wherein, each of the first cases corresponds to a row of elements or a column of elements.

[0184] Optionally, the center point determination module 403 is specifically configured to:

[0185] Calculate the mean vector of the word vectors corresponding to the same structured data of each of the first cases; wherein, each of the mean vectors is the vector of the center point;

[0186] The relevance determination module 404 is specifically configured to:

[0187] Determine the distance between each of the second cases and the center point according to the vector of the center point and each of the second vector matrices;

[0188] Determine the relevance between each of the second cases and the target case set based on each of the distances.

[0189] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.

[0190] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0191] The embodiments of this application also provide a terminal device. Refer to Figure 9 , the terminal device 500 may include: at least one processor 510, a memory 520, and a computer program stored in the memory 520 and executable on the at least one processor 510. When the processor 510 executes the computer program, it implements the steps in any of the foregoing method embodiments, such as Figure 2 the steps S101 to S105 in the illustrated embodiment. Alternatively, when the processor 510 executes the computer program, it implements the functions of each module / unit in the foregoing device embodiments, such as Figure 8 the functions of the illustrated modules 401 to 405.

[0192] Exemplarily, the computer program can be divided into one or more modules / units. One or more modules / units are stored in the memory 520 and executed by the processor 510 to complete this application. The one or more modules / units can be a series of computer program segments capable of performing specific functions, and these program segments are used to describe the execution process of the computer program in the terminal device 500.

[0193] Those skilled in the art can understand that Figure 9 this is only an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or combine some components, or different components, such as input / output devices, network access devices, buses, etc.

[0194] The processor 510 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0195] The memory 520 may be an internal storage unit of the terminal device or an external storage device of the terminal device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory 520 is used to store the computer program and other programs and data required by the terminal device. The memory 520 may also be used to temporarily store data that has been output or is to be output.

[0196] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application do not limit to only one bus or one type of bus.

[0197] The case relevance detection method provided by the embodiments of this application can be applied to terminal devices such as computers, mobile phones, wearable devices, in-vehicle devices, tablet computers, laptop computers, netbooks, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, mobile phones, etc. The embodiments of this application do not impose any restrictions on the specific types of terminal devices.

[0198] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the various embodiments of the above-mentioned case relevance detection method can be implemented.

[0199] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal is enabled to execute the steps in the various embodiments of the above-mentioned case relevance detection method.

[0200] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.

[0201] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0202] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0203] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0204] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0205] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for detecting the relevance of cases, characterized in that, Including: Extracting multiple structured data of each first case in the case library that meets the preset requirements; wherein, all the first cases constitute a target case set, and each first case is a case handled by the user or a case selected by the user in the case library; Converting the structured data of each first case into corresponding word vectors to obtain a first vector matrix corresponding to the target case set; Determining the center point of the target case set in the multi-dimensional vector space based on the first vector matrix; Determining a second vector matrix of each second case in the case library, and determining the relevance between each second case and the target case set based on the second vector matrix and the center point; wherein, the second case is a case in the case library other than the target case set; Sending the link information and / or relevant files of the second cases whose relevance meets the preset requirements to the user terminal; Wherein, the extracting multiple structured data of each first case in the target case set includes: Generating a regular expression according to the position relationship between the text information corresponding to the structured data and the context information related to the text information; Dividing the text of the first case into multiple single sentences, and determining target single sentences that match the regular expression among the multiple single sentences; Extracting the structured data from the target single sentences; The structured data is text data or a numerical value. The converting the structured data of each first case into corresponding word vectors to obtain a first vector matrix corresponding to the target case set includes: Converting the text data in the structured data of each first case into word vectors; Taking the word vectors and numerical values of each first case as elements of the first vector matrix; wherein, each first case corresponds to a row of elements or a column of elements.

2. The case relevance detection method according to claim 1, characterized in that The extracting multiple structured data of each first case in the case library that meets the preset requirements includes: Annotating the training text data; wherein, the annotated information includes multiple types of information; Inputting the annotated training text data into a target network to determine loss functions corresponding to various types of information; Training the target network based on the loss functions corresponding to various types of information; Inputting the text of each first case into the trained target network to extract the structured data.

3. The case relevance detection method according to claim 2, characterized in that The target network includes multiple sub-networks, and each sub-network corresponds to one type of information; The training the target network based on the loss functions corresponding to various types of information includes: Training the corresponding sub-network based on the loss functions corresponding to various types of information.

4. The case relevance detection method according to claim 2, wherein The training the target network based on the loss functions corresponding to various types of information includes: Calculating the proportion of the text quantities corresponding to various types of information; Determining a total loss function according to the sum of the products of each loss function and the corresponding proportion; Training the target network based on the total loss function.

5. The case relevance detection method according to claim 1, wherein The extracting structured data from the target single sentences includes: Performing field division on the target single sentences; wherein, each target single sentence is divided into at least one field, and each field corresponds to a field name and a field value; For each relevant field with the same field content but different field names or field values, normalize the field names or field values; Extract structured data from each of the fields after the normalization process.

6. The case relevance detection method according to any one of claims 1 to 5, characterized in that The determining of the center point of the target case set in the multi-dimensional vector space based on the first vector matrix includes: Calculating the mean vector of the word vectors corresponding to the same structured data of each first case; wherein, each of the mean vectors is the vector of the center point; The determining of the relevance between each second case and the target case set based on the second vector matrix and the center point includes: Determining the distance between each second case and the center point according to the vector of the center point and each of the second vector matrices; Determining the relevance between each second case and the target case set based on each of the distances.

7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 6 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Live broadcast identifier recommendation method and related equipment

    CN109462778A