Case Retrieval Method, System and Device Based on XLNet Model

The Xlnet model preprocesses case text and integrates features, solving the problem of insufficient efficiency and accuracy of case search in the existing technology, and achieving more efficient and accurate case search.

CN114490946BActive Publication Date: 2025-07-25CENT SOUTH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210142076.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-07-25
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

The existing similar case search methods have poor search efficiency, accuracy and adaptability, especially the inability to effectively process long text information based on keyword rule matching and deep learning technology, and the Bert model has limited text length, resulting in the omission of key information.

Method used

The Xlnet model is used for case search, and the text similarity characteristics and semantic characteristics are calculated by preprocessing text data, and then fusing them into a fully connected neural network to output the search results. Specific steps include character unification, stop word removal, word participle annotation, calculating jaccard similarity, edit distance and tf-idf cosine distance, and extracting semantic features using the Xlnet model's sorting language model, Attention Mask mechanism and dual-stream self-attention mechanism.

Benefits of technology

The efficiency, accuracy and adaptability of similar case searches are improved. Through data cleaning and semantic feature fusion, the model's understanding of long text is improved, noise interference is reduced, training samples is increased, and the accuracy of prediction results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114490946B_ABST
    Figure CN114490946B_ABST
Patent Text Reader

Abstract

In an embodiment of the present disclosure, a method, system, and device for retrieving similar cases based on the Xlnet model are provided, belonging to the technical field of data processing. Specifically, it includes: preprocessing the target case text and the text in the case retrieval database; calculating the case text similarity features between the preprocessed target case text and the text in the case retrieval database according to a preset algorithm, and extracting semantic features using the Xlnet model; fusing the case text similarity features and the semantic features and inputting them into a fully connected neural network to output the retrieval result. Through the solution of the present disclosure, data cleaning is performed during the preprocessing of case text data to make the information contained in the original data more standardized and accurate. Then, the case text similarity features are calculated, and the Xlnet model is used to convert the text into word vectors to obtain semantic features and fuse them, and the retrieval result is obtained by inputting them into a fully connected neural network, improving the efficiency, accuracy, and adaptability of similar case retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of data processing, and in particular, to a case retrieval method, system, and device based on the Xlnet model. Background Art

[0002] Currently, when understood in a broad sense, "similar cases" can be regarded as "categorized cases", and when understood in a narrow sense, "similar cases" can be regarded as "similar cases". With the development of informatization, society has paid great attention to the trial results of cases, and people's need for case retrieval has become increasingly strong. It is very necessary to provide an intelligent retrieval method that reduces labor costs.

[0003] Currently, the methods for similar case retrieval mainly include techniques based on keyword rule matching and deep learning. Using keyword rule matching technology requires establishing a series of rules, which requires a large amount of labor costs. Moreover, it is difficult for users to summarize a keyword to retrieve the cases they want, and mechanical matching cannot meet the needs of users. The deep learning based on the CNN model or LSTM model has limited learning effects on long texts and cannot effectively represent the information of case texts. The self-attention mechanism of the Bert model can well learn the context information of texts, but the Bert model has certain limitations on the input length of texts. Usually, key parts of long case texts need to be intercepted, which is likely to miss key information. Unilaterally using only the deep learning method or the statistical feature method cannot comprehensively judge the similarity degree of case texts.

[0004] It can be seen that there is an urgent need for a precise, efficient, and adaptable case retrieval method based on the Xlnet model. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide a case retrieval method, system, and device based on the Xlnet model, which at least partially solve the problems of poor retrieval efficiency, accuracy, and adaptability in the prior art.

[0006] In a first aspect, embodiments of the present disclosure provide a case retrieval method based on the Xlnet model, including:

[0007] Preprocessing the target case text and the text in the case retrieval database;

[0008] Calculating the case text similarity features between the preprocessed target case text and the text in the case retrieval database according to a preset algorithm, and extracting the semantic features of the preprocessed target case text and the text in the case retrieval database by using the Xlnet model;

[0009] Fusing the case text similarity features and the semantic features and inputting them into a fully connected neural network to output a retrieval result.

[0010] According to a specific implementation of the embodiment of the present disclosure, the step of preprocessing the target case text and the text in the case search database includes:

[0011] Unifying the characters of the target case text and the text in the case search database;

[0012] Perform stop word removal operation on the target case text after character unification and the text in the case search database;

[0013] The target case text after the stop word removal operation and the text in the case search database are segmented, and each segmented word is tagged with a part of speech.

[0014] According to a specific implementation of the embodiment of the present disclosure, the step of calculating the case text similarity feature between the preprocessed target case text and the text in the case retrieval database according to a preset algorithm includes:

[0015] Calculate the jaccard similarity between the preprocessed target case text and each case text in the preprocessed case retrieval database;

[0016] Calculate the edit distance between the preprocessed target case text and each case text in the preprocessed case retrieval database;

[0017] Calculate the tf-idf cosine distance between the preprocessed target case text and each case text in the preprocessed case retrieval database;

[0018] The jaccard similarity, the edit distance and the tf-idf cosine distance are used as the case text similarity features.

[0019] According to a specific implementation of an embodiment of the present disclosure, the Xlnet model includes a sorting language model, an Attention Mask mechanism and a dual-stream self-attention mechanism.

[0020] According to a specific implementation of the embodiment of the present disclosure, the step of extracting semantic features of the preprocessed target case text and the text in the case retrieval database using the Xlnet model includes:

[0021] Arrange the word order of the preprocessed target case text and each case text in the preprocessed case retrieval database respectively, and randomly sample and predict the sorting results;

[0022] Constructing a mask matrix for the preprocessed target case text and each case text in the preprocessed case retrieval database according to the prediction results;

[0023] The AR language model is pre-trained according to the mask matrix and the two-stream self-attention mechanism to generate the semantic features.

[0024] According to a specific implementation of the embodiment of the present disclosure, before the step of fusing the case text similarity feature with the semantic feature and inputting the result into a fully connected neural network and outputting the search result, the method further includes:

[0025] Construct the fully connected neural network, wherein the fully connected neural network includes an input layer, a hidden layer and an output layer, and a ReLu function is used as an activation function between the input layer and the hidden layer.

[0026] According to a specific implementation of the embodiment of the present disclosure, the step of fusing the case text similarity feature with the semantic feature and inputting the result into a fully connected neural network and outputting the search result includes:

[0027] Concatenate the jaccard similarity, the edit distance and the tf-idf cosine distance into a statistical feature vector;

[0028] splicing the statistical feature vector and the semantic feature into a fused feature vector of the same size as the input layer;

[0029] Input the fused feature vector into the input layer and the hidden layer, perform fitting using a random dropout method, and obtain an output vector through the output layer;

[0030] The output vector is normalized and processed using a preset function to obtain the search result.

[0031] In a second aspect, the present disclosure provides a similar case retrieval system based on the Xlnet model, including:

[0032] A preprocessing module, used to preprocess the target case text and the text in the case retrieval database;

[0033] A calculation module, used to calculate the case text similarity features between the preprocessed target case text and the text in the case retrieval database according to a preset algorithm, and to extract the semantic features between the preprocessed target case text and the text in the case retrieval database using an Xlnet model;

[0034] The fusion module is used to fuse the case text similarity feature with the semantic feature and input the result into a fully connected neural network to output the retrieval result.

[0035] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:

[0036] at least one processor; and,

[0037] A memory communicatively connected to the at least one processor; wherein,

[0038] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the case retrieval method based on the Xlnet model in the foregoing first aspect or any implementation manner of the first aspect.

[0039] The case retrieval solution based on the Xlnet model in the embodiments of the present disclosure includes: preprocessing the target case text and the text in the case retrieval database; calculating the case text similarity features between the preprocessed target case text and the text in the case retrieval database according to a preset algorithm, and extracting the semantic features of the preprocessed target case text and the text in the case retrieval database by using the Xlnet model; fusing the case text similarity features and the semantic features and inputting them into a fully connected neural network to output a retrieval result.

[0040] The beneficial effects of the embodiments of the present disclosure are as follows: Through the solution of the present disclosure, data cleaning is performed during the preprocessing of case text data, so that the information contained in the original data is more standardized and accurate. Then, the case text similarity features are calculated, and the Xlnet model is used to convert the text into word vectors to obtain semantic features and fuse them, and the retrieval result is obtained by inputting into a fully connected neural network, which improves the efficiency, accuracy, and adaptability of case retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 It is a schematic flowchart of a case retrieval method based on the Xlnet model provided by the embodiments of the present disclosure;

[0043] Figure 2 It is a schematic flowchart of a data preprocessing provided by the embodiments of the present disclosure;

[0044] Figure 3 It is a schematic diagram of predictions in different orders generated according to factorization provided by the embodiments of the present disclosure;

[0045] Figure 4 It is a mask matrix provided by the embodiments of the present disclosure;

[0046] Figure 5 It is a schematic diagram of the calculation principle of the dual-stream self-attention mechanism provided by the embodiments of the present disclosure;

[0047] Figure 6 This is the overall workflow diagram of Xlnet provided by the embodiments of the present disclosure;

[0048] Figure 7 This is the fusion network model diagram provided by the embodiments of the present disclosure;

[0049] Figure 8 This is a schematic structural diagram of a case retrieval system based on the Xlnet model provided by the embodiments of the present disclosure;

[0050] Figure 9 This is a schematic diagram of an electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0051] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0052] The following uses specific specific examples to illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0053] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0054] It should also be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. The diagrams only show the components related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0055] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0056] An embodiment of the present disclosure provides a method for retrieving similar cases based on the Xlnet model. The method can be applied to the process of retrieving and comparing case texts in the scenario of text data processing.

[0057] See Figure 1 , which is a schematic flowchart of a method for retrieving similar cases based on the Xlnet model provided by an embodiment of the present disclosure. As Figure 1 shown, the method mainly includes the following steps:

[0058] S101, preprocess the target case text and the text in the case retrieval database;

[0059] Optionally, the preprocessing of the target case text and the text in the case retrieval database in step S101 includes:

[0060] Unify the characters of the target case text and the text in the case retrieval database;

[0061] Perform stop word removal on the target case text and the text in the case retrieval database after unifying the characters;

[0062] Perform word segmentation on the target case text and the text in the case retrieval database after the stop word removal operation, and perform part-of-speech tagging on each segmented word.

[0063] In specific implementation, as Figure 2 shown, considering that the case samples are all Chinese data, their length is usually more than 3000 words, which contains a lot of unhelpful information, such as place names and some specific structural expressions in court judgments, which will cause very large interference to the learning cost and learning accuracy of the deep learning model, and structured data is required for statistical analysis, calculation processing, and model training. To solve the above problems, the data can be preprocessed through the following steps:

[0064] Step 1: Due to reasons such as different input methods, the same Chinese character may be considered different characters by the computer due to the difference between half-width and full-width. Therefore, it is necessary to uniformly convert full-width to half-width.

[0065] Step 2: Due to the particularity of Chinese texts, some stop words will be retained. Removing stop words can not only reduce the length of the text, but also increase the density of useful information words, improve the running speed of the model, and make it easier and more accurate for the model to learn semantics.

[0066] Step 3: Since legal texts contain a series of words such as place names and personal names that are irrelevant to the judgment of case similarity, which will also cause certain interference to the feature vectors extracted by the model, this method first uses the jieba word segmentation tool to segment the case text, and then performs part-of-speech tagging on each word segment. If the word segment is tagged as a noun and appears in the prepared list of useless nouns such as place names and personal names, this method will remove it.

[0067] Step 4: Since the number of original legal case texts is limited and the label distribution is uneven, data augmentation can generate more data from the limited data. For example, if data A and B are similar, and B and C are similar, according to the transitivity of closure, then there is a similarity relationship between A and C. Through this method, the training samples and the diversity of samples can be increased, and the robustness of the model can be improved.

[0068] Compared with the traditional text processing method, the preprocessing can not only shorten the length of the case text, remove most of the noise interference, but also increase the training samples and the diversity of samples, which greatly improves the prediction results.

[0069] S102. Calculate the case text similarity features between the preprocessed target case text and the texts in the case retrieval database according to a preset algorithm, and extract the semantic features between the preprocessed target case text and the texts in the case retrieval database by using the Xlnet model;

[0070] Further, as described in step S102, calculating the case text similarity features between the preprocessed target case text and the texts in the case retrieval database according to a preset algorithm includes:

[0071] Calculate the jaccard similarity between the preprocessed target case text and each case text in the preprocessed case retrieval database;

[0072] Calculate the edit distance between the preprocessed target case text and each case text in the preprocessed case retrieval database;

[0073] Calculate the tf-idf cosine distance between the preprocessed target case text and each case text in the preprocessed case retrieval database;

[0074] Use the jaccard similarity, the edit distance, and the tf-idf cosine distance as the case text similarity features.

[0075] Optionally, the Xlnet model includes a ranking language model, an Attention Mask mechanism, and a two-stream self-attention mechanism.

[0076] Based on the above embodiments, in step S102, extracting the semantic features of the preprocessed target case text and the text in the case retrieval database by using the Xlnet model includes:

[0077] Arrange the word orders of the preprocessed target case text and each case text in the preprocessed case retrieval database respectively, and perform random sampling and prediction on the sorting results;

[0078] Construct a mask matrix for the preprocessed target case text and each case text in the preprocessed case retrieval database according to the prediction results;

[0079] Pre-train the AR language model according to the mask matrix and the two-stream self-attention mechanism to generate the semantic features.

[0080] In specific implementation, the deep learning network can extract the semantic information of the text, but the statistical features, distance features, and similarity functions of the text can also carry information about whether the case texts match. Therefore, this patent proposes the above three aspects of feature information as part of the input of the feature fusion network. The specific steps are as follows:

[0081] Step 1: First, the jaccard similarity of the text can be calculated. The jaccard similarity is used to compare the similarity and difference between finite sample sets. The larger the jaccard value, the higher the similarity of the case texts. Given two case texts A and B, the calculation process of the jaccard coefficient is shown in formula (1):

[0082]

[0083] Step 2: The edit distance between the case texts A and B also reflects the similarity between the case texts. The edit distance e_distance can be calculated by using the dynamic programming method, and the recurrence formula is shown in (2):

[0084]

[0085] Step 3: Calculate the tf-idf cosine distance of the case text. The calculation process of the statistical word frequency is shown in formula (3):

[0086]

[0087] The term frequency (tf) refers to the frequency of a given word in the file. In the above formula, ni,j is the number of times the word appears in file j, and the denominator is the sum of the number of times all words appear in file j. The calculation of inverse document frequency (IDF) is shown in formula (4):

[0088]

[0089] The num in this formula represents the total number of documents. i,j Represents the number of documents in which word j appears. The calculation of tf-idf is shown in formula (5):

[0090] tf-idf=tf*idf# (5)

[0091] Finally, tf-idf is used to calculate the cosine distance of the feature vector, as shown in formula (6):

[0092]

[0093] Advantages: Compared with directly using deep neural networks to extract feature information, this method contains more comprehensive information. In some relatively simple case texts, statistical features and distance features can be used to directly measure the similarity between case texts. The encoding constructed by this method can contain its semantic information, and the length of the encoded sequence is also more suitable. Increasing the weight of statistical features and distance features in the judgment of case text similarity can improve the prediction accuracy of the model.

[0094] Then, considering that the traditional one-hot encoding belongs to mutually orthogonal basis vectors in Euclidean space, it is impossible to calculate the similarity, and when the vocabulary is too large, the one-hot vector will be very sparse, increasing the computational cost. The Word2Vec model is a static word vector model. The trained word vector and the word have a one-to-one correspondence, which makes the model unable to recognize polysemous words. In addition, Word2Vec only considers the local information of the word and has a weak ability to obtain contextual information. The Bert model is an autoencoding language model, but it ignores the correlation between masked words and has certain limitations. The Xlnet model can solve the above problems very well. Xlnet consists of three parts: sorting language model, Attention Mask mechanism, and dual-stream self-attention mechanism, as follows:

[0095] 1. Sorting Language Model

[0096] The Xlnet model is based on the autoregressive (AR) language model, but the AR language model cannot understand the context semantics at the same time. In order to solve this problem, a ranking language model is proposed. The specific steps are as follows:

[0097] Step 1: Permute all word orders. The permutation language model (PLM) obtains different word order structures by fully permuting the sentences of the case text. Suppose a given sequence (x1, x2, x3, x4) is fully permuted to get orders such as 3→2→4→1, 2→4→3→1, 1→4→3→2, etc. Taking the prediction of x3 as an example, for the order 3→2→4→1, since x3 is at the first position, only its hidden state men can be used for prediction. For the order 1→4→3→2, x1 and x4 can be used for prediction. The prediction process is as Figure 3 shown, where (a) and (b) respectively represent predicting the value of x3 according to different orders generated by factorization;

[0098] Step 2: Randomly sample the permutations. For a sequence x of a given length T, there are T! results for its full permutation. When the sequence length is too long, it will lead to too high algorithm complexity. So only random sampling is performed on it. The random sampling optimization of the full permutation of the sequence is achieved through formula (7):

[0099]

[0100] Step 3: Perform partial prediction on the sequence. For the randomly sampled sequence, words at the later positions of the sequence are preferred. Words at the later positions have a higher probability of obtaining more context information.

[0101] 2. Attention Mask mechanism

[0102] PLM enables the model to see the semantic information of the context. At the same time, the Attention Mask mechanism is introduced, making the model still unidirectionally modeled from left to right from the outside. Inside the transformer, the part to be predicted is masked so that it is ignored during the prediction process. For example, if the original sentence arrangement is 1→2→3→4 and the sampled sequence is 3→2→4→1, the constructed mask matrix is as Figure 4 shown. Each row in the matrix diagram represents the x1, x2, x3, x4 sequences in turn. The black masks in each column represent the semantic information that can be seen. The mask in the first row is 2, 3, 4, indicating that x1 can only obtain the semantic information of x2, x3, x4, and so on.

[0103] 3. Two-Stream Self Attention mechanism

[0104] The Two-Stream Self Attention mechanism can solve the problem that the PLM model cannot obtain the position information of words after being scrambled. The two streams are the Content Stream and the Query Stream. The objective function of the traditional AR language model for a sequence x of length T is as formula (8):

[0105]

[0106] Among them, z is a sequence randomly sampled from all permutations of the sequence x with length T, and z t represents the serial number at the t position of the sampled sequence, x is the predicted word, and e(x) is the embedding of x; the content hidden state h θ (x z<t ) encodes both the content of x and the information of its previous context, but does not contain position information. It can be interpreted as the prediction of the word at the t position of the sorted sequence, and the probability is calculated from the words corresponding to the serial numbers before the t position. Since the PLM will shuffle the sequence order, position word information needs to be added, as shown in formula (9):

[0107]

[0108] Among them, g θ (x z<t , z t ) represents the query hidden state, which contains the words before the t position and the position information of the word x to be predicted. It only encodes the context and position information of the predicted word x, but cannot encode the content information of x.

[0109] The content hidden state h θ (x z<t ) and the query hidden state g θ (x z<t , z t ) are updated as shown in the following formulas (10) and (11):

[0110]

[0111]

[0112] Among them, m represents the number of layers of the network. Usually, at the 0th layer, the query hidden state g0 is initialized as a variable w, and the content hidden state h(0) is initialized as the embedding of the word, that is, e(x). Calculate the data of the first layer according to the 0th layer, and calculate layer by layer backward. Specifically, as Figure 5 shown, where (a) is the working principle diagram of the Content Stream, (b) is the working principle diagram of the Query Stream, h and g respectively represent the content information and the query information, Figure 6 is the overall working flow chart of Xlnet.

[0113] The model pre-trained through three parts: the sorted language model, Attention Mask, and two-stream self-attention can consider the relationships between masked words, overcome this problem existing in the Bert model, and can better understand and express the context semantics of case texts. Compared with traditional AR regression language models, the sorting mechanism of Xlnet overcomes the problem that AR models can only predict from left to right or from right to left. For full sorting, random sampling and predicting the words at the end of the sequence can reduce the algorithm complexity and accelerate the convergence speed.

[0114] S103. After fusing the case text similarity feature and the semantic feature, input them into a fully connected neural network to output a retrieval result.

[0115] Optionally, before the step S103 of fusing the case text similarity feature and the semantic feature and inputting them into a fully connected neural network to output a retrieval result, the method further includes:

[0116] Construct the fully connected neural network, where the fully connected neural network includes an input layer, a hidden layer, and an output layer, and a ReLu function is used as the activation function between the input layer and the hidden layer.

[0117] Further, the step S103 of fusing the case text similarity feature and the semantic feature and inputting them into a fully connected neural network to output a retrieval result includes:

[0118] Concatenate the jaccard similarity, the edit distance, and the tf-idf cosine distance into a statistical feature vector;

[0119] Concatenate the statistical feature vector and the semantic feature into a fusion feature vector with the same size as the input layer;

[0120] Input the fusion feature vector into the input layer and the hidden layer, and use the dropout method for fitting to obtain an output vector through the output layer;

[0121] Normalize the output vector and process it using a preset function to obtain the retrieval result.

[0122] Specifically in implementation, the fully connected neural network can be constructed first. As Figure 7 shown, the fully connected neural network can be composed of an input layer Input Layer of 1×771, a hidden layer Hidden Layer of 1×256, and an output layer Output Layer of 1×4. The ReLu is used as the activation function between the input layer and the hidden layer.

[0123] Considering that only using the traditional text similarity judgment method is likely to ignore the semantic information in the case text, while only using the deep learning method is likely to ignore the most original feature information of the case text, this method proposes a fusion network model that fuses the calculated case text similarity features with the semantic feature vectors extracted by Xlnet as the input of the fully connected neural network. The fusion network model consists of a feature fusion stage and a fully connected stage. Step 1: Construct the statistical feature vector V a Concatenate the Jaccard coefficient, edit distance (editdistance), and cosine distance of tf-idf of the case text into a 1×3 statistical feature vector V a .

[0124] Step 2: Construct the fusion feature vector V merge Concatenate the statistical feature vector V a Concatenate the 1×3 statistical feature vector and the 1×768 semantic feature vector V extracted by Xlnet b into a 1×771 fusion feature vector V merge , which is used as the input vector for the fully connected stage in the fusion network.

[0125] Step 3: Use the Dropout method to randomly inactivate certain neurons to reduce overfitting. The inactivation probability p = 0.3. After the feature vector V merge passes through the entire fully connected stage, an output vector V of size 1×4 is obtained output . Perform normalization processing (BatchNormalization) on V output by first subtracting the mean and then dividing by the maximum value, and then process it with the sigmoid function target . Judge the similarity degree of two case texts according to the four-dimensional vector V target . This method divides the text similarity degree into 4 levels.

[0126] Advantages: Compared with only using the traditional text similarity judgment method and only using the features extracted by deep learning, using the fusion network can make the information contained in the feature vector more comprehensive, improve the quality of the feature vector, and greatly improve the accuracy of data prediction.

[0127] The case retrieval method based on the Xlnet model provided in this embodiment preprocesses the case text data by data cleaning to make the information contained in the original data more standardized and accurate, then calculates the case text similarity features, and uses the Xlnet model to convert the text into word vectors, obtains semantic features and fuses them, and inputs them into the fully connected neural network to obtain the retrieval result, improving the efficiency, accuracy, and adaptability of case retrieval.

[0128] Corresponding to the above method embodiments, refer to Figure 8 , an embodiment of the present disclosure further provides a case retrieval system 80 based on the Xlnet model, including:

[0129] A preprocessing module 801, configured to preprocess the target case text and the text in the case retrieval database;

[0130] A calculation module 802, configured to calculate the case text similarity features between the preprocessed target case text and the text in the case retrieval database according to a preset algorithm, and extract the semantic features of the preprocessed target case text and the text in the case retrieval database by using the Xlnet model;

[0131] A fusion module 803, configured to fuse the case text similarity features and the semantic features and input them into a fully connected neural network to output a retrieval result.

[0132] Figure 8 The system shown can correspondingly execute the content in the above method embodiments. For parts not described in detail in this embodiment, refer to the content recorded in the above method embodiments and will not be elaborated here.

[0133] Refer to Figure 9 , an embodiment of the present disclosure further provides an electronic device 90, which includes: at least one processor and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the case retrieval method based on the Xlnet model in the foregoing method embodiments.

[0134] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, which stores computer instructions for causing the computer to execute the case retrieval method based on the Xlnet model in the foregoing method embodiments.

[0135] An embodiment of the present disclosure further provides a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute the case retrieval method based on the Xlnet model in the foregoing method embodiments.

[0136] Next, refer to Figure 9, which shows a schematic structural diagram of an electronic device 90 suitable for implementing the embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0137] As Figure 9 shown, the electronic device 90 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 90 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0138] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 90 to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device 90 having various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively.

[0139] Particularly, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.

[0140] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0141] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.

[0142] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device can execute the relevant steps of the above-mentioned method embodiments.

[0143] Alternatively, the above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device can execute the relevant steps of the above-mentioned method embodiments.

[0144] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0146] The units described in the embodiments of this disclosure can be implemented in software or in hardware.

[0147] It should be understood that the various parts of this disclosure can be implemented in hardware, software, firmware, or a combination thereof.

[0148] As described above, the above is only the specific implementation manner of this disclosure, but the protection scope of this disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by this disclosure should be covered by the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. A method for retrieving similar cases based on the Xlnet model, characterized in that, include: Preprocess the target case text and the text in the case search database; Calculating case text similarity features between the preprocessed target case text and the text in the case retrieval database according to a preset algorithm, and extracting semantic features between the preprocessed target case text and the text in the case retrieval database using an Xlnet model, wherein the Xlnet model includes a sorting language model, an Attention Mask mechanism, and a dual-stream self-attention mechanism; The step of calculating the case text similarity feature between the preprocessed target case text and the text in the case search database according to a preset algorithm includes: Calculate the jaccard similarity between the preprocessed target case text and each case text in the preprocessed case retrieval database; Calculate the edit distance between the preprocessed target case text and each case text in the preprocessed case retrieval database; Calculate the tf-idf cosine distance between the preprocessed target case text and each case text in the preprocessed case retrieval database; Using the jaccard similarity, the edit distance and the tf-idf cosine distance as the case text similarity features; The step of extracting semantic features of the preprocessed target case text and the text in the case retrieval database using the Xlnet model includes: Arrange the word order of the preprocessed target case text and each case text in the preprocessed case retrieval database respectively, and randomly sample and predict the sorting results; Constructing a mask matrix for the preprocessed target case text and each case text in the preprocessed case retrieval database according to the prediction results; Pre-training an AR language model according to the mask matrix and the two-stream self-attention mechanism to generate the semantic features; The case text similarity feature is fused with the semantic feature and input into a fully connected neural network to output the retrieval result.

2. The method according to claim 1, wherein ,The step of preprocessing the target case text and the text in the case retrieval database includes: Unifying the characters of the target case text and the text in the case search database; Perform stop word removal operation on the target case text after character unification and the text in the case search database; The target case text after the stop word removal operation and the text in the case search database are segmented, and each segmented word is tagged with a part of speech.

3. The method according to claim 1, characterized in that , before the step of fusing the case text similarity feature with the semantic feature and inputting the result into a fully connected neural network and outputting the search result, the method further includes: Construct the fully connected neural network, wherein the fully connected neural network includes an input layer, a hidden layer and an output layer, and a ReLu function is used as an activation function between the input layer and the hidden layer.

4. The method according to claim 3, wherein The step of fusing the case text similarity feature with the semantic feature and inputting the result into a fully connected neural network and outputting the search result includes: Concatenate the jaccard similarity, the edit distance and the tf-idf cosine distance into a statistical feature vector; splicing the statistical feature vector and the semantic feature into a fused feature vector of the same size as the input layer; Input the fused feature vector into the input layer and the hidden layer, perform fitting using a random dropout method, and obtain an output vector through the output layer; The output vector is normalized and processed using a preset function to obtain the search result.

5. A similar case retrieval system based on the Xlnet model, characterized in that, include: A preprocessing module, used to preprocess the target case text and the text in the case retrieval database; A calculation module, used to calculate the case text similarity features between the preprocessed target case text and the text in the case retrieval database according to a preset algorithm, and to extract the semantic features of the preprocessed target case text and the text in the case retrieval database using an Xlnet model, wherein the Xlnet model includes a sorting language model, an Attention Mask mechanism, and a two-stream self-attention mechanism; The specific process of the calculation module includes: Calculate the jaccard similarity between the preprocessed target case text and each case text in the preprocessed case retrieval database; Calculate the edit distance between the preprocessed target case text and each case text in the preprocessed case retrieval database; Calculate the tf-idf cosine distance between the preprocessed target case text and each case text in the preprocessed case retrieval database; Using the jaccard similarity, the edit distance and the tf-idf cosine distance as the case text similarity features; Arrange the word order of the preprocessed target case text and each case text in the preprocessed case retrieval database respectively, and randomly sample and predict the sorting results; Constructing a mask matrix for the preprocessed target case text and each case text in the preprocessed case retrieval database according to the prediction results; Pre-training an AR language model according to the mask matrix and the two-stream self-attention mechanism to generate the semantic features; The fusion module is used to fuse the case text similarity feature with the semantic feature and input the result into a fully connected neural network to output the retrieval result.

6. An electronic device, characterized in that, The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the similar case retrieval method based on the Xlnet model described in any one of the preceding claims 1-4.

Citation Information

Patent Citations

  • Similar case retrieval method, similar case retrieval device and electronic equipment

    CN110928994A