Method, apparatus and electronic device for determining legal event judgment information

By using multiple rounds of clustering and confidence score calculation in the legal consultation system, the accuracy of the legal event information entered by the user is determined, and the detection deviation caused by insufficient spoken information is solved, and more accurate legal event judgment information is provided.

CN119398180BActive Publication Date: 2025-07-18NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510005410.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-07-18
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

In the legal consultation scenario, the spoken consultation information entered by the user is not sufficient to provide sufficient information for legal event detection, resulting in deviations in the detection results.

Method used

By obtaining the first event information input by the user, determining the associated multiple preset texts using the preset database, performing multiple rounds of clustering processing, determining candidate seed search pairs, and calculating confidence scores based on the correlation degree and quantity, selecting the target seed search pairs as event determination information.

Benefits of technology

Improve the accuracy of legal event detection, ensure the accuracy of event information, and guide users to supplement information when there is insufficient information to provide more accurate legal advice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119398180B_ABST
    Figure CN119398180B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technologies, and discloses a method, an apparatus, and an electronic device for determining legal event determination information. The method includes: obtaining first event information corresponding to first question information; determining, in a preset database based on the first event information, a plurality of preset texts associated with the first event information; processing a first preset text among the plurality of preset texts to obtain a plurality of retrieval pairs corresponding to the first preset text; performing clustering processing on the plurality of retrieval pairs corresponding to the first preset text to obtain a plurality of candidate seed retrieval pairs; determining a target confidence score for each candidate seed retrieval pair among the plurality of candidate seed retrieval pairs; determining, as target seed retrieval pairs, the candidate seed retrieval pairs among the plurality of candidate seed retrieval pairs whose target confidence scores are greater than or equal to a preset score; and determining the retrieval identifiers and retrieval contents corresponding to the target seed retrieval pairs as event determination information. The present application can ensure the accuracy of the first event information based on the event determination information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a method, an apparatus, and an electronic device for determining legal event determination information. Background Art

[0002] In a legal consultation scenario, a user can consult information related to a legal case through a legal consultation system. When generating a corresponding response according to the consultation information input by the user, the legal consultation system needs to detect legal events involved in the consultation information. Among them, a legal event refers to a legal relationship that a legal subject may be involved in. Thus, the legal consultation system can further process the consultation information based on the determined legal events (such as retrieving relevant cases and legal bases, etc.) to provide accurate consultation opinions for the user. Summary of the Invention

[0003] Embodiments of this application provide a method, an apparatus, and an electronic device for determining legal event determination information, which can determine event determination information based on first event information corresponding to first question information, so as to ensure the accuracy of the first event information corresponding to the first question information identified based on the event determination information. Specifically, the embodiments of this application disclose the following technical solutions:

[0004] In a first aspect of the embodiments of this application, a method for determining legal event determination information is provided. The method includes: obtaining first event information corresponding to first question information input by a user, and determining a plurality of preset texts associated with the first event information in a preset database based on the first event information; processing a first preset text to obtain a plurality of retrieval pairs corresponding to the first preset text; where the first preset text is any one of the plurality of preset texts, and each of the plurality of retrieval pairs includes a retrieval identifier and retrieval content; performing clustering processing on the plurality of retrieval pairs corresponding to the first preset text to obtain a plurality of candidate seed retrieval pairs; determining a target confidence score for each of the plurality of candidate seed retrieval pairs; where the target confidence score is related to the degree of association between each candidate seed retrieval pair and the first event information, and the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair; determining the candidate seed retrieval pairs with a target confidence score greater than or equal to a preset score among the plurality of candidate seed retrieval pairs as target seed retrieval pairs; and determining the retrieval content of the target seed retrieval pairs as the event determination information corresponding to the first question information.

[0005] In some embodiments, clustering the multiple retrieval pairs corresponding to the first preset text to obtain multiple candidate seed retrieval pairs, including: in the clustering process of the current round, obtaining multiple first seed retrieval pairs and multiple unclustered first retrieval pairs obtained in the clustering process of the previous round; wherein, the multiple retrieval pairs include the multiple first seed retrieval pairs and the multiple first retrieval pairs, and the sum of the number of the multiple first seed retrieval pairs and the multiple first retrieval pairs is equal to a preset number; clustering the multiple first seed retrieval pairs and the multiple first retrieval pairs to obtain multiple second retrieval pairs corresponding to each of the multiple legal categories; among the multiple second retrieval pairs corresponding to each of the legal categories, determining a second seed retrieval pair corresponding to each of the legal categories, and determining multiple unclustered second retrieval pairs in the multiple retrieval pairs according to the preset number, the multiple first retrieval pairs, and the multiple second seed retrieval pairs; in the clustering process of the next round, clustering the multiple second seed retrieval pairs and the multiple second retrieval pairs until there are no unclustered retrieval pairs in the multiple retrieval pairs, and obtaining the multiple candidate seed retrieval pairs determined in the last round of clustering process; wherein, each of the multiple candidate seed retrieval pairs corresponds to a different legal category.

[0006] In some embodiments, among the multiple second retrieval pairs corresponding to each of the legal categories, determining a second seed retrieval pair corresponding to each of the legal categories, including: in each of the legal categories, determining a semantic vector corresponding to each of the second retrieval pairs; according to the semantic vectors corresponding to each of the second retrieval pairs, determining the total distance between each first semantic vector and a second semantic vector in each of the legal categories; wherein, the first semantic vector is the semantic vector corresponding to any one of the multiple second retrieval pairs corresponding to each of the legal categories, and the second semantic vector is the semantic vector corresponding to a second retrieval pair other than the second retrieval pair corresponding to the first semantic vector; determining a target semantic vector with the shortest total distance from the multiple first semantic vectors to the second semantic vector, and determining the second retrieval pair corresponding to the target semantic vector as the second seed retrieval pair corresponding to each of the legal categories.

[0007] In some embodiments, determining the target confidence score for each of the multiple candidate seed retrieval pairs includes: determining a first confidence score for each of the candidate seed retrieval pairs according to the degree of association between each of the candidate seed retrieval pairs and the first event information; wherein, the higher the degree of association between each of the candidate seed retrieval pairs and the first event information, the higher the first confidence score of each of the candidate seed retrieval pairs; determining the number of retrieval pairs in the legal category corresponding to each of the candidate seed retrieval pairs, and determining a second confidence score for each of the candidate seed retrieval pairs according to the number of retrieval pairs corresponding to each of the candidate seed retrieval pairs; wherein, the more the number of retrieval pairs corresponding to each of the candidate seed retrieval pairs, the higher the second confidence score of each of the candidate seed retrieval pairs; based on the first confidence score and the second confidence score of each of the candidate seed retrieval pairs, determining the target confidence score for each of the candidate seed retrieval pairs.

[0008] In some embodiments, determining the retrieval content of the target seed retrieval pair as the event determination information corresponding to the first question information includes: if the ratio of the number of the target seed retrieval pairs to the number of the multiple candidate seed retrieval pairs is greater than or equal to a preset ratio, determining the retrieval content of the target seed retrieval pair as the event determination information.

[0009] In some embodiments, the method further includes: if the ratio of the number of the target seed retrieval pairs to the number of the multiple candidate seed retrieval pairs is less than the preset ratio, and the number of times of loop execution of clustering is less than or equal to the preset number of execution times, continuing to perform clustering processing on the multiple retrieval pairs until the ratio of the number of the target seed retrieval pairs to the number of the multiple candidate seed retrieval pairs is greater than or equal to the preset ratio, or the number of times of loop execution of clustering is greater than the preset number of execution times.

[0010] In some embodiments, the method further includes: if the number of times of loop execution of clustering is greater than the preset number of execution times, processing a second preset text among the multiple preset texts to obtain multiple retrieval pairs corresponding to the second preset text; wherein, the second preset text is any one of the multiple preset texts other than the first preset text; performing clustering processing on the multiple retrieval pairs corresponding to the second preset text to determine the target seed retrieval pair corresponding to the second preset text, and determining the retrieval content of the target seed retrieval pair corresponding to the second preset text as the event determination information corresponding to the first question information.

[0011] In some embodiments, the above event determination information is used to determine whether the above first event information holds; when it is determined that the above first event information holds, the above first event information is the target event information corresponding to the above first question information; when it is impossible to determine whether the above first event information holds, the above event determination information is further used to determine a second question information to obtain a response information of the user to the above second question information; the above response information is used to determine a second event information corresponding to the above first question information, and the above second event information is the target event information corresponding to the above first question information.

[0012] In a second aspect of the embodiments of the present application, a device for determining legal event determination information is provided. The device includes: an acquisition module configured to acquire first event information corresponding to a first question information input by a user, and determine a plurality of preset texts associated with the above first event information in a preset database based on the above first event information; a processing module configured to process a first preset text to obtain a plurality of retrieval pairs corresponding to the above first preset text; wherein the above first preset text is any one of the above plurality of preset texts, and each of the above plurality of retrieval pairs includes a retrieval identifier and a retrieval content; a clustering module configured to perform clustering processing on the plurality of retrieval pairs corresponding to the above first preset text to obtain a plurality of candidate seed retrieval pairs; a first determination module configured to determine a target confidence score for each of the above candidate seed retrieval pairs among the above plurality of candidate seed retrieval pairs; wherein the above target confidence score is related to the degree of association between each of the above candidate seed retrieval pairs and the above first event information, and the number of retrieval pairs in the legal category corresponding to each of the above candidate seed retrieval pairs; a second determination module configured to determine the candidate seed retrieval pairs among the above plurality of candidate seed retrieval pairs whose target confidence scores are greater than or equal to a preset score as target seed retrieval pairs; a third determination module configured to determine the retrieval content of the above target seed retrieval pairs as the event determination information corresponding to the above first question information.

[0013] In a third aspect of the embodiments of the present application, an electronic device is provided, including: one or more processors and a memory, the memory being configured to: store one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining legal event determination information described in the foregoing first aspect.

[0014] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The storage medium stores computer program instructions, and when the computer reads the instructions, it executes the method for determining legal event determination information described in the foregoing first aspect.

[0015] A fifth aspect of an embodiment of the present application provides a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to execute the method for determining legal event determination information described in the foregoing first aspect. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of a method for determining legal event determination information provided by some embodiments of the present application;

[0018] Figure 2 It is a flowchart of another method for determining legal event determination information provided by some embodiments of the present application;

[0019] Figure 3 It is a flowchart of yet another method for determining legal event determination information provided by some embodiments of the present application;

[0020] Figure 4 It is a flowchart of yet another method for determining legal event determination information provided by some embodiments of the present application;

[0021] Figure 5 It is a flowchart of yet another method for determining legal event determination information provided by some embodiments of the present application;

[0022] Figure 6 It is a schematic diagram of a device for determining legal event determination information provided by some embodiments of the present application;

[0023] Figure 7 It is a schematic diagram of an electronic device provided by some embodiments of the present application. Detailed Embodiments

[0024] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention and make the above-mentioned objects, features, and advantages of the embodiments of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be further described in detail below with reference to the drawings.

[0025] When determining the legal events involved in the consultation information input by the user, since the consultation information input by the user is usually more colloquial, there may not be enough information in the consultation information for legal event detection, which may lead to certain deviations in the detection of legal events.

[0026] Based on the above technical problems, the present application provides a method, device, and electronic device for determining legal event determination information. After identifying the legal events involved in the consultation information input by the user, the event determination information can be further determined according to the identified legal event information, so that it can be determined whether the above legal event information is accurate based on the event determination information and the consultation information subsequently.

[0027] It should be noted that the training samples used in the training process of the neural network model involved in the present application are all from legally authorized legal documents, judgments, case descriptions, etc., and the conclusions obtained by the method for determining legal event determination information provided by the present application are only used to form consultation opinions for the user's reference.

[0028] The following provides a detailed description of the method for determining legal event determination information provided by the present application.

[0029] Figure 1 It is a flowchart of a method for determining legal event determination information provided by some embodiments of the present application. As Figure 1 shown, the method for determining legal event determination information may include steps 110 to 140.

[0030] Step 110, obtain the first event information corresponding to the first question information input by the user, and determine a plurality of preset texts associated with the first event information in a preset database.

[0031] In some embodiments, the user may input legal consultation information (hereinafter referred to as the first question information) in the user interface provided by the legal consultation system (hereinafter referred to as the central control system). The central control system may call corresponding legal tools in response to the first question information input by the user to process the first question information to obtain a processing result, and then summarize the processing results obtained by each legal tool to obtain a final consultation opinion and output it to the user.

[0032] Among them, the legal tool may include a legal event detection model and a legal event establishment condition determination model (hereinafter simply referred to as the establishment condition determination model). The legal event detection model can preliminarily identify the legal events involved in the first question information input by the user to obtain a legal event detection result (i.e., the first event information). Exemplarily, the legal event detection model can extract the context relationships corresponding to each entity in the first question information input by the user (i.e., the initial features corresponding to each entity), and update the initial features corresponding to each entity by means of hierarchical propagation and aggregation of neighbor node information, so as to finally obtain the core information corresponding to each entity and the core relationships between each entity; furthermore, based on the core information corresponding to each entity and the core relationships between each entity, the legal event information corresponding to the first question information (i.e., the first event information) is determined.

[0033] Since the content of the first question information input by the user is usually more colloquial and may not necessarily provide comprehensive information for legal event detection, at this time, it is still impossible to determine whether the first event information is consistent with the legal event actually consulted by the user. Therefore, the establishment condition determination model can determine the event determination information corresponding to the first event information, so as to determine whether the first question information is accurate through the event determination information.

[0034] In some embodiments, after obtaining the first event information output by the legal event detection model, the establishment condition determination model can retrieve the judgment documents (i.e., the preset texts) related to the first event information in the preset data. For example, if the first question information input by the user is "I rented a storefront from Party A and signed a three-year lease contract, but now I have lost money in business due to an accident. Can I request to terminate the lease contract in advance", and after the legal event detection model processes the first question information, the obtained first event information is "Contract termination in case of contract deadlock", then the establishment condition determination model can retrieve the judgment documents related to "Contract termination in case of contract deadlock" in the preset database.

[0035] Step 120: Process the first preset text to obtain multiple retrieval pairs corresponding to the first preset text.

[0036] In some embodiments, then, the establishment condition determination model can analyze any one of the above-mentioned multiple preset texts (i.e., the first preset text) through a large language model (LLM) to obtain multiple retrieval pairs corresponding to the first preset text.

[0037] Among them, each retrieval pair in multiple retrieval pairs includes a retrieval identifier and retrieval content. The retrieval identifier is used to represent a legal act or legal event, such as "theft", "contract rescission", "tortious act", etc.; the retrieval content is used to represent an entity associated with the legal act or legal event, such as "Zhang San", "5,000 yuan", "malicious breach of contract", etc. In the preset text, each legal act may correspond to one or more entities, and multiple different retrieval pairs can be determined according to each legal act and each entity.

[0038] Step 130, perform clustering processing on multiple retrieval pairs corresponding to the first preset text to obtain multiple candidate seed retrieval pairs.

[0039] In some embodiments, after determining multiple retrieval pairs corresponding to the above-mentioned first preset text, the condition determination model can perform clustering processing on the multiple retrieval pairs through the large language model LLM to determine multiple candidate seed retrieval pairs. Exemplarily, to avoid an excessive burden on the model caused by clustering multiple retrieval pairs at one time, a multi-round clustering method can be used to perform clustering processing on the multiple retrieval pairs, and the multi-round clustering can also enable the clustering process of the next round to learn the clustering results of the previous round, thereby improving the clustering effect.

[0040] Exemplarily, in each round of the clustering process, a preset number of retrieval pairs (including the seed retrieval pairs obtained from the previous round of clustering and the unclustered retrieval pairs) can be selected and clustered to obtain multiple retrieval pairs corresponding to each legal category in one or more legal categories, and then the seed retrieval pairs are determined from the multiple retrieval pairs corresponding to each legal category. Among them, the seed retrieval pair is a representative retrieval pair among the multiple retrieval pairs corresponding to each legal category. After completing the clustering of all retrieval pairs, the final multiple candidate seed retrieval pairs can be obtained. Among them, each candidate seed retrieval pair corresponds to a legal category.

[0041] Among them, the legal category can be determined according to the legal field to which the legal case belongs or the specific type of the legal case, etc. For example, the legal category can include criminal law categories (such as intentional injury crime, theft crime), civil law categories (such as contract disputes, property inheritance), administrative law categories (such as administrative penalties), intellectual property law categories (such as patents, trademarks, copyrights), labor law categories (such as labor contracts, work injury compensation), etc.; or include contract categories (such as lease contract disputes, sales contract disputes, service contract disputes), tort categories (such as personal injury compensation, traffic accident liability), labor disputes (such as labor contract rescission, unpaid overtime wages), etc. It should be noted that the above specific content of the legal category is only an example, and this embodiment does not limit it.

[0042] Step 140: Determine a target seed retrieval pair that meets a preset condition among multiple candidate seed retrieval pairs, and determine the retrieval content of the target seed retrieval pair as the event determination information corresponding to the first question information.

[0043] In some embodiments, after determining the multiple candidate seed retrieval pairs, a target seed retrieval pair that meets the preset condition can be determined from the multiple candidate seed retrieval pairs. Among them, the target seed retrieval pair can be a candidate retrieval pair with a relatively high semantic association degree with the first event information, and the number of target seed retrieval pairs can be one or more. Then, the event determination information corresponding to the first question information is determined according to the retrieval identifier and retrieval content of the target seed retrieval pair.

[0044] Exemplarily, if the first question information input by the user is "Can I terminate the lease contract of the store?", the first event information obtained after detecting the legal event for the first question information is the termination of the contract in the case of a contract deadlock, and the retrieval identifier of the determined target seed retrieval pair is "contract termination", and the retrieval content is "It is manifestly unfair for the breaching party to continue to perform the contract", then the retrieval content of the target seed retrieval pair can be determined as the event determination information corresponding to the first question information.

[0045] Through the above solution, after identifying the legal event involved in the consultation information input by the user through the legal event detection model and obtaining the first event information, the corresponding event determination information can be further determined based on the first event information, so that it can be determined whether the above first event information is accurate according to the event determination information and the consultation information subsequently.

[0046] Figure 2 For another flowchart of a method for determining legal event determination information provided by some embodiments of the present application, as Figure 2 shown, the above step 130 may include steps 210 to 240.

[0047] Step 210: Obtain multiple first seed retrieval pairs obtained in the previous round of clustering process and multiple unclustered first retrieval pairs during the clustering process of the current round.

[0048] In some embodiments, during the process of performing multiple rounds of clustering on multiple retrieval pairs corresponding to the first preset text, the clustering process of the current round first obtains multiple seed retrieval pairs (i.e., the first seed retrieval pairs) obtained in the previous round of clustering process and multiple retrieval pairs that have not been clustered (i.e., the first retrieval pairs).

[0049] Exemplarily, take the number of retrieval pairs as 1000 and the preset number as 100 as an example. In the first round of clustering, 100 retrieval pairs are selected from 1000 retrieval pairs for unsupervised clustering to obtain multiple retrieval pairs corresponding to each legal category among multiple legal categories, and the first seed retrieval pairs corresponding to each legal category are determined from the multiple retrieval pairs corresponding to each legal category. If the number of the first seed retrieval pairs is 30, then in the second round of clustering, 70 unclustered first retrieval pairs are selected from the remaining 900 retrieval pairs (that is, the sum of the number of the first seed retrieval pairs and the first retrieval pairs is equal to the preset number of 100), and then the above 50 first seed retrieval pairs and 70 first retrieval pairs are clustered again.

[0050] Step 220: Cluster the multiple first seed retrieval pairs and the multiple first retrieval pairs to obtain multiple second retrieval pairs corresponding to each legal category among multiple legal categories.

[0051] In some embodiments, after obtaining the multiple first seed retrieval pairs and the multiple first retrieval pairs through the clustering process of the previous round, the multiple first seed retrieval pairs and the multiple first retrieval pairs are subjected to unsupervised clustering through the clustering process of the current round to obtain multiple second retrieval pairs corresponding to each legal category among multiple legal categories. Continuing with the above example, after clustering the above 50 first seed retrieval pairs and 70 first retrieval pairs, multiple second retrieval pairs corresponding to each legal category can be obtained, such as multiple second retrieval pairs corresponding to 20 legal categories respectively.

[0052] In some examples, an upper limit value of the number of legal categories obtained in each round of the clustering process can be set. For example, when clustering 100 retrieval pairs in each round, the number of legal categories obtained is at most 60.

[0053] In some examples, during the clustering process, different weights can also be set for the retrieval identifier and the retrieval content respectively. Exemplarily, the weight corresponding to the retrieval identifier can be set to be less than the weight corresponding to the retrieval content (such as setting the weight corresponding to the retrieval identifier to 0.4 and the weight corresponding to the retrieval content to 0.6), so that the basis for the condition determination model to cluster multiple retrieval pairs will be more biased towards the retrieval content of each retrieval pair.

[0054] Step 230: Among the multiple second retrieval pairs corresponding to each legal category, determine the second seed retrieval pairs corresponding to each legal category, and determine multiple unclustered second retrieval pairs among the multiple retrieval pairs according to the preset number, the multiple first retrieval pairs, and the multiple second seed retrieval pairs.

[0055] In some embodiments, after obtaining the multiple second retrieval pairs corresponding to each legal category through the clustering process of the current round, representative second seed retrieval pairs are determined among the multiple second retrieval pairs corresponding to each legal category.

[0056] Exemplarily, the method for determining the second seed retrieval pairs corresponding to each legal category can be determined based on the distance (such as Euclidean distance) between each second retrieval pair, or can be obtained by semantic summarization of multiple second retrieval pairs corresponding to each legal category by a large language model. This embodiment does not limit this.

[0057] In some embodiments, after completing the clustering of the current round, a plurality of unclustered second retrieval pairs are obtained from the plurality of retrieval pairs according to a preset quantity, and the sum of the quantities of the plurality of second retrieval pairs is equal to the preset quantity. Continuing with the above example, if a plurality of second retrieval pairs corresponding to 20 legal categories are obtained, the number of the determined plurality of second seed retrieval pairs is 20 (that is, each legal category corresponds to one second seed retrieval pair), and then 80 unclustered second retrieval pairs are selected from the remaining 800 retrieval pairs (that is, the sum of the number of second seed retrieval pairs and the number of second retrieval pairs is equal to the preset quantity of 100), and then the above 20 second seed retrieval pairs and 80 second retrieval pairs are subjected to the next round of clustering.

[0058] Step 240, in the clustering process of the next round, cluster the plurality of second seed retrieval pairs and the plurality of second retrieval pairs until there are no unclustered retrieval pairs in the plurality of retrieval pairs, and obtain the plurality of candidate seed retrieval pairs determined in the last round of clustering process.

[0059] In some embodiments, in the clustering process of the next round, cluster the above-mentioned plurality of second seed retrieval pairs and the plurality of second retrieval pairs, and so on, until all the plurality of retrieval pairs are clustered (that is, there are no unclustered retrieval pairs in the plurality of retrieval pairs), and obtain the plurality of candidate seed retrieval pairs determined in the last round of clustering process. Among them, each candidate seed retrieval pair in the plurality of candidate seed retrieval pairs corresponds to a different legal category.

[0060] Through the above solution, the plurality of retrieval pairs corresponding to the first preset text can be clustered by multiple rounds of clustering, so as to improve the clustering effect, that is, a more accurate classification result can be obtained, and further, more accurate event determination information can be determined according to the seed retrieval pairs.

[0061] Figure 3 It is a flowchart of another method for determining legal event determination information provided by some embodiments of the present application. As Figure 3 shown, "determine the second seed retrieval pairs corresponding to each legal category among the plurality of second retrieval pairs corresponding to each legal category" in the above step 230 may include steps 310 to 330.

[0062] Step 310, in each legal category, determine the semantic vectors corresponding to each second retrieval pair.

[0063] In some embodiments, when determining the second seed retrieval pairs corresponding to each legal category based on the distances (such as Euclidean distances) between each pair of second retrievals, the semantic vectors corresponding to each pair of second retrievals in each legal category can be determined first. Exemplarily, the semantic vectors corresponding to the retrieval contents of each pair of second retrievals can be determined.

[0064] For example, in the legal category of "contract rescission", the retrieval contents of the second retrieval pairs and the corresponding semantic vectors of each retrieval content are: S1 = "Although the breaching party breached the contract first, it was not a malicious breach", V1 = [0.8, 0.5]; S2 = "It is grossly unfair for the breaching party to continue to perform the contract", V2 = [0.7, 0.6]; S3 = "The observing party refuses to rescind the contract, which violates the principle of good faith", V3 = [0.9, 0.4]. It should be noted that the specific values of the semantic vectors corresponding to the above retrieval contents are only examples.

[0065] Step 320: Determine the total distance between each first semantic vector and each second semantic vector in each legal category according to the semantic vectors corresponding to each pair of second retrievals.

[0066] In some embodiments, after determining the semantic vectors corresponding to each pair of second retrievals in each legal category, the distances (such as Euclidean distances) between the semantic vectors of each pair of second retrievals in each legal category and the semantic vectors of other pairs of second retrievals in the same legal category are determined, that is, the distances between each first semantic vector and each second semantic vector in each legal category are determined, and the distances between each first semantic vector and each second semantic vector are added up to obtain the total distance between each first semantic vector and each second semantic vector. Among them, the first semantic vector is the semantic vector corresponding to any one of the multiple pairs of second retrievals corresponding to each legal category, and the second semantic vector is the semantic vector corresponding to the pair of second retrievals other than the pair of second retrievals corresponding to the first semantic vector.

[0067] Continuing with the above example, among the above second retrieval pairs S1, S2, and S3, the Euclidean distance between the semantic vector V1 of the second retrieval pair S1 and the semantic vector V2 of S2 ; the Euclidean distance between the semantic vector V1 of the second retrieval pair S1 and the semantic vector V3 of S3 . Then the total distance between the semantic vector V1 of the second retrieval pair S1 and the semantic vectors of the second retrieval pairs S2 and S3 .

[0068] Similarly, the total distance between the semantic vector V2 of the second retrieval pair S2 and the semantic vectors of the second retrieval pairs S1 and S3 can be determined ; and the total distance between the semantic vector V3 of the second retrieval pair S3 and the semantic vectors V1 of the second retrieval pair S1 and V2 of the second retrieval pair S2 .

[0069] Step 330: Determine the target semantic vector with the shortest total distance from the multiple first semantic vectors to the second semantic vector, and determine the second retrieval pair corresponding to the target semantic vector as the second seed retrieval pair corresponding to each legal category.

[0070] In some embodiments, after obtaining the total distances corresponding to the above first semantic vectors, the first semantic vector with the shortest total distance can be determined as the target semantic vector, and the second retrieval pair corresponding to the target semantic vector is the second seed retrieval pair corresponding to its legal category.

[0071] Continuing with the above example, in the total distance corresponding to the semantic vector V1 of the second retrieval pair S1 , the total distance corresponding to the semantic vector V2 of the second retrieval pair S2 and the total distance corresponding to the semantic vector V3 of the second retrieval pair S3 , the total distance has the smallest value. Therefore, the semantic vector V1 of the second retrieval pair S1 can be determined as the target semantic vector, and the second retrieval pair S1 is the second seed retrieval pair in the legal category of "contract rescission".

[0072] Through the above solution, when determining the seed retrieval pair, by determining the sum of the distances between the semantic vectors of each retrieval pair in each legal category and the semantic vectors of all other retrieval pairs, and selecting the retrieval pair with the smallest sum of distances as the seed retrieval pair, the representativeness and centrality of the seed retrieval pair can be ensured, thereby ensuring the accuracy of the event determination information determined based on the seed retrieval pair.

[0073] Figure 4 is a flowchart of another method for determining legal event determination information provided by some embodiments of the present application. As Figure 4 shown, the "determine the target seed retrieval pair that meets the preset conditions among the multiple candidate seed retrieval pairs" in the above step 140 may include steps 410 to 420.

[0074] Step 410: Determine the target confidence score of each candidate seed retrieval pair among the multiple candidate seed retrieval pairs.

[0075] In some embodiments, after finally obtaining the candidate seed retrieval pairs corresponding to each legal category through multiple rounds of clustering, the target confidence score of each candidate seed retrieval pair can be determined according to a preset scoring rule.

[0076] Among them, the target confidence score may be related to the degree of association between each candidate seed retrieval pair and the first event information, as well as the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair. For example, the higher the degree of association between each candidate seed retrieval pair and the first event information, and the larger the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair, the higher the target confidence score of each candidate seed retrieval pair; conversely, the lower the degree of association between each candidate seed retrieval pair and the first event information, and the smaller the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair, the lower the target confidence score of each candidate seed retrieval pair.

[0077] Step 420: Determine the candidate seed retrieval pairs with a target confidence score greater than or equal to a preset score among the multiple candidate seed retrieval pairs as the target seed retrieval pairs.

[0078] In some embodiments, then, candidate seed retrieval pairs with a target confidence score greater than or equal to a preset score can be determined among the multiple candidate seed retrieval pairs, and these candidate seed retrieval pairs are determined as the target seed retrieval pairs. For example, if the target confidence score corresponding to candidate seed retrieval pair S1 is 80, the target confidence score corresponding to candidate seed retrieval pair S2 is 60, the target confidence score corresponding to candidate seed retrieval pair S3 is 50, the target confidence score corresponding to candidate seed retrieval pair S4 is 30, the target confidence score corresponding to candidate seed retrieval pair S5 is 70, and the preset score is 60, then the candidate seed retrieval pairs S1, S2, and S5 with a target confidence score greater than or equal to 60 are the target seed retrieval pairs.

[0079] Through the above solution, based on the degree of association between each candidate seed retrieval pair and the first event information and the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair, the confidence score of each candidate seed retrieval pair is determined, so that the target seed retrieval pairs determined based on the confidence scores of each candidate seed retrieval pair can be more relevant to the first event information and can more accurately represent each legal category, thereby further improving the accuracy of determining whether the first event information holds based on the target seed retrieval pairs in the subsequent process.

[0080] Figure 5 The flowchart of another method for determining legal event determination information provided by some embodiments of this application is as Figure 5 shown, and the above step 410 may include steps 510 to 530.

[0081] Step 510: Determine the first confidence score of each candidate seed retrieval pair according to the degree of association between each candidate seed retrieval pair and the first event information.

[0082] In some embodiments, when determining the target confidence score of each candidate seed retrieval pair, the first confidence score of each candidate seed retrieval pair may be determined according to the degree of association between each candidate seed retrieval pair and the first event information. Among them, the higher the degree of association between each candidate seed retrieval pair and the first event information, the higher the first confidence score of each candidate seed retrieval pair.

[0083] Step 520, determine the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair, and determine the second confidence score of each candidate seed retrieval pair according to the number of retrieval pairs corresponding to each candidate seed retrieval pair.

[0084] In some embodiments, and, the second confidence score of each candidate seed retrieval pair may be determined according to the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair. Among them, the more the number of retrieval pairs corresponding to each candidate seed retrieval pair, the higher the second confidence score of each candidate seed retrieval pair.

[0085] Step 530, based on the first confidence score and the second confidence score of each candidate seed retrieval pair, determine the target confidence score of each candidate seed retrieval pair.

[0086] In some embodiments, after determining the first confidence score and the second confidence score corresponding to each candidate seed retrieval pair, the target confidence score corresponding to each candidate seed retrieval pair may be determined according to the sum of the first confidence score and the second confidence score. For example, if the first confidence score score1 corresponding to the candidate seed retrieval pair S1 is 80 and the second confidence score score2 is 60, then the target confidence score corresponding to the candidate seed retrieval pair S1 is score = score1 + score2 = 80 + 60 = 140.

[0087] Exemplarily, different weights may also be set for the first confidence score and the second confidence score, and the target confidence score is determined based on the first confidence score, the second confidence score, and the corresponding weights. Continuing with the above example, if the weight of the first confidence score is set to 0.6 and the weight of the second confidence score is set to 0.4, then the target confidence score corresponding to the candidate seed retrieval pair S1 is score = 0.6×score1 + 0.4×score2 = 0.6×80 + 0.4×60 = 72.

[0088] Through the above solution, determining the confidence score of each candidate seed retrieval pair based on the degree of association between each candidate seed retrieval pair and the first event information can improve the accuracy when determining the target seed retrieval pair; determining the confidence score of each candidate seed retrieval pair based on the number of retrieval pairs corresponding to each candidate seed retrieval pair can improve the universality when determining the target seed retrieval pair, improve the generalization ability, and reduce the problem of overfitting.

[0089] In some embodiments, step 140 described above includes: if the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is greater than or equal to a preset ratio, then determine the retrieval content of the target seed retrieval pairs as event determination information.

[0090] In some embodiments, after obtaining multiple candidate seed retrieval pairs and determining the target confidence score of each candidate seed retrieval pair, it is possible to further determine whether the number of target seed retrieval pairs reaches a certain proportion among all candidate seed retrieval pairs. Exemplarily, the ratio of the number of target seed retrieval pairs to the number of all candidate seed retrieval pairs can be determined, and it can be determined whether this ratio is greater than or equal to a preset ratio. If the proportion of the target seed retrieval pairs among all candidate seed retrieval pairs is greater than the preset ratio, then determine the retrieval content of the target seed retrieval pairs as the event determination information corresponding to the first question information.

[0091] For example, if the total number of candidate seed retrieval pairs is 5, the number of target seed retrieval pairs is 3, and the preset ratio is 0.6, then the proportion of the target seed retrieval pairs among the candidate seed retrieval pairs is 3 / 5 = 0.6, that is, the proportion of the target seed retrieval pairs is equal to the preset ratio. Therefore, the retrieval content of the target seed retrieval pairs can be determined as the event determination information corresponding to the first question information. If the retrieval content of target seed retrieval pair S1 is "Although the breaching party breached the contract first, it was not a malicious breach", the retrieval content of target seed retrieval pair S2 is "It is grossly unfair for the breaching party to continue to perform the contract", and the retrieval content of target seed retrieval pair S3 is "The non-breaching party's refusal to terminate the contract violates the principle of good faith", then the event determination information corresponding to the first question information is the above retrieval content.

[0092] In some embodiments, the above method further includes: if the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is less than the preset ratio, and the number of times of looping and performing clustering is less than or equal to the preset number of executions, then continue to perform clustering processing on the multiple retrieval pairs until the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is greater than or equal to the preset ratio, or the number of times of looping and performing clustering is greater than the preset number of executions.

[0093] In some embodiments, if the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is less than the preset ratio, and the number of times of re-performing multiple rounds of clustering (i.e., the number of times of looping and performing clustering) is less than or equal to the preset number of executions, then it is possible to re-perform multiple rounds of clustering on the multiple retrieval pairs corresponding to the first preset text until the ratio of the number of the obtained target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is greater than or equal to the preset ratio, or the number of times of re-performing multiple rounds of clustering is greater than the preset number of executions.

[0094] For example, if the total number of candidate seed retrieval pairs corresponding to the first preset text is 10, the number of target seed retrieval pairs is 2, and the preset ratio is 0.6, and the preset number of execution times is 5 times, then the proportion of the target seed retrieval pairs in the candidate seed retrieval pairs is 2 / 10 = 0.2, that is, the proportion of the target seed retrieval pairs is less than the preset ratio. Then, when the number of times of re - executing multiple rounds of clustering is less than or equal to 5 times, multiple rounds of clustering can be performed again on the multiple retrieval pairs corresponding to the first preset text until the proportion of the target seed retrieval pairs corresponding to the first preset text in the candidate seed retrieval pairs is greater than the preset ratio of 0.6, or the number of times of re - executing multiple rounds of clustering is greater than 5 times.

[0095] In some embodiments, the above - mentioned method further includes: if the number of times of loop - executing clustering is greater than the preset number of execution times, then process the second preset text among the multiple preset texts to obtain multiple retrieval pairs corresponding to the second preset text; perform clustering processing on the multiple retrieval pairs corresponding to the second preset text to determine the target seed retrieval pairs corresponding to the second preset text, and determine the retrieval content of the target seed retrieval pairs corresponding to the second preset text as the event determination information corresponding to the first question information.

[0096] In some embodiments, if after multiple re - clusterings (such as the number of times of loop - executing clustering is greater than the preset number of execution times), the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is still less than the preset ratio, then any preset text other than the first preset text (i.e., the second preset text) can be re - selected from the multiple preset texts.

[0097] Similarly, the condition determination model processes the second preset text to obtain multiple retrieval pairs corresponding to the second preset text; and performs multiple rounds of clustering on the multiple retrieval pairs corresponding to the second preset text to determine multiple candidate seed retrieval pairs and target seed retrieval pairs corresponding to the second preset text. If the ratio of the number of target seed retrieval pairs corresponding to the second preset text to the number of multiple candidate seed retrieval pairs corresponding to the second preset text is greater than or equal to the preset ratio, then determine the retrieval content of the target seed retrieval pairs as the event determination information; if the ratio of the number of target seed retrieval pairs corresponding to the second preset text to the number of multiple candidate seed retrieval pairs is less than the preset ratio, and the number of times of loop - executing clustering is less than or equal to the preset number of execution times, then continue to perform clustering processing on the multiple retrieval pairs. And so on, until the ratio of the number of target seed retrieval pairs corresponding to the second preset text to the number of multiple candidate seed retrieval pairs is greater than or equal to the preset ratio, or the number of times of re - executing multiple rounds of clustering is greater than the preset number of execution times.

[0098] If a target seed retrieval pair that can be used as the event determination information corresponding to the first question information cannot be obtained from the second preset text, continue to obtain any preset text (i.e., the third preset text) other than the first preset text and the second preset text from multiple preset texts, and process the third preset text... and so on, until a target seed retrieval pair corresponding to any preset text in the multiple preset texts can be used as the event determination information corresponding to the first question information, or until all preset texts cannot provide the event determination information corresponding to the first question information.

[0099] In some embodiments, the above method further includes: the event determination information is used to determine whether the first event information is established; in the case where it is determined that the first event information is established, the first event information is the target event information corresponding to the first question information; in the case where it cannot be determined whether the first event information is established, the event determination information is further used to determine the second question information to obtain the reply information of the user to the second question information; the reply information is used to determine the second event information corresponding to the first question information, and the second event information is the target event information corresponding to the first question information.

[0100] In some embodiments, after determining the event determination information corresponding to the first question information based on the preset text, it can be determined whether the first event information determined by the legal event detection model is accurate based on the event determination information. When the first question information does not lack necessary information, it can be determined whether the first event information is established according to the first question information and the event determination information, where the necessary information refers to the information used to accurately identify legal events. If the first event information is established, the first event information is determined as the target event information corresponding to the first question information.

[0101] When the first question information lacks necessary information, it is impossible to determine whether the first event information is established according to the first question information and the event determination information. At this time, the central control system can determine and output the second question information (or called information slot) based on the event determination information, and obtain the reply information input by the user to the second question information; thus, the legal event detection model can supplement and improve the first question information according to the reply information, and perform legal event detection on the supplemented first question information again to obtain the second event information. The second event information determined by the legal event detection model may be the same as the first event information or different from the first event information. If the second event information is the same as the first event information, the second event information or the first event information can be determined as the target event information; if the second event information is different from the first event information, the newly detected second event information is determined as the target event information.

[0102] In some embodiments, after determining the final target event information, the central control system may call other legal tools to perform legal case retrieval or legal provision retrieval based on the target event information to provide advisory opinions to users.

[0103] Taking the first question information entered by the user in the user interface of the central control system as "I signed a house lease contract, but the landlord did not deliver the house on the agreed time. I have notified the landlord to terminate the contract, but I have already suffered a loss of 2,000 yuan. Can I ask the landlord to compensate for the loss?" as an example, if the first event information corresponding to the determined first question information is "contract termination and compensation", and the event determination information corresponding to the determined first question information is: "The tenant signed a house lease contract with the landlord", "The landlord did not deliver the house on time", "The tenant has notified the landlord to terminate the contract", "The tenant has economic losses", then based on the first question information and the above event determination information, it can be determined that there is no lack of necessary information in the first question information, that is, the content corresponding to the above event determination information can be found in the first question information; and, the content in the first question information conforms to all the event determination information, so it can be determined that the above first event information is established, that is, the first event information is accurate. Thus, the central control system can call other legal tools to retrieve and integrate legal cases or legal provisions related to "contract termination and compensation", obtain the final legal advisory opinion and output it to the user.

[0104] Taking the first question information entered by the user in the user interface of the central control system as "I signed a three-year commercial lease contract with Party A to rent the other party's storefront, but because I can't make money now, can I ask to end the lease contract in advance?" as an example, in this first question information, since the reason for the business loss is unknown, the basis for the parties to request contract termination may be contract stalemate or change of circumstances. Therefore, after the first event information corresponding to the first question information determined by the legal event detection model (such as determining that the first event information is "contract termination in case of contract stalemate"), the event determination information corresponding to "contract termination in case of contract stalemate" can be determined by continuing to use the establishment condition determination model, so as to determine the necessary information missing in the first question information through this event determination information.

[0105] If the event determination information corresponding to "Contract rescission in case of contract deadlock" determined by the condition determination model for establishment is: 1) Although the breaching party breached the contract first, it was not a malicious breach; 2) If the breaching party continues to perform the contract, it will be unfair to the breaching party; 3) The non-breaching party's refusal to rescind the contract violates the principle of good faith. And according to the above first question information, it can be determined that the user is the breaching party, but the reason for the user's business loss cannot be determined. Therefore, it cannot be determined whether the user belongs to a malicious breach, nor can it be determined whether it is unfair to the user to continue to perform the contract. Therefore, the central control system can generate corresponding second question information based on the above 1) and 2) event determination information and output it to the user. For example, the second question information can be "May I ask what caused your business loss?" and "Have you fully negotiated with the other party?" etc.

[0106] If the reply information input by the user to the above second question information is: "Because my capital chain broke and I couldn't continue operating, and after negotiating with the other party, he insisted that I pay three years' rent", then the legal event detection model can supplement the first question information based on the reply information to obtain the target question information. For example, "I signed a three-year shop lease contract with Party A to rent the other party's storefront, but because my capital chain broke and I couldn't continue operating, and after negotiating with the other party, he insisted that I pay three years' rent. Can I request to end the lease contract in advance?" Then, the legal event detection model can determine that the second event information corresponding to the target question information is "Contract rescission in case of contract deadlock". Therefore, since the second event information is the same as the first event information, it can be determined that the target event information is "Contract rescission in case of contract deadlock". Finally, the central control system can call other legal tools to retrieve and integrate legal cases or legal provisions related to "Contract rescission in case of contract deadlock", and obtain the final consultation opinion and output it to the user.

[0107] Applying the technical solution of the present application can determine the event determination information related to the first event information corresponding to the first question information input by the user in the preset text by means of multi-round clustering, determining confidence scores, etc., so as to improve the accuracy when determining the event determination information, and further accurately judge whether the first event information holds based on the event determination information; and when the necessary information is missing in the first question information and it is impossible to judge whether the first event information holds, an information slot can be generated according to the event determination information to guide the user to supplement the necessary information in the first question information, so that the corresponding legal event can be determined based on the complete question information, and then a more accurate legal consultation opinion can be provided to the user based on the legal event.

[0108] Figure 6 Schematic diagram of a device for determining legal event determination information provided by some embodiments of the present application. As Figure 6As shown in the figure, the determination device 600 for legal event determination information includes an acquisition module 610, a processing module 620, a clustering module 630, a first determination module 640, a second determination module 650, and a third determination module 660.

[0109] The acquisition module 610 is configured to acquire first event information corresponding to first question information input by a user, and determine a plurality of preset texts associated with the first event information in a preset database based on the first event information.

[0110] The processing module 620 is configured to process a first preset text among the plurality of preset texts to obtain a plurality of retrieval pairs corresponding to the first preset text.

[0111] Wherein, each retrieval pair among the plurality of retrieval pairs includes a retrieval identifier and retrieval content.

[0112] The clustering module 630 is configured to perform clustering processing on the plurality of retrieval pairs corresponding to the first preset text to obtain a plurality of candidate seed retrieval pairs.

[0113] The first determination module 640 is configured to determine a target confidence score for each candidate seed retrieval pair among the plurality of candidate seed retrieval pairs.

[0114] Wherein, the target confidence score is related to the degree of association between each candidate seed retrieval pair and the first event information, and the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair.

[0115] The second determination module 650 is configured to determine candidate seed retrieval pairs with a target confidence score greater than or equal to a preset score among the plurality of candidate seed retrieval pairs as target seed retrieval pairs.

[0116] The third determination module 660 is configured to determine the retrieval content of the target seed retrieval pair as event determination information corresponding to the first question information.

[0117] In some embodiments, the clustering module 630 is specifically configured to: in the clustering process of the current round, obtain a plurality of first seed retrieval pairs and a plurality of unclustered first retrieval pairs obtained in the clustering process of the previous round; wherein, the plurality of retrieval pairs include a plurality of first seed retrieval pairs and a plurality of first retrieval pairs, and the sum of the number of the plurality of first seed retrieval pairs and the plurality of first retrieval pairs is equal to a preset number; perform clustering processing on the plurality of first seed retrieval pairs and the plurality of first retrieval pairs to obtain a plurality of second retrieval pairs corresponding to each legal category among the plurality of legal categories; among the plurality of second retrieval pairs corresponding to each legal category, determine a second seed retrieval pair corresponding to each legal category, and determine a plurality of unclustered second retrieval pairs in the plurality of retrieval pairs according to the preset number, the plurality of first retrieval pairs, and the plurality of second seed retrieval pairs; perform clustering processing on the plurality of second seed retrieval pairs and the plurality of second retrieval pairs in the next round of clustering process until there are no unclustered retrieval pairs in the plurality of retrieval pairs, and obtain a plurality of candidate seed retrieval pairs determined in the last round of clustering process; wherein, each candidate seed retrieval pair in the plurality of candidate seed retrieval pairs corresponds to a different legal category.

[0118] In some embodiments, the clustering module 630 is specifically configured to: in each legal category, determine a semantic vector corresponding to each second retrieval pair; according to the semantic vector corresponding to each second retrieval pair, determine the total distance between each first semantic vector and the second semantic vector in each legal category; wherein, the first semantic vector is a semantic vector corresponding to any second retrieval pair among the plurality of second retrieval pairs corresponding to each legal category, and the second semantic vector is a semantic vector corresponding to a second retrieval pair other than the second retrieval pair corresponding to the first semantic vector; determine a target semantic vector with the shortest total distance from the plurality of first semantic vectors to the second semantic vector, and determine the second retrieval pair corresponding to the target semantic vector as the second seed retrieval pair corresponding to each legal category.

[0119] In some embodiments, the first determination module 640 is specifically configured to: determine a first confidence score for each candidate seed retrieval pair according to the correlation degree between each candidate seed retrieval pair and the first event information; wherein, the higher the correlation degree between each candidate seed retrieval pair and the first event information, the higher the first confidence score of each candidate seed retrieval pair; determine the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair, and determine a second confidence score for each candidate seed retrieval pair according to the number of retrieval pairs corresponding to each candidate seed retrieval pair; wherein, the more the number of retrieval pairs corresponding to each candidate seed retrieval pair, the higher the second confidence score of each candidate seed retrieval pair; based on the first confidence score and the second confidence score of each candidate seed retrieval pair, determine the target confidence score of each candidate seed retrieval pair.

[0120] In some embodiments, the third determination module 660 is specifically configured to: if the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is greater than or equal to a preset ratio, determine the retrieval content of the target seed retrieval pair as event determination information.

[0121] In some embodiments, the clustering module 630 is further configured to: if the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is less than the preset ratio, and the number of times of loop execution of clustering is less than or equal to the preset execution times, continue to perform clustering processing on the multiple retrieval pairs until the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is greater than or equal to the preset ratio, or the number of times of loop execution of clustering is greater than the preset execution times.

[0122] In some embodiments, the processing module 620 is further configured to: if the number of times of loop execution of clustering is greater than the preset execution times, process the second preset text in the multiple preset texts to obtain multiple retrieval pairs corresponding to the second preset text; where the second preset text is any one of the multiple preset texts except the first preset text; the clustering module 630 is further configured to: perform clustering processing on the multiple retrieval pairs corresponding to the second preset text to determine the target seed retrieval pair corresponding to the second preset text, and determine the retrieval content of the target seed retrieval pair corresponding to the second preset text as the event determination information corresponding to the first question information.

[0123] In some embodiments, the event determination information is used to determine whether the first event information is established; in the case where it is determined that the first event information is established, the first event information is the target event information corresponding to the first question information; in the case where it is impossible to determine whether the first event information is established, the event determination information is further used to determine the second question information to obtain the response information of the user to the second question information; the response information is used to determine the second event information corresponding to the first question information, and the second event information is the target event information corresponding to the first question information.

[0124] Figure 7 The figure is a schematic diagram of an electronic device provided in some embodiments of the present application. In some embodiments, the electronic device includes one or more processors and a memory. The memory is configured to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining legal event determination information in the above embodiments.

[0125] As Figure 7 shown, the electronic device 700 includes: a processor 701 and a memory 702. Exemplarily, the electronic device 700 may further include: a communication interface 703 and a communication bus 704.

[0126] Among them, the processor 701, the memory 702, and the communication interface 703 complete their mutual communication through the communication bus 704. The communication interface 703 is used to communicate with network elements of other devices such as clients or other servers.

[0127] In some embodiments, the processor 701 is used to execute the program 705, and specifically can execute the relevant steps in the embodiments of the method for determining legal event determination information described above. Specifically, the program 705 may include program code, and the program code includes computer-executable instructions.

[0128] Exemplarily, the processor 701 may be a central processing unit CPU, or a specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. One or more processors included in the electronic device 700 may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0129] In some embodiments, the memory 702 is used to store the program 705. The memory 702 may include high-speed RAM memory, and may also include non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory.

[0130] The program 705 can specifically be called by the processor 701 to cause the electronic device 700 to perform the operations of the method for determining legal event determination information.

[0131] Some embodiments of the present application provide a computer-readable storage medium. The computer-readable storage medium stores at least one executable instruction. When the executable instruction runs on the electronic device 700, it causes the electronic device 700 to execute the method for determining legal event determination information in the above embodiments.

[0132] The executable instruction can specifically be used to cause the electronic device 700 to perform the operations of the method for determining legal event determination information.

[0133] For example, the computer-readable storage medium may be a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0134] The beneficial effects achievable by the computer-readable storage medium provided in some embodiments of the present application can refer to the beneficial effects in the corresponding method for determining legal event determination information provided above, and will not be elaborated here.

[0135] It should be noted that in the application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0136] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0137] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus or device), or used in conjunction with these instruction execution systems, apparatus or devices.

[0138] For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus or device.

[0139] More specific examples of computer-readable media (non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM).

[0140] In addition, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, deciphering or otherwise processing as appropriate, and then stored in a computer memory. It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof.

[0141] In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0142] The above-described embodiments of the present application do not constitute a limitation on the protection scope of the present application.

Claims

1. A method for determining legal event judgment information, characterized in that Including: Obtain first event information corresponding to first question information input by a user, and determine a plurality of preset texts associated with the first event information in a preset database based on the first event information; Process a first preset text to obtain a plurality of retrieval pairs corresponding to the first preset text; wherein, the first preset text is any one of the plurality of preset texts, and each of the plurality of retrieval pairs includes a retrieval identifier and retrieval content; Perform clustering processing on the plurality of retrieval pairs corresponding to the first preset text to obtain a plurality of candidate seed retrieval pairs; Determine a target confidence score for each of the plurality of candidate seed retrieval pairs; wherein, the target confidence score is related to the degree of association between each of the candidate seed retrieval pairs and the first event information, and the number of retrieval pairs in the legal category corresponding to each of the candidate seed retrieval pairs; Determine candidate seed retrieval pairs with a target confidence score greater than or equal to a preset score among the plurality of candidate seed retrieval pairs as target seed retrieval pairs; Determine the retrieval content of the target seed retrieval pairs as event determination information corresponding to the first question information; Wherein, the event determination information is used to determine whether the first event information is established; in the case where it is determined that the first event information is established, the first event information is the target event information corresponding to the first question information; in the case where it is impossible to determine whether the first event information is established, the event determination information is further used to determine second question information to obtain a reply message of the user to the second question information; the reply message is used to determine second event information corresponding to the first question information, and the second event information is the target event information corresponding to the first question information.

2. The method according to claim 1, wherein The performing clustering processing on the plurality of retrieval pairs corresponding to the first preset text to obtain a plurality of candidate seed retrieval pairs includes: In the clustering process of the current round, obtain a plurality of first seed retrieval pairs obtained in the clustering process of the previous round and a plurality of first retrieval pairs that have not been clustered; wherein, the plurality of retrieval pairs include the plurality of first seed retrieval pairs and the plurality of first retrieval pairs, and the sum of the number of the plurality of first seed retrieval pairs and the plurality of first retrieval pairs is equal to a preset number; Perform clustering processing on the plurality of first seed retrieval pairs and the plurality of first retrieval pairs to obtain a plurality of second retrieval pairs corresponding to each of the plurality of legal categories; Among the plurality of second retrieval pairs corresponding to each of the legal categories, determine second seed retrieval pairs corresponding to each of the legal categories, and determine a plurality of second retrieval pairs that have not been clustered in the plurality of retrieval pairs according to the preset number, the plurality of first retrieval pairs, and the plurality of second seed retrieval pairs; In the next round of clustering process, perform clustering processing on the multiple second seed retrieval pairs and the multiple unclustered second retrieval pairs until there are no unclustered retrieval pairs among the multiple retrieval pairs, and obtain the multiple candidate seed retrieval pairs determined in the last round of clustering process; wherein, each of the multiple candidate seed retrieval pairs corresponds to a different legal category.

3. The method according to claim 2, wherein Determining the second seed retrieval pair corresponding to each legal category among the multiple second retrieval pairs corresponding to each legal category includes: In each legal category, determine the semantic vector corresponding to each second retrieval pair; According to the semantic vectors corresponding to each second retrieval pair, determine the total distance between each first semantic vector and the second semantic vector in each legal category; wherein, the first semantic vector is the semantic vector corresponding to any second retrieval pair among the multiple second retrieval pairs corresponding to each legal category, and the second semantic vector is the semantic vector corresponding to the second retrieval pair other than the second retrieval pair corresponding to the first semantic vector; Determine the target semantic vector with the shortest total distance from the multiple first semantic vectors to the second semantic vector, and determine the second retrieval pair corresponding to the target semantic vector as the second seed retrieval pair corresponding to each legal category.

4. The method according to claim 1, wherein Determining the target confidence score of each candidate seed retrieval pair among the multiple candidate seed retrieval pairs includes: According to the degree of association between each candidate seed retrieval pair and the first event information, determine the first confidence score of each candidate seed retrieval pair; wherein, the higher the degree of association between each candidate seed retrieval pair and the first event information, the higher the first confidence score of each candidate seed retrieval pair; Determine the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair, and determine the second confidence score of each candidate seed retrieval pair according to the number of retrieval pairs corresponding to each candidate seed retrieval pair; wherein, the more the number of retrieval pairs corresponding to each candidate seed retrieval pair, the higher the second confidence score of each candidate seed retrieval pair; Based on the first confidence score and the second confidence score of each candidate seed retrieval pair, determine the target confidence score of each candidate seed retrieval pair.

5. The method according to claim 1, wherein Determining the retrieval content of the target seed retrieval pair as the event determination information corresponding to the first question information includes: If the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is greater than or equal to a preset ratio, then determine the retrieval content of the target seed retrieval pair as the event determination information.

6. The method according to claim 5, characterized in that, The method further includes: If the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is less than the preset ratio, and the number of times of loop execution of clustering is less than or equal to the preset number of execution times, then continue to perform clustering processing on the multiple retrieval pairs until the ratio of the number of target seed retrieval pairs to the number of multiple candidate seed retrieval pairs is greater than or equal to the preset ratio, or the number of times of loop execution of clustering is greater than the preset number of execution times.

7. The method according to claim 6, characterized in that, The method further includes: If the number of times the loop performs clustering is greater than the preset number of executions, process the second preset text among the multiple preset texts to obtain multiple retrieval pairs corresponding to the second preset text; wherein, the second preset text is any preset text among the multiple preset texts other than the first preset text. Perform clustering processing on the multiple retrieval pairs corresponding to the second preset text to determine the target seed retrieval pair corresponding to the second preset text, and determine the retrieval content of the target seed retrieval pair corresponding to the second preset text as the event determination information corresponding to the first question information.

8. A determining device for legal event determination information, characterized in that, Including: An acquisition module, configured to acquire the first event information corresponding to the first question information input by the user, and determine multiple preset texts associated with the first event information in the preset database based on the first event information. A processing module, configured to process the first preset text to obtain multiple retrieval pairs corresponding to the first preset text; wherein, the first preset text is any preset text among the multiple preset texts, and each of the multiple retrieval pairs includes a retrieval identifier and retrieval content. A clustering module, configured to perform clustering processing on the multiple retrieval pairs corresponding to the first preset text to obtain multiple candidate seed retrieval pairs. A first determination module, configured to determine the target confidence score of each of the candidate seed retrieval pairs among the multiple candidate seed retrieval pairs; wherein, the target confidence score is related to the degree of association between each candidate seed retrieval pair and the first event information, and the number of retrieval pairs in the legal category corresponding to each candidate seed retrieval pair. A second determination module, configured to determine the candidate seed retrieval pairs among the multiple candidate seed retrieval pairs whose target confidence scores are greater than or equal to the preset score as the target seed retrieval pairs. A third determination module, configured to determine the retrieval content of the target seed retrieval pair as the event determination information corresponding to the first question information. Wherein, the event determination information is used to determine whether the first event information holds; in the case where it is determined that the first event information holds, the first event information is the target event information corresponding to the first question information; in the case where it is impossible to determine whether the first event information holds, the event determination information is further used to determine a second question information to obtain a reply information of the user to the second question information; the reply information is used to determine a second event information corresponding to the first question information, and the second event information is the target event information corresponding to the first question information.

9. An electronic device, characterized in that, Including: One or more processors; And A memory, configured to: store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining legal event determination information according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, the method for determining legal event determination information according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Legal document generation method and system based on reading understanding and intention recognition model

    CN114297342A

  • Keyword mining method and device suitable for long document and medium

    CN115858773A