A method and apparatus for focus analysis

By adopting a multi-case analysis model in the auxiliary trial analysis system, the existing system's waste of resources and differences in effect due to the diversity of case causes is solved, and the generation of dispute focus across cases and the annotation of dialogue content is achieved, which reduces deployment costs and resource occupation.

CN113934810BActive Publication Date: 2025-07-01ALIBABA CLOUD COMPUTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010601373.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-29
Publication Date
2025-07-01
Estimated Expiration
2040-06-29

AI Technical Summary

Technical Problem

The existing auxiliary trial analysis system needs to set up multiple analysis models according to different cases, resulting in significant differences in resource waste and application effects.

Method used

Using a multi-case analysis model, through training of case samples with multiple preset cases, we can generate dispute focus across cases and label relevant dialogue content.

Benefits of technology

There is no need to deploy a large number of analytical models with designated cases, saving system resources, reducing deployment costs, and improving the versatility and efficiency of auxiliary trial analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113934810B_ABST
    Figure CN113934810B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for focus analysis, which relates to the technical field of computer artificial intelligence. The main objective of the present invention is to generate dispute focuses for cases of different case types and label relevant conversation content. The main technical solution of the present invention is as follows: collecting the designated case type of the current case and the conversation content; using a multi-case-type analysis model to analyze at least one dispute focus of the current case from the conversation content according to the designated case type, where the multi-case-type analysis model is trained by case samples of multiple preset case types; and obtaining focus sentences related to the at least one dispute focus in the conversation content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer artificial intelligence, and particularly to a method and device for focus analysis. Background Art

[0002] With the reform of the case-filing registration system, the number of cases has increased significantly, and the contradiction between "too many cases and too few staff" has become prominent. The existing trial system, trial capacity, and judicial service capacity of the court are no longer able to adapt to this situation. There is an urgent need to further improve the informatization level of the people's court, deepen the intensity of judicial disclosure, promote the reengineering of the trial process, and solve the problem of "too many cases and too few staff" in the people's court.

[0003] In the existing auxiliary trial analysis system during the trial process, relying on text and image recognition and semantic analysis technologies, it can automatically generate dispute focuses based on the conversations of roles such as judges, plaintiffs, and defendants in the trial according to the specified case type, and mark the conversation content related to the dispute focuses for the subsequent production of judgment documents. In this way, by providing intelligent auxiliary support and services, the case handling efficiency can be improved. However, the existing auxiliary trial analysis system needs to set corresponding analysis models according to different case types. Thus, during actual deployment, due to the large number of case types covered by judicial business, a large number of analysis models corresponding to different case types need to be deployed, consuming a large amount of machine resources. Moreover, due to the uneven sample data volumes of different case types, the actual application effects of analysis models for different case types also vary significantly. Summary of the Invention

[0004] In view of the above problems, the present invention proposes a method and device for focus analysis, and the main purpose is to generate dispute focuses for cases of different case types and mark the relevant conversation content.

[0005] To achieve the above object, the present invention mainly provides the following technical solutions:

[0006] On the one hand, the present invention provides a method for focus analysis, specifically including:

[0007] Collect the specified case type and conversation content of the current case;

[0008] Use a multi-case analysis model to analyze at least one dispute focus of the current case from the conversation content according to the specified case type, where the multi-case analysis model is trained by case samples of multiple preset case types;

[0009] Obtain focus statements related to the at least one dispute focus in the conversation content.

[0010] On the other hand, the present invention provides a device for focus analysis, specifically including:

[0011] An acquisition unit for acquiring the specified case type and conversation content of the current case;

[0012] An analysis unit for analyzing at least one controversial focus of the current case from the conversation content according to the specified case type obtained by the acquisition unit by using a multi-case-type analysis model, where the multi-case-type analysis model is trained by case samples of multiple preset case types;

[0013] A marking unit for obtaining focus statements related to at least one controversial focus obtained by the analysis unit from the conversation content.

[0014] On the other hand, the present invention provides a processor for running a program, where when the program runs, it executes the above-mentioned method for focus analysis.

[0015] By means of the above technical solution, a method and device for focus analysis provided by the present invention mainly realize cross-case-type generation of controversial focuses and marking of relevant conversation content by deploying a multi-case-type analysis model. During the actual court trial process, courts hearing cases of different case types can all call the multi-case-type analysis model to assist in trial analysis, automatically generate the controversial focuses of the current case according to the conversation content in the trial in combination with the specified case type, and obtain the statements related to the controversial focuses in the conversation content as focus statements. Compared with the existing auxiliary trial analysis system that customizes analysis models for specified case types, the present invention does not need to deploy a large number of analysis models for specified case types, and saves a large amount of system resources by deploying a multi-case-type analysis model.

[0016] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Description of the Drawings

[0017] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0018] Figure 1 Shows a flowchart of a method for focus analysis proposed by an embodiment of the present invention;

[0019] Figure 2 Shows a flowchart of the process of vectorizing the conversation content proposed by an embodiment of the present invention;

[0020] Figure 3Shows the flowchart of the process of analyzing the controversial focus of a case by the multi-case analysis model proposed in the embodiments of the present invention;

[0021] Figure 4 Shows the flowchart of the process of annotating focus statements proposed in the embodiments of the present invention;

[0022] Figure 5 Shows the block diagram of the composition of a focus analysis device proposed in the embodiments of the present invention;

[0023] Figure 6 Shows the block diagram of the composition of another focus analysis device proposed in the embodiments of the present invention. Detailed implementation manners

[0024] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0025] A focus analysis method provided by the embodiments of the present invention is mainly applied to an auxiliary court trial analysis system. The main purpose is to increase the generality of the analysis model, so that the analysis model can analyze the controversial focus of court trial cases across case types, and reduce the occupancy of system resources by the analysis model in actual deployment, thereby reducing the deployment cost of the auxiliary court trial analysis system. Specifically, the implementation steps of this method in the actual court trial process are as Figure 1 shown, including:

[0026] Step 101, collect the specified case type of the current case and the conversation content.

[0027] Among them, the specified case type is generally set manually according to the current court trial case. The conversation content refers to the conversation information during the trial of the current court trial case, and the conversation content can generally be obtained through voice conversion technology or obtained by manual recording by the court clerk.

[0028] It should be noted that the conversation content in this step includes not only the statement information of the conversation, but also the role information of the speaker. Therefore, each piece of conversation content consists of two parts: role information and conversation statement. Among them, the role information generally includes judges, plaintiffs, defendants, etc. Adding role information to the conversation content helps to understand the semantics of the conversation statements, so as to more accurately analyze the controversial focus between the plaintiff and the defendant.

[0029] Step 102, use the multi-case analysis model to analyze at least one controversial focus of the current case from the conversation content according to the specified case type.

[0030] Among them, the multi-case analysis model is used to analyze the dispute focus and the corresponding dispute focus content of the case based on the conversation content during the court trial. The difference between it and the customized designated case analysis model is that the multi-case analysis model is applicable to multiple preset case types. When multiple courts hear cases of different case types, the multi-case analysis model can be used for auxiliary analysis without searching for the analysis model for the designated case type. In the actual court trial process, by calling the multi-case analysis model, there is no need to search for and select the analysis model required for the designated case type. Instead, the case type of the case being tried is input into the multi-case analysis model, and the multi-case analysis model dynamically analyzes the conversation content during the court trial according to the case type, determines and outputs the dispute focus and the dispute focus content.

[0031] In addition, when training the multi-case analysis model, it is not necessary to classify the case samples according to the case type. Instead, case samples that conform to the preset case types can all be used to train the multi-case analysis model. Among them, the preset case types are the case types of cases that can be processed specified in advance by the model, and the number of preset case types can also be adjusted according to requirements during the actual application process.

[0032] The multi-case analysis model in this step can first determine multiple dispute foci related to the designated case type of the current court trial case, and then, by analyzing the conversation content, determine the probability that it falls within these dispute foci. Thus, the dispute focus with the highest probability is taken as the dispute focus of the current court trial case. Among them, the specific method for determining the dispute focus is not limited to setting a probability threshold, determining the one greater than the probability threshold as the dispute focus, or taking the one with the maximum probability value as the dispute focus, or comprehensively setting the determination conditions of the above two methods to determine the dispute focus.

[0033] Furthermore, according to the determined dispute focus, the dispute focus content can be further generated using the text generation algorithm based on the conversation content, and the dispute focus content corresponds to the dispute focus one by one.

[0034] Step 103: Obtain focus statements related to at least one dispute focus from the conversation content.

[0035] The purpose of obtaining the focus statements in this step is to assist the trial personnel in viewing the key content in the conversation during the court trial, so as to check the generated dispute focus and the content of the dispute focus. Specifically, the acquisition of the focus statements in this step can be converted into an operation of classifying the conversation statements in the conversation content. The classification categories are determined according to the number of dispute foci, including dispute foci and non-dispute foci. For example, if there are 2 dispute foci, then the corresponding categories are 3 types, namely Dispute Focus 1, Dispute Focus 2, and Non-dispute Focus. According to the specific classification categories, use a neural network model to classify the conversation statements in the conversation content, and then mark the conversation statements related to each dispute focus, that is, the focus statements.

[0036] Through the above Figure 1 It can be seen from the steps in the above embodiments that the embodiments of the present invention are mainly applied to the court trial process. According to the conversation content during the trial, it assists the trial personnel in determining the dispute focus of the case and marking the focus statements related to the dispute focus, which is convenient for the trial personnel to check. The difference between the embodiments of the present invention and the existing auxiliary trial system is that there is no need to manually select an analysis model for a specified case type. The multi-case-type analysis model in this embodiment can analyze across case types. The trial personnel only need to input the case type of the current case into the multi-case-type analysis model. In this way, only this multi-case-type analysis model needs to be deployed during actual deployment, reducing the resource requirements of the system, that is, reducing the deployment cost. At the same time, when the embodiments of the present invention collect the conversation content, the role information of the conversation is added to the conversation content, so that when generating the dispute focus, the semantics of the conversation statements can be more accurately understood based on the role of the conversation person, thereby more accurately determining the dispute focus of the case.

[0037] Furthermore, for the above Figure 1 It can be seen from the method of the above-mentioned focus analysis that the execution of the embodiments of the present invention requires first constructing a multi-case-type analysis model and training the multi-case-type analysis model with the existing case samples marked with case types, so that it can accurately generate the dispute focus and the corresponding content of the dispute focus. For this reason, the specific process of constructing the multi-case-type analysis model in the embodiments of the present invention is as follows:

[0038] First, determine the case type set, which includes multiple preset case types. That is, determine the types of case types applicable to the multi-case-type analysis model. This case type set can be custom-set according to user needs.

[0039] After that, construct a multi-case-type analysis model based on multiple preset case types. This multi-case-type analysis model can be constructed based on a neural network. The input of the model is the case sample of the specified case type, that is, the conversation content during the trial of the case and the specified case type of the case. It should be noted that the specified case type is one of the preset case types in the case type set. The output of the model is the dispute focus of the case sample and the corresponding content of the dispute focus.

[0040] Finally, the multi-case analysis model is trained using case samples marked with case causes. Since it can analyze cases of multiple preset case causes, in order to avoid interference, during the training of the multi-case analysis model, it is necessary to mask other case causes of the multiple preset case causes according to the designated case cause of the case sample, so as to train the loss function of the multi-case analysis model under the designated case cause. Specifically, during the backpropagation iterative learning process of the neural network, only the loss under the designated case cause is iterated, and for other case causes, through the masking operation (i.e., multiplying by 0), it is removed from the overall loss, and then backpropagation and gradient calculation are performed to determine the relevant parameters of the loss function.

[0041] Furthermore, the multi-case analysis model in the embodiment of the present invention analyzes the dispute focus according to the dialogue content for the designated case cause. Therefore, it is necessary to pre-declare the dispute focuses corresponding to different case causes. In this embodiment, the correspondence between the case cause and the dispute focus is realized by constructing a legal knowledge graph. In the case of multiple case causes, each case cause needs to sort out its own legal knowledge graph. Usually, the legal knowledge graph has a tree-like or network-like topological structure for expressing complex legal relationships. Each legal knowledge graph includes legal entities and legal relationships. Among them, the legal entities are used to represent the dispute focuses of the designated case cause. Generally, one case cause has multiple dispute focuses, and the legal relationships are used to represent the association relationships between different dispute focuses.

[0042] The legal knowledge graph based on the case cause can be constructed manually by legal experts or automatically by algorithms, including but not limited to algorithms such as TransE and TransR to construct the graph. Through the constructed legal knowledge graph, vector representation of each legal entity can be realized, that is, vector representation of each dispute focus corresponding to different case causes.

[0043] According to the multi-case analysis model constructed according to the above embodiments and the pre-set correspondence between the designated case cause and the dispute focus (legal knowledge graph), the embodiment of the present invention will further Figure 1 illustrate its preferred implementation manner step by step for the embodiments shown:

[0044] First, for step 101, the designated case cause of the current case and the dialogue content are collected. Among them, the designated case cause is the information input manually, and the dialogue content is to convert the dialogue information in the current court trial process into vectorized information recognizable by the multi-case analysis model. The specific process is as Figure 2 shown, including:

[0045] Step 201, obtain the dialogue content of the current case.

[0046] The content of the conversation includes at least one conversation statement and the corresponding preset role identifier. Among them, the preset role identifier can represent multiple roles. Assuming that the roles participating in the conversation in the conversation content include: judge, plaintiff 1, plaintiff 2, plaintiff's attorney, defendant's attorney, and the preset role identifiers are A, B, C, then the A identifier represents the judge, the B identifier represents plaintiff 1, plaintiff 2, and the plaintiff's attorney, and the C identifier represents the defendant's attorney. The finally obtained conversation content is represented as {(role A, conversation statement 1), (role B, conversation statement 2), (role A, conversation statement 3)…(role X, conversation statement n)}, where X is one of the roles A, B, C.

[0047] Step 202: Generate sentence vectors based on each conversation statement and the corresponding preset role identifier.

[0048] This step is to represent the conversation content in a vectorized manner. Among them, the conventional vectorized representation method is to segment the conversation statements, convert each segment into a word vector through a dictionary, and represent the preset role identifier in a preset manner to obtain the vector representation of the conversation statements, such as a vector sequence {e1, e2, e3, …, em}, where e represents a segment and m represents the number of segments. Then, the vector of the preset role identifier is concatenated with the vector sequence of the conversation statements, that is, the vector of the preset role identifier is concatenated with each word vector, so as to obtain a sentence vector with role information, which is assumed to be represented as {E1, E2, E3, …, Em}.

[0049] It should be noted that the method of integrating role information with the information of conversation statements is not limited to the above vector concatenation, and can also be other methods such as vector addition and multiplication.

[0050] Step 203: Use a neural network model to encode the sentence vectors corresponding to at least one conversation statement to obtain the conversation vector corresponding to the conversation content.

[0051] Among them, the neural network model can adopt a deep learning neural network, including but not limited to neural network models such as CNN, LSTM, Transformer, and Attention. The sentence vectors {E1, E2, E3, …, Em} in the previous step are encoded through the neural network model to obtain a vector v. According to the sentence vectors corresponding to the obtained historical conversation content, the conversation vectors representing the overall conversation content can be obtained, {v1, v2, v3, …, vn}, where n represents the number of sentence vectors.

[0052] Step 204: Perform semantic encoding on the conversation vector according to the context relationship of the conversation statements in the conversation content to obtain the conversation semantic vector corresponding to the conversation content.

[0053] This step is to perform secondary encoding on the dialogue vectors obtained in the previous step based on the context relationship between dialogue statements, mapping {v1, v2, v3, …, vn} to dialogue semantic vectors {h1, h2, h3, …, hn}, where v and h correspond one by one, and h is a further representation of v. It can be seen that more context semantic information of dialogue statements is incorporated into the dialogue semantic vectors.

[0054] As can be seen from the above steps, when the embodiments of the present invention perform vector representation on dialogue content, not only role information is incorporated into dialogue statements, but also encoding is performed based on the context relationship of dialogue statements to obtain dialogue semantic vectors containing context semantic information, so that the vector representation of dialogue content can more accurately express the semantics in the dialogue, thereby providing more abundant material information for subsequent analysis.

[0055] Secondly, for step 102, the specific process of analyzing the dispute focus of a case by the multi-case analysis model is as Figure 3 shown, including:

[0056] Step 301: Determine multiple dispute focuses of the specified case type according to the legal knowledge graph.

[0057] This step is to use the constructed legal knowledge graph to search for the dispute focuses associated with the specified case type according to the input specified case type, that is, the case type of the current court trial.

[0058] Step 302: The multi-case analysis model analyzes the dialogue semantic vectors corresponding to the dialogue content and determines the mapping probabilities of the dialogue semantic vectors on multiple dispute focuses.

[0059] This step requires pooling the dialogue semantic vectors {h1, h2, h3, …, hn}, including but not limited to pooling methods such as attention, max-pool, mean-pool, etc., to aggregate information to obtain a dialogue representation, represented by vector H, which can be understood as a representation that synthesizes all the above dialogue information. This representation can theoretically strengthen key information, such as the main involved plot and the defense process, etc., and weaken some irrelevant information, such as some flowing words. Then, use the fully connected layer to process vector H to obtain a mapping probability vector, and the probability value of each dimension in the mapping probability vector is the mapping probability of the dialogue semantic vector on the dispute focus corresponding to that dimension. Among them, the dimension of the mapping probability vector is the number of dispute focuses associated with the specified case type.

[0060] Step 303: Determine the dispute focus contained in the dialogue content according to the mapping probability.

[0061] This step can generally be determined by setting a threshold. That is, when there is a mapping probability greater than this threshold among the mapping probabilities corresponding to each controversial focus, the corresponding controversial focus is determined to be the controversial focus contained in the current conversation content of the case (that is to say, as the trial conversation content continues to increase, this controversial focus may increase or change). However, by using the threshold method, when the threshold is set too high, it may lead to the inability to determine the controversial focus.

[0062] In this regard, this step adds a determination condition for the maximum mapping probability on the basis of setting the threshold. That is, when it is judged that there is no mapping probability greater than the threshold among the mapping probabilities corresponding to each controversial focus, the controversial focus corresponding to the maximum mapping probability is determined to be the controversial focus contained in this conversation content.

[0063] Furthermore, after the multi-case analysis model determines the controversial focus of the case, the embodiment of the present invention can also generate corresponding controversial focus content according to the determined controversial focus in combination with the conversation content. The controversial focus content corresponds one-to-one with the controversial focus, and the specific process of its generation includes:

[0064] Step 304: Determine the focus vector corresponding to each controversial focus according to the legal knowledge graph.

[0065] This step is based on the determined controversial focus to obtain the corresponding focus vector, and this focus vector is the vector representation of the corresponding legal entity in the legal knowledge graph.

[0066] Step 305: Fuse the dialogue semantic vector with each focus vector respectively to obtain at least one new dialogue content vector.

[0067] Among them, the ways of fusing the dialogue semantic vector {h1, h2, h3, …, hn} with the focus vector include but are not limited to gating mechanism, splicing and other ways. In this way, a dialogue representation incorporating the prior information of the controversial focus is obtained, that is, the new dialogue content vector, denoted as {H1, H2, H3, …, Hn}. When there are multiple controversial foci, multiple corresponding new dialogue content vectors will be obtained.

[0068] Step 306: Use the text generation algorithm to process the new dialogue content vector to generate the corresponding controversial focus content for each controversial focus.

[0069] The text generation algorithm in this step can adopt the seq2seq model of deep learning to process the new dialogue content vector and output the controversial focus content corresponding to the controversial focus. Similarly, when there are multiple controversial foci, multiple corresponding controversial focus contents will also be obtained.

[0070] However, it should be noted that the execution of the above steps is a dynamic analysis process. That is, as the number of dialogue statements increases, the focus of controversy in the entire dialogue content will also change accordingly, and the change in the focus of controversy will further affect the corresponding content of the focus of controversy.

[0071] Finally, for step 103, the specific process of obtaining the focus statements related to at least one focus of controversy is as Figure 4 shown, including:

[0072] Step 401: Determine the statement classification labels according to at least one focus of controversy.

[0073] Among them, the statement classification labels include the focus statements corresponding to the focus of controversy and non-focus statements.

[0074] In this embodiment, the process of annotating focus statements can be converted into the process of classifying dialogue statements. To perform classification, it is first necessary to determine the categories of classification, that is, the statement classification labels. Suppose there are 2 foci of controversy currently determined in the court trial case, namely Focus 1 and Focus 2. Then, the statement classification labels determined in this step are 3, namely Focus 1 statements, Focus 2 statements, and non-focus statements.

[0075] Step 402: Process the dialogue semantic vectors using a neural network model to determine the statement classification labels for each dialogue statement in the dialogue content.

[0076] Among them, the neural network model is used to classify each dialogue statement in the dialogue semantic vectors {h1, h2, h3,..., hn}, that is, to determine the probability of each dialogue statement corresponding to each statement classification label, and determine the statement classification label with the highest probability as the category to which the dialogue statement belongs.

[0077] Step 403: Highlight the dialogue statements with the statement classification label of focus statements according to the focus of controversy.

[0078] This step is to prominently display the statements related to the focus of controversy selected from the dialogue statements for the convenience of real-time viewing and verification by the trial personnel. Among them, the focus statements of different foci of controversy can be highlighted in different colors or brightness to distinguish different foci of controversy.

[0079] Furthermore, after marking the focus statements, this embodiment can also mark the focus of controversy and the corresponding content of the focus of controversy on the dialogue content, that is, mark the focus of controversy and the content of the focus of controversy corresponding to the highlighted focus statements. In this way, when the trial personnel view the dialogue content, they can not only determine the focus of controversy and the content of the focus of controversy according to the color or brightness of the focus statements, but also view the focus of controversy and the content of the focus of controversy through more intuitive marked information.

[0080] With the improvement and popularization of China's legal system, the focus analysis method proposed in the embodiments of the present invention can be applied to various scenarios where there are legal disputes or non-legal rule disputes. For example, when applied to an e-commerce platform, it can judge for both buyers and sellers according to the platform's rules, and can also be used to handle various complaints and make handling results such as penalties and mediations. It can also be applied to relevant legal institutions, such as courts, procuratorates, law firms, and the legal affairs, mediation and arbitration departments, neighborhood committees, civil affairs bureaus, etc. of companies and enterprises. In courts and procuratorates, the embodiments of the present invention can assist in case trials and form pre-plans. In law firms and the legal affairs of companies and enterprises, they can simulate and review the court trial process to find important viewpoints and relevant content in the case. In government departments such as mediation and arbitration departments, neighborhood committees, and civil affairs bureaus, the multi-case analysis model can also be customized and trained according to the specific laws and regulations specified by each department, so as to assist in handling specific cases or events in different departments in a targeted manner, and assist in marking or highlighting the key content in the case.

[0081] Through the above Figures 2 - 4 As can be seen from the descriptions corresponding to the steps of the above embodiments, a focus analysis method proposed in the embodiments of the present invention realizes cross-case auxiliary court

[0082] trial analysis based on the constructed legal knowledge graph and multi-case analysis model. By analyzing the conversations of roles such as judges, plaintiffs, and defendants, it automatically generates the dispute focus corresponding to the specified case type, as well as the content of the dispute focus, and highlights the focus statements related to the dispute focus. At the same time, it marks the dispute focus and the content of the dispute focus on the focus statements correspondingly for easy viewing and verification. When the multi-case analysis model is actually applied and deployed, even when cases of multiple different case types are being tried simultaneously, the multi-case analysis model can be called simultaneously to assist in court trial analysis, so there is no need to deploy multiple analysis models for different case types, which can effectively reduce the occupation of system resources and lower the deployment cost.

[0083] Furthermore, as an implementation of the above Figures 1 - 4 shown method, the embodiments of the present invention provide a focus analysis device. The main purpose of this device is to generate dispute focuses for cases of different case types and mark relevant conversation content. For the convenience of reading, the details of the foregoing method embodiments will not be described one by one in the embodiments of this device, but it should be clear that the device in this embodiment can correspondingly implement all the content of the foregoing method embodiments. The device is as Figure 5 shown and specifically includes:

[0084] A collection unit 51, configured to collect the specified case type of the current case and the conversation content;

[0085] An analysis unit 52, configured to analyze at least one disputed focus of the current case from the conversation content according to a specified case type obtained by the acquisition unit 52 by using a multi-case-type analysis model, where the multi-case-type analysis model is trained by case samples of multiple preset case types;

[0086] A marking unit 53, configured to obtain focus statements related to at least one disputed focus obtained by the analysis unit 53 from the conversation content.

[0087] Further, as Figure 6 shown, the apparatus further includes:

[0088] A generation unit 54, configured to generate corresponding disputed focus content according to the disputed focus obtained by the analysis unit 52 and the conversation content;

[0089] The marking unit 53 is further configured to mark the corresponding disputed focus and the disputed focus content obtained by the generation unit 55 on the focus statements.

[0090] Further, as Figure 6 shown, the apparatus further includes:

[0091] A modeling unit 55, configured to build a multi-case-type analysis model based on multiple preset case types, where the input of the multi-case-type analysis model is a case sample of a specified case type, the specified case type is one of the multiple preset case types, and the output of the multi-case-type analysis model is the disputed focus of the case sample and the corresponding disputed focus content;

[0092] A training unit 56, configured to mask other case types according to the specified case type of the case sample during the process of training the multi-case-type analysis model built by the modeling unit 55, so as to train the loss function of the multi-case-type analysis model under the specified case type.

[0093] Further, as Figure 6 shown, the apparatus further includes:

[0094] A construction unit 57, configured to build a case-type-based legal knowledge graph according to the multiple preset case types, where the legal knowledge graph includes legal entities and legal relationships, the legal entities are used to represent the disputed focuses of the case types, and the legal relationships are used to represent the association relationships between different disputed focuses;

[0095] A determination unit 58, configured to determine the vector representation of the disputed focuses of multiple preset case types according to the legal knowledge graph obtained by the construction unit 57.

[0096] Further, as Figure 6 shown, the acquisition unit 51 includes:

[0097] An acquisition module 511, configured to acquire the conversation content of the current court trial case, where the conversation content includes at least one conversation statement and a corresponding preset role identifier;

[0098] A generation module 512, configured to generate a statement vector based on each conversation statement obtained by the acquisition module 511 and the corresponding preset role identifier;

[0099] An encoding module 513, configured to encode the statement vectors corresponding to at least one conversation statement obtained by the generation module 512 by using a neural network model to obtain a conversation vector corresponding to the conversation content;

[0100] A mapping module 514, configured to perform semantic encoding on the conversation vector obtained by the encoding module 513 according to the context relationship of the conversation statements in the conversation content to obtain a conversation semantic vector corresponding to the conversation content.

[0101] Further, as Figure 6 shown, the analysis unit 52 includes:

[0102] An acquisition module 521, configured to determine multiple controversial focuses of the specified case type according to a legal knowledge graph;

[0103] An analysis module 522, configured to analyze the conversation semantic vector corresponding to the conversation content by using a multi-case type analysis model to determine the mapping probability of the conversation semantic vector on the multiple controversial focuses obtained by the acquisition module 521;

[0104] A determination module 523, configured to determine the controversial focuses included in the conversation content according to the mapping probability obtained by the analysis module 522.

[0105] Further, as Figure 6 shown, the analysis unit 52 further includes:

[0106] The acquisition module 521 is further configured to, after the determination module determines the controversial focuses included in the conversation content, determine a focus vector corresponding to each controversial focus according to the legal knowledge graph;

[0107] A fusion module 524, configured to fuse the conversation semantic vector with each focus vector determined by the acquisition module 521 to obtain at least one new conversation content vector;

[0108] A generation module 525, configured to process the new conversation content vector obtained by the fusion module 524 by using a text generation algorithm to generate corresponding controversial focus content for each controversial focus.

[0109] Further, the determination module 523 is specifically configured to:

[0110] Determine whether there is a mapping probability greater than the threshold among the mapping probabilities corresponding to each controversial focus;

[0111] If there is, determine the controversial focus corresponding to the mapping probability as the controversial focus contained in the conversation content;

[0112] If not, determine the controversial focus corresponding to the maximum mapping probability as the controversial focus contained in the conversation content.

[0113] Further, as Figure 6 shown, the annotation unit 53 includes:

[0114] A determination module 531, configured to determine a statement classification label according to the at least one controversial focus, where the statement classification label includes a focus statement and a non-focus statement corresponding to the controversial focus;

[0115] A classification module 532, configured to process the conversation semantic vector by using a neural network model, and determine, for each conversation statement in the conversation content, the statement classification label obtained by the determination module 531;

[0116] A processing module 533, configured to perform corresponding highlighting processing on the conversation statements whose statement classification labels are focus statements after being processed by the classification module 532 according to the controversial focus.

[0117] In addition, an embodiment of the present invention further provides a processor, where the processor is used to run a program, and when the program runs, it executes the method for focus analysis provided in any one of the above embodiments.

[0118] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0119] It can be understood that the relevant features in the above methods and devices can be referred to each other. In addition, the "first", "second", etc. in the above embodiments are used to distinguish each embodiment, and do not represent the advantages and disadvantages of each embodiment.

[0120] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.

[0121] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. A variety of general-purpose systems may also be used in conjunction with the teachings based hereon. The structure required to construct such systems will be apparent from the above description. Additionally, the present invention is not directed to any particular programming language. It should be appreciated that the teachings of the present invention described herein can be implemented in a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the preferred embodiments of the present invention.

[0122] In addition, the memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0123] Those skilled in the art will appreciate that the embodiments of the present application may be provided as a method, system, or computer program product. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0124] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the flows or multiple flows and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0125] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in Figure 1 one or more of the flows or multiple flows and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0127] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0128] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0129] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0130] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0131] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0132] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for focus analysis, the method comprising: Collecting the designated case type and conversation content of the current case, wherein the conversation content includes at least one conversation statement and the corresponding preset role identifier; Using a multi-case-type analysis model to analyze at least one controversial focus of the current case from the conversation content according to the designated case type, wherein the multi-case-type analysis model is trained by case samples of multiple preset case types, and the corresponding relationship between the multiple preset case types and the controversial focuses of the preset case type is determined based on a legal knowledge graph; Obtaining focus statements related to the at least one controversial focus in the conversation content; Wherein, the method further comprises: during the process of training the multi-case-type analysis model, masking other case types according to the designated case type of the case sample to train the loss function of the multi-case-type analysis model under the designated case type.

2. The method according to claim 1, wherein The method further comprises: Generating corresponding controversial focus content according to the controversial focus and the conversation content; Annotating the corresponding controversial focus and the controversial focus content on the focus statements.

3. The method according to claim 1, wherein Training the multi-case-type analysis model in the following manner: The input of the multi-case-type analysis model is a case sample of a designated case type, the designated case type is one of multiple preset case types, and the output of the multi-case-type analysis model is the controversial focus of the case sample and the corresponding controversial focus content.

4. The method according to claim 1, wherein The method further comprises: Constructing a case-type-based legal knowledge graph according to the multiple preset case types, the legal knowledge graph includes legal entities and legal relationships, the legal entities are used to represent the controversial focuses of the case type, and the legal relationships are used to represent the association relationships between different controversial focuses; Determining the vector representations of the controversial focuses of the multiple preset case types according to the legal knowledge graph.

5. The method according to any one of claims 1-4, characterized in that, Collecting the designated case type and conversation content of the current case, including: Obtaining the conversation content of the current case; Generating sentence vectors based on each conversation statement and the corresponding preset role identifier; Using a neural network model to encode the sentence vectors corresponding to the at least one conversation statement to obtain a conversation vector corresponding to the conversation content; Performing semantic encoding on the conversation vector according to the context relationship of the conversation statements in the conversation content to obtain a conversation semantic vector corresponding to the conversation content.

6. The method according to claim 5, characterized in that, Using a multi-case-type analysis model to analyze at least one controversial focus of the current case from the conversation content according to the designated case type, including: Determining multiple controversial focuses of the designated case type according to the legal knowledge graph; The multi-case-type analysis model analyzes the conversation semantic vector corresponding to the conversation content to determine the mapping probabilities of the conversation semantic vector on the multiple controversial focuses; Determining the controversial focuses included in the conversation content according to the mapping probabilities.

7. The method according to claim 6, wherein After determining the controversial focuses included in the conversation content, the method further comprises: Determining the focus vector corresponding to each controversial focus according to the legal knowledge graph; Fusing the conversation semantic vector with each focus vector respectively to obtain at least one new conversation content vector; Using a text generation algorithm to process the new conversation content vector to generate corresponding controversial focus content for each controversial focus.

8. The method according to claim 6, wherein Determine the controversial focuses contained in the conversation content according to the mapping probability, including: Judge whether there is a mapping probability greater than the threshold among the mapping probabilities corresponding to each controversial focus; If there is, determine the controversial focus corresponding to the mapping probability as the controversial focus contained in the conversation content; If not, determine the controversial focus corresponding to the maximum mapping probability as the controversial focus contained in the conversation content.

9. The method according to claim 5, characterized in that, Obtain focus sentences related to the at least one controversial focus in the conversation content, including: Determine sentence classification labels according to the at least one controversial focus, where the sentence classification labels include focus sentences and non-focus sentences corresponding to the controversial focus; Use a neural network model to process the conversation semantic vector to determine sentence classification labels for each conversation sentence in the conversation content; Highlight the conversation sentences with the sentence classification label of focus sentence according to the controversial focus.

10. A device for focus analysis, the device includes: A collection unit for collecting the specified case type of the current case and the conversation content, where the conversation content includes at least one conversation sentence and the corresponding preset role identifier; An analysis unit for using a multi-case-type analysis model to analyze at least one controversial focus of the current case from the conversation content according to the specified case type obtained by the collection unit, where the multi-case-type analysis model is trained by case samples of multiple preset case types, and the corresponding relationship between the multiple preset case types and the controversial focuses of the preset case type is determined based on a legal knowledge graph; A labeling unit for obtaining focus sentences related to the at least one controversial focus obtained by the analysis unit in the conversation content; Wherein, the device for focus analysis further includes a training unit for masking other case types according to the specified case type of the case sample during the training of the multi-case-type analysis model to train the loss function of the multi-case-type analysis model under the specified case type.

11. The device according to claim 10, characterized in that, The device further includes: A generation unit for generating corresponding controversial focus content according to the controversial focus obtained by the analysis unit and the conversation content; The labeling unit is further configured to label the corresponding controversial focus and the controversial focus content obtained by the generation unit on the focus sentences.

12. The device according to claim 10, characterized in that, The device further includes: A modeling unit for constructing a multi-case-type analysis model based on multiple preset case types, where the input of the multi-case-type analysis model is a case sample of a specified case type, the specified case type is one of the multiple preset case types, and the output of the multi-case-type analysis model is the controversial focus of the case sample and the corresponding controversial focus content.

13. The device according to claim 10, characterized in that, The device further includes: A construction unit for constructing a case-type-based legal knowledge graph according to the multiple preset case types, where the legal knowledge graph includes legal entities and legal relationships, the legal entities are used to represent the controversial focuses of the case type, and the legal relationships are used to represent the association relationships between different controversial focuses; A determination unit for determining the vector representation of the controversial focuses of the multiple preset case types according to the legal knowledge graph obtained by the construction unit.

14. A processor, characterized in that, The processor is used to run a program, wherein, when the program runs, it executes the method for focus analysis described in any one of claims 1-10.

Citation Information

Patent Citations

  • Method and system for processing court trial records

    CN110895568A

  • Platform for managing social media content throughout an organization

    WO2019018689A1