Method and system for digitally displaying law cases
Through automated digital display methods and systems of legal cases, the problem of low efficiency in displaying existing legal case data is solved, automated processing and accurate display are realized, and work efficiency and display accuracy are improved.
Patent Information
- Application Number
- CN202510089999.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The display of existing legal case data still requires manual analysis and sorting, which is inefficient and requires analysts to have strong legal professional knowledge.
Provide a method and system for displaying legal cases digitally, and automatically extract, image description generation, semantic recognition and template matching by obtaining text data and image data, realizing automatic digital display of legal cases.
It improves the efficiency of legal case data processing, reduces the need for manual analysis, ensures the accuracy of the presentation content, and can customize digital display templates for different types of case displays.
Smart Images

Figure CN120012745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of law teaching technology, and more specifically to a method and system for digitally displaying law cases. Background Art
[0002] Jurisprudence, also known as legal studies or legal science, is a science that studies law, legal phenomena, and their regularities. It is a specialized discipline that studies issues related to law, and it is also a knowledge and theoretical system about legal issues. Its core lies in the study of order and justice, and it is the study of order and justice. In the classroom teaching of law, it is often necessary to explain the legal provisions in combination with actual legal cases. Therefore, case teaching is one of the important components of legal teaching. The people, things, and objects involved in legal case teaching have their own unique professional characteristics. When teaching and presenting, it is necessary to analyze and organize the data of legal cases for the convenience of teaching. However, the analysis and organization of the existing legal case data display is still done manually, which requires the analyst to have strong legal professional knowledge, and the data processing efficiency is low. Therefore, how to provide a digital display method and system for legal cases is an urgent problem to be solved by technical personnel in this field. Summary of the invention
[0003] In view of this, the present invention provides a method and system for digitally displaying legal cases, which automatically organize and display legal case materials and improve data processing efficiency.
[0004] In order to achieve the above object, the present invention provides the following technical solutions:
[0005] A method for digitally displaying legal cases, comprising the following steps:
[0006] S1. Obtain text data and image data of legal cases;
[0007] S2, dividing the paragraphs in the text data of the legal cases into long paragraphs and short paragraphs based on the number of sentences contained, extracting summaries from the long paragraphs to obtain summary texts;
[0008] S3, generating image description text based on the image data of legal cases;
[0009] S4. Perform semantic recognition on summary texts and short paragraphs to extract legal-related text keywords and their attribute tags; perform semantic recognition on image description texts to extract image keywords and their attribute tags;
[0010] S5, constructing a digital display template of the legal case, and matching the image keywords of the image data and the text keywords of the text data with the digital display template;
[0011] S6. Based on the matching result of the digital display template, fill the corresponding information into the template to complete the digital display of the legal case.
[0012] Optionally, S2 is specifically:
[0013] S21, dividing the paragraphs in the text data of legal cases into long paragraphs and short paragraphs based on the number of sentences contained;
[0014] S22, extracting feature vectors of sentences in the long paragraph text. If the text data of the legal case contains a title, extracting feature vectors of the title at the same time and calculating the similarity between all sentences in the long paragraph text and the title;
[0015] S23, respectively calculating the domain word frequency weights and special word weights of all sentences in the long paragraph text, and if the text data of the legal case contains a title, also calculating the title similarity weight, and obtaining the comprehensive weight of the sentence based on the above weights;
[0016] S24. Sort all the sentences based on their comprehensive weights, select and filter the top-ranked sentences to obtain summary texts.
[0017] Optional, S3 is:
[0018] S31, preprocessing the image data of the legal case, and then extracting the image feature vector;
[0019] S32, extracting a positive feature state and a reverse feature state of the image based on a feature vector of the image;
[0020] S33, jointly predicting a description label sequence of the image based on the forward feature state and the reverse feature state of the image to obtain an image description text.
[0021] Optionally, S4 is specifically:
[0022] S41, converting summary text, short paragraphs and image description text into feature vectors;
[0023] S42, extracting local features and context information of the feature vector, performing keyword label prediction on the extracted features, and obtaining entity attributes of the keyword;
[0024] S43. Determine whether there is a relationship between the keywords based on the predefined attribute relationship of the keywords. If so, take the keywords and the keyword attribute relationship as an information group, and output the information group and the individual keywords as the result of semantic recognition.
[0025] Optionally, the digital display template constructed in S5 includes:
[0026] Set a text matching word and a corresponding text display scheme for the text keyword. When the text keyword matches a word in the text matching word, display the text data using the text display scheme corresponding to the text matching word.
[0027] A plurality of image matching phrases and corresponding image display schemes are set, and when an image keyword matches a word in the image matching phrase, the image keyword is displayed through the image display scheme corresponding to the image matching phrase.
[0028] Optionally, S5 matches the keyword with the digital display template as follows:
[0029] The semantic similarity between the attribute tag of the keyword and the matching word set in the digital display template is calculated. If the similarity is greater than a preset similarity threshold, the match is confirmed.
[0030] A digital display system for legal cases, which executes the digital display method for legal cases, comprises:
[0031] Abstract extraction module, used to extract abstract text of long paragraphs in the text data of legal cases;
[0032] Description generation module, used to generate image description text of legal cases;
[0033] The semantic recognition module performs semantic recognition on the text and image data of legal cases and extracts legal-related keywords;
[0034] Template building module, used to build digital display templates of legal cases;
[0035] The information matching module is used to match keywords and digital display templates to complete information filling.
[0036] It can be seen from the above technical solution that compared with the prior art, the present invention provides a method and system for digital display of legal cases, which has the following beneficial effects: the present invention can automatically complete the digital display of legal cases based on the text data and image data in the legal case materials, thereby improving work efficiency, and ensuring the accuracy of the displayed content through precise identification and matching; the present invention can customize the matching rules of the digital display template, which can be applied to different types of case displays. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0038] Figure 1 This is a flow chart of the digital display method of legal cases of the present invention. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] The embodiment of the present invention discloses a method for digitally displaying legal cases, such as Figure 1 As shown, the following steps are included:
[0041] S1. Obtain text data and image data of legal cases;
[0042] S2, dividing the paragraphs in the text data of the legal cases into long paragraphs and short paragraphs based on the number of sentences contained, extracting summaries from the long paragraphs to obtain summary texts;
[0043] S3, generating image description text based on the image data of legal cases;
[0044] S4. Perform semantic recognition on summary texts and short paragraphs to extract legal-related text keywords and their attribute tags; perform semantic recognition on image description texts to extract image keywords and their attribute tags;
[0045] S5, constructing a digital display template of the legal case, and matching the image keywords of the image data and the text keywords of the text data with the digital display template;
[0046] S6. Based on the matching result of the digital display template, fill the corresponding information into the template to complete the digital display of the legal case.
[0047] Furthermore, S2 is specifically:
[0048] S21, dividing the paragraphs in the text data of legal cases into long paragraphs and short paragraphs based on the number of sentences contained;
[0049] S22, extracting feature vectors of sentences in the long paragraph text. If the text data of the legal case contains a title, extracting feature vectors of the title at the same time and calculating the similarity between all sentences in the long paragraph text and the title;
[0050] S23, respectively calculating the domain word frequency weights and special word weights of all sentences in the long paragraph text, and if the text data of the legal case contains a title, also calculating the title similarity weight, and obtaining the comprehensive weight of the sentence based on the above weights;
[0051] S24. Sort all the sentences based on their comprehensive weights, select and filter the top-ranked sentences to obtain summary texts.
[0052] The text data of legal cases may include case introduction, case focus, related laws, plaintiff information, defendant information, court information, trial results, etc. Different contents have different lengths. When semantic extraction is performed on long paragraphs such as case introduction and trial results, its summary text needs to be extracted first.
[0053] In the embodiment of the present invention, the cosine similarity between the word vectors extracted by Word2Vec is first calculated, and then the average value of the similarities of all words in the sentence and the title is used as the similarity between all sentences and the title in the long paragraph text;
[0054] The calculation of each weight is as follows:
[0055] Domain word frequency weight W1:
[0056]
[0057] In the formula, c1 is the number of keywords contained in the sentence, and c2 is the number of all words in the sentence;
[0058] Special word weight: If the sentence contains a preset special word, and the special word is set to a professional term related to law, then its special word weight W2 is set to 1;
[0059] Title similarity weight W3:
[0060]
[0061] In the formula, s represents the similarity between the sentence and the title;
[0062] Comprehensive weight W:
[0063] W=α1×W1+α2×W2+α3×W3
[0064] Where α1, α2, and α3 are the influence coefficients corresponding to weights W1, W2, and W3, respectively.
[0065] In S24, candidate sentences are obtained by selecting the sentences with the highest ranking. The similarities between all candidate sentences are calculated in the same way. If the similarity between two sentences exceeds a set value, the sentences with the lowest ranking are deleted. After the screening is completed, a summary text is obtained.
[0066] Furthermore, S3 is specifically:
[0067] S31, preprocessing the image data of the legal case, and then extracting the image feature vector;
[0068] S32, extracting a positive feature state and a reverse feature state of the image based on a feature vector of the image;
[0069] S33, jointly predicting a description label sequence of the image based on the forward feature state and the reverse feature state of the image to obtain an image description text.
[0070] In an embodiment of the present invention, after preprocessing the image data of the legal case, the image feature vector is extracted by a convolutional neural network, and then the positive feature state and the reverse feature state of the image are respectively extracted by a bidirectional long short-term memory network, and the bidirectional long short-term memory network includes two LSTM networks;
[0071] The LSTM network consists of three gate structures, namely the forget gate, input gate and output gate;
[0072] The forget gate is used to determine how much of each state to retain:
[0073] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0074] The input gate is used to control the input at the current moment and update the unit state:
[0075] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0076]
[0077] Determine the update status in the candidate value based on the update value:
[0078]
[0079] The output gate is used to generate output values:
[0080] o t =σ(W o ·[h t-1 ,x t ]+b o )
[0081] Get the hidden state based on the output value and update state:
[0082] h t =o t *tanh(C t )
[0083] Where W f , W i , W C , W o are the weight matrices of the forget gate, input gate, candidate value, and output gate, respectively, and b f 、b i 、b C 、b o are the bias vectors of the forget gate, input gate, candidate value, and output gate, respectively. σ(·) represents the Sigmoid activation function. t-1 is the hidden state at the previous moment, h t is the hidden state at the current moment, x t is the current input, is the candidate value, i t is the updated value, o t is the output value, c t-1 is the state at the previous moment, f t is the degree of retention, c t To update the status.
[0084] Bidirectional long short-term memory network gets positive feature state and reverse feature state Then the positive feature state of the combined image and reverse feature state The probability of the description label sequence is calculated through the Softmax function, and finally the text description is obtained through the pre-constructed dictionary library.
[0085] Furthermore, S4 is specifically:
[0086] S41, converting summary text, short paragraphs and image description text into feature vectors;
[0087] S42, extracting local features and context information of the feature vector, performing keyword label prediction on the extracted features, and obtaining entity attributes of the keyword;
[0088] S43. Determine whether there is a relationship between the keywords based on the predefined attribute relationship of the keywords. If so, take the keywords and the keyword attribute relationship as an information group, and output the information group and the individual keywords as the result of semantic recognition.
[0089] In an embodiment of the present invention, Word2Vec is used to convert summary text, short paragraphs and image description text into feature vectors, and then local features are extracted through a convolutional neural network, context information is extracted through an LSTM network, and finally the probability of label output is calculated based on a CRF network.
[0090] Furthermore, the digital display template constructed in S5 includes:
[0091] Set a text matching word and a corresponding text display scheme for the text keyword. When the text keyword matches a word in the text matching word, display the text data using the text display scheme corresponding to the text matching word.
[0092] A plurality of image matching phrases and corresponding image display schemes are set, and when an image keyword matches a word in the image matching phrase, the image keyword is displayed through the image display scheme corresponding to the image matching phrase.
[0093] In an embodiment of the present invention, when setting a display plan, matching can be performed on information groups and individual keywords respectively. If the display plan requires information group matching, the keywords in the information group and the keyword attribute relationship need to be matched separately during matching; if the display plan does not require a complete match of the information group, it can be matched with individual keywords or keywords in the information group separately.
[0094] In the embodiment of the present invention, the display scheme for long paragraphs can be set to display the original text or summary text after information matching; the text display scheme can also be set to include keyword frequency statistical analysis and other content.
[0095] Furthermore, S5 matches the keywords with the digital display template as follows:
[0096] The semantic similarity between the attribute tag of the keyword and the matching word set in the digital display template is calculated. If the similarity is greater than a preset similarity threshold, the match is confirmed.
[0097] In the embodiment of the present invention, the semantic similarity is calculated using cosine similarity:
[0098]
[0099] In the formula, S AB Represents the similarity between word vector A and word vector B.
[0100] and Figure 1 Corresponding to the method, the embodiment of the present invention further discloses a digital display system for legal cases, which executes the digital display method for legal cases, including:
[0101] Abstract extraction module, used to extract abstract text of long paragraphs in the text data of legal cases;
[0102] Description generation module, used to generate image description text of legal cases;
[0103] The semantic recognition module performs semantic recognition on the text and image data of legal cases and extracts legal-related keywords;
[0104] Template building module, used to build digital display templates of legal cases;
[0105] The information matching module is used to match keywords and digital display templates to complete information filling.
[0106] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0107] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for digital display of legal cases, characterized in that: The following steps are involved: S1. Obtain text data and image data of legal cases; S2, dividing the paragraphs in the text data of the legal cases into long paragraphs and short paragraphs based on the number of sentences contained, extracting summaries from the long paragraphs to obtain summary texts; S3, generating image description text based on the image data of legal cases; S4. Perform semantic recognition on summary texts and short paragraphs to extract legal-related text keywords and their attribute tags; perform semantic recognition on image description texts to extract image keywords and their attribute tags; S5, constructing a digital display template of the legal case, and matching the image keywords of the image data and the text keywords of the text data with the digital display template; S6. Based on the matching result of the digital display template, fill the corresponding information into the template to complete the digital display of the legal case.
2. A method for digital display of legal cases according to claim 1, characterized in that: S2 is specifically: S21, dividing the paragraphs in the text data of legal cases into long paragraphs and short paragraphs based on the number of sentences contained; S22, extracting feature vectors of sentences in the long paragraph text. If the text data of the legal case contains a title, extracting feature vectors of the title and calculating similarities between all sentences in the long paragraph text and the title; S23, respectively calculating the domain word frequency weights and special word weights of all sentences in the long paragraph text, and if the text data of the legal case contains a title, also calculating the title similarity weight, and obtaining the comprehensive weight of the sentence based on the above weights; S24. Sort all the sentences based on their comprehensive weights, select and filter the top-ranked sentences to obtain summary texts.
3. A method for digital display of legal cases according to claim 1, characterized in that: S3 is specifically: S31, preprocessing the image data of the legal case, and then extracting the image feature vector; S32, extracting a positive feature state and a reverse feature state of the image based on a feature vector of the image; S33, jointly predicting a description label sequence of the image based on the forward feature state and the reverse feature state of the image to obtain an image description text.
4. A method for digital display of legal cases according to claim 1, characterized in that: S4 is specifically: S41, converting summary text, short paragraphs and image description text into feature vectors; S42, extracting local features and context information of the feature vector, performing keyword label prediction on the extracted features, and obtaining entity attributes of the keyword; S43. Determine whether there is a correlation between the keywords based on the predefined attribute relationship of the keywords. If so, take the keywords and the keyword attribute relationship as an information group, and output the information group and the individual keywords as the result of semantic recognition.
5. A method for digital display of legal cases according to claim 1, characterized in that: The digital display templates constructed in S5 include: Set a text matching word and a corresponding text display scheme for the text keyword. When the text keyword matches a word in the text matching word, display the text data using the text display scheme corresponding to the text matching word. A plurality of image matching phrases and corresponding image display schemes are set, and when an image keyword matches a word in the image matching phrase, the image keyword is displayed through the image display scheme corresponding to the image matching phrase.
6. A method for digital display of legal cases according to claim 5, characterized in that: In S5, the keywords are matched with the digital display template as follows: The semantic similarity between the attribute tag of the keyword and the matching word set in the digital display template is calculated. If the similarity is greater than a preset similarity threshold, the match is confirmed.
7. A digital display system for legal cases, characterized in that: The method for digitally displaying legal cases according to any one of claims 1 to 6 comprises: Abstract extraction module, used to extract abstract text of long paragraphs in the text data of legal cases; Description generation module, used to generate image description text of legal cases; The semantic recognition module performs semantic recognition on the text and image data of legal cases and extracts legal-related keywords; Template building module, used to build digital display templates of legal cases; The information matching module is used to match keywords and digital display templates to complete information filling.