Big data-based judicial document analysis method and system

Through the judicial document analysis method based on big data, the word vector data and large language model are used to solve the problem of poor extraction of keywords in judicial document, and more efficient and accurate judicial document quality analysis is achieved.

CN120146029AActive Publication Date: 2025-06-13GUANGDONG BOWEI CHUANGYUAN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510592684.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-13
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The prior art has poor results in the extraction of keywords of judicial instruments, and it is difficult to adapt to the diversity and complexity of judicial instruments content. The method based on large language models requires powerful computing resources, which may lead to resource bottlenecks and inefficient analysis.

Method used

The judicial document analysis method based on big data is adopted, and by inputting word vector data into the preset classification model, word thermal data is obtained, central words and central statements are located, and keywords are extracted and quality analysis reports are generated using the large language model.

Benefits of technology

It improves the accuracy and efficiency of keyword extraction in judicial documents, reduces the impact of irrelevant text information, reduces the resource occupation of large language models in the calculation and reasoning process, and improves the accuracy and efficiency of quality analysis reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146029A_ABST
    Figure CN120146029A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text processing, in particular to a judicial document analysis method and system based on big data, starting from word vector data, word thermodynamic data is output by using a preset classification model, and correlation between segmented words and word categories is quantitatively displayed in a thermodynamic value form. Therefore, the segmented words related to the word types in the target judicial document are accurately positioned, then the head word, the center statement and the keyword are determined in a layer-by-layer progressive mode, a rich and comprehensive data basis is provided for subsequent quality analysis, the accuracy of a quality analysis report is improved, and the quality analysis efficiency is improved. And irrelevant text information in the target judicial document can be effectively filtered out in the keyword extraction process, so that subsequent data extraction and analysis are focused on core contents of the target judicial document, resources occupied by a large language model in the calculation and reasoning process are reduced, and the efficiency of the quality analysis process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and in particular to a judicial document analysis method and system based on big data. Background Art

[0002] In the judicial field, the quality of judicial documents directly affects judicial credibility and case handling efficiency. Therefore, analyzing the quality of judicial documents helps to correct errors in judicial documents in a timely manner, facilitates judicial personnel to quickly and accurately understand the content of the case, and avoids repeated communication and misunderstandings caused by unclear expressions and confusing logic in the documents. It plays an important role in maintaining the authority of the judiciary and improving the efficiency of judicial management.

[0003] The quality analysis of judicial documents relies on keyword extraction. Keyword extraction methods in the prior art include rule-based keyword extraction methods, which mainly rely on pre-defined rules and templates as the basis for keyword extraction, and are difficult to adapt to the diversity and complexity of the content of judicial documents, resulting in poor keyword extraction effects; and statistical-based keyword extraction methods, which determine keywords by calculating the frequency of occurrence of words in documents and their rarity in the entire corpus, and only consider the statistical characteristics of words, ignoring the semantic relationship and contextual information between words, and are prone to extracting words that are irrelevant to the core content of the case but have a high word frequency, or omitting important but less frequent keywords, resulting in poor keyword extraction effects, and thus resulting in low accuracy of the quality analysis results of judicial documents.

[0004] In order to improve the effect of keyword extraction in judicial documents, the existing technology adopts a keyword extraction method based on traditional machine learning, and uses machine learning models such as naive Bayes and support vector machines to extract keywords. However, this method relies on a large amount of labeled data for training, and professional legal knowledge is required in the annotation of judicial documents, resulting in high cost and low efficiency of the above method; and keyword extraction of judicial documents based on large language models requires word-by-word analysis and processing of the entire input judicial document, which involves complex calculation and reasoning processes, and running a large language model requires powerful computing resource support. When extracting keywords from a large number of judicial documents, resource bottlenecks may occur, causing the model to run slowly or even fail to run normally, affecting the accuracy and efficiency of the quality analysis of judicial documents.

[0005] Therefore, how to improve the accuracy and efficiency of quality analysis of judicial documents has become an urgent problem to be solved. Summary of the invention

[0006] In view of the above technical problems, the technical solution adopted by the present invention is a judicial document analysis method based on big data, which includes the following steps: S1. Input the word vector data corresponding to each judicial document segment of the target judicial document into a preset classification model to obtain the word heat data of each judicial document segment corresponding to each channel of the preset classification model. Each channel of the preset classification model corresponds to a word category, and each word heat data is used to represent the correlation between each word segmentation in the corresponding judicial document segment and the corresponding word category through a heat value.

[0007] S2. For any word heat data, determine the several segmentation words corresponding to several heat centers in the judicial document segment corresponding to the current word heat data as the several central words corresponding to the judicial document segment corresponding to the current word heat data. The heat center refers to the position of the segmentation word corresponding to the heat value greater than the preset heat value threshold in the word heat data in the corresponding judicial document segment.

[0008] S3. Extract the central sentence corresponding to each central word from the judicial document segment corresponding to the current word heat data.

[0009] S4. Input each central sentence corresponding to the current word heat data and the word category corresponding to the current word heat data into a preset large language model to obtain several keywords in each central sentence corresponding to the current word heat data and the word category corresponding to each keyword.

[0010] S5. Traverse all word heat data to obtain all keywords corresponding to the target judicial document and the word category corresponding to each keyword.

[0011] S6. Input all keywords corresponding to the target judicial document and the word category corresponding to each keyword into a preset large language model to obtain the quality analysis report corresponding to the target judicial document.

[0012] The present invention also provides a judicial document analysis system based on big data. The judicial document analysis system based on big data includes: A word classification module, which is used to input the word vector data corresponding to each judicial document segment of the target judicial document into a preset classification model to obtain the word heat data of each judicial document segment corresponding to each channel of the preset classification model. Each channel of the preset classification model corresponds to a word category, and each word heat data is used to represent the correlation between each word segmentation in the corresponding judicial document segment and the corresponding word category through a heat value.

[0013] The central word screening module is used to, for any word heat data, determine, in the judicial document segment corresponding to the current word heat data, several word segmentations corresponding to several heat centers in the current word heat data as several central words corresponding to the judicial document segment corresponding to the current word heat data, where a heat center refers to the position of the word segmentation corresponding to the heat value greater than the preset heat value threshold in the word heat data in the corresponding judicial document segment.

[0014] The central sentence extraction module is used to extract, from the judicial document segment corresponding to the current word heat data, the central sentence corresponding to each central word.

[0015] The first keyword extraction module is used to input each central sentence corresponding to the current word heat data and the word category corresponding to the current word heat data into a preset large language model to obtain several keywords in each central sentence corresponding to the current word heat data and the word category corresponding to each keyword.

[0016] The second keyword extraction module is used to traverse all the word heat data to obtain all the keywords corresponding to the target judicial document and the word category corresponding to each keyword.

[0017] The quality analysis module is used to input all the keywords corresponding to the target judicial document and the word category corresponding to each keyword into a preset large language model to obtain the quality analysis report corresponding to the target judicial document.

[0018] The present invention has at least the following beneficial effects: By starting from the word vector data, using a preset classification model to output word heat data, and quantitatively displaying the correlation between the word segmentation and the word category in the form of a heat value, the word segmentations related to each word category in the target judicial document can be accurately located. Furthermore, through a step-by-step manner, central words, central sentences, and keywords are determined, providing a rich and comprehensive data basis for subsequent quality analysis, improving the accuracy of the quality analysis report, and being able to effectively filter out the irrelevant text information in the target judicial document during the keyword extraction process, enabling subsequent data extraction and analysis to focus on the core content of the target judicial document, reducing the resources occupied by the large language model during the calculation and reasoning process, and improving the efficiency of the quality analysis process. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1The flowchart of a big data-based judicial document analysis method provided by Embodiment 1 of the present invention; Figure 2 The schematic diagram of a big data-based judicial document analysis system provided by Embodiment 2 of the present invention. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It can be understood that, under appropriate circumstances, the above terms used to distinguish similar objects can be interchanged so that the present invention can also implement other embodiments other than the above-described illustrated embodiments or described embodiments. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0023] Embodiment 1 Embodiment 1 of the present invention provides a big data-based judicial document analysis method. The big data-based judicial document analysis method includes the following steps, as Figure 1 shown: S1. Input the word vector data corresponding to each judicial document segment of the target judicial document into a preset classification model, and obtain the word heat data of each judicial document segment corresponding to each channel of the preset classification model. Among them, each channel of the preset classification model corresponds to a word category, and each word heat data is used to represent the correlation between each word segmentation in the corresponding judicial document segment and the corresponding word category through a heat value.

[0024] Among them, the target judicial document refers to a specific judicial document that needs to be analyzed for quality and analysis, such as various legal documents with legal effect or related to judicial procedures, such as judgments, rulings, mediation documents, indictments, etc.

[0025] The judicial document segment is a smaller text unit obtained by splitting the target judicial document according to certain rules. Each judicial document segment relatively independently contains a part of semantic information, which is convenient for subsequent model processing and analysis.

[0026] Word vector data refers to data that converts words in judicial documents into numerical vector forms. In the vector space, the dimensions and values of word vectors reflect the semantic features and context relationships of words, providing a feature basis for word classification.

[0027] A preset classification model is a machine learning or deep learning model that has been pre-trained. The internal structure and parameters of the preset classification model have been optimized through training on a large amount of relevant data. It is used to classify the input word vector data into different word categories and can identify and judge the correlation between words and different categories.

[0028] A channel refers to a logical unit in the preset classification model, and each channel corresponds to a specific word category. Specifically, when the preset classification model processes the input word vector data, it will perform calculations and analyses for each channel separately and output the correlation results of the corresponding word categories.

[0029] Word heat data refers to the result data output by the preset classification model, which presents the degree of correlation between each word segment and the corresponding word category in the form of heat values. The numerical values of the heat data intuitively characterize the importance distribution of different word segments under different word categories.

[0030] Word categories refer to the pre-defined word set classifications with specific semantics by the implementer according to the content and analysis requirements of judicial documents, and are used for semantic classification and analysis of word segments in judicial documents. For example, the "party information" category includes the plaintiff, defendant, third party, legal representative, entrusted agent, witness, expert witness, etc., which are used to clarify the identities of various personnel participating in the judicial process. The "judicial procedure" category refers to words related to the judicial litigation procedure, such as filing a case, acceptance, hearing, trial, judgment, ruling, appeal, protest, execution, etc., which are used to reflect different stages and operations of a case in the judicial process. The "legal basis" category involves specific names of laws and regulations, articles, and related legal concepts, such as the "Tort Liability Law", "Principle of Fault Liability", "Contract Breach Clause", etc. The "legal relationship" category is used to define the nature of the legal relationship between the parties, such as contract relationship, tort relationship, marriage relationship, inheritance relationship, etc.

[0031] The heat value is a quantitative indicator in the word heat data, which is used to measure the strength of the correlation between a word segment and the corresponding word category. The heat value usually takes values between 0 and 1. The closer the value is to 1, the stronger the correlation between the word segment and the word category. The closer the value is to 0, the weaker the correlation between the word segment and the word category.

[0032] As described above, through the word vector technology, the word segmentation in the judicial document fragments is mapped into a vector space, so that the semantic relationship between word segmentations can be quantitatively represented by means of the distance and operation of vectors, which is convenient for the preset classification model to understand and process the semantic information of words. Through the feature patterns and mapping relationships between different word categories and word vectors learned by the preset classification model during the training stage, the possibility that each word segmentation in the judicial document fragment belongs to different word categories is judged, and the correlation result is output in the form of a heat value, realizing the probabilistic judgment and quantitative display of the semantic categories of word segmentations, and providing a data basis for subsequent keyword extraction and quality analysis.

[0033] In a specific embodiment, S1 includes the following steps: S11, perform word segmentation on each initial statement in the target judicial document to obtain a word segmentation set corresponding to each initial statement, where the word segmentations in each word segmentation set are arranged in the order of their positions before and after in the corresponding initial statement.

[0034] S12, according to the target input length corresponding to the preset classification model and the preset length of each word vector, obtain the number of word vector inputs corresponding to the preset classification model.

[0035] S13, according to the number of word vector inputs and the number of word segmentations in the word segmentation set corresponding to each initial statement, group all the initial statements to obtain several initial statement combinations and a word segmentation set combination corresponding to each initial statement combination, where the total number of word segmentations corresponding to each initial statement combination is less than or equal to the number of word vector inputs.

[0036] S14, determine each initial statement combination as a judicial document fragment.

[0037] S15, perform word vector conversion on each word segmentation in the word segmentation set combination corresponding to each judicial document fragment to obtain the word vector data corresponding to each judicial document fragment.

[0038] Wherein, the initial statement refers to a complete sentence separated by punctuation marks such as a period, a question mark, an exclamation mark, etc. in the target judicial document.

[0039] With the help of a word segmentation tool, each initial statement in the target judicial document is split into individual word segmentations, and then a word segmentation set corresponding to each initial statement is obtained. Those skilled in the art know that any word segmentation tool in the prior art falls within the protection scope of the present invention, and will not be elaborated here.

[0040] Calculate the number of word vector inputs that the preset classification model can accept according to the target input length required by the preset classification model and the preset length of each word vector set when tokenizing according to the preset tokenization tool. Specifically, the number of word vector inputs is equal to the ratio of the target input length to the preset length of each word vector.

[0041] The initial statements in each initial statement combination are arranged in the order of their positions before and after in the target legal document. The grouping principle of the initial statements is that the total number of tokens corresponding to several initial statements in each initial statement combination does not exceed the number of word vector inputs, and in the order of the positions before and after of each initial statement in the target legal document, the total number of tokens corresponding to several initial statements in each initial statement combination plus the number of tokens corresponding to the next initial statement exceeds the number of word vector inputs, so as to obtain several initial statement combinations and the corresponding token set combinations for each combination.

[0042] Convert each token in the token set combination corresponding to each legal document fragment into a word vector through a pre-trained word vector model, and finally obtain the word vector data corresponding to each legal document fragment. Correspondingly, the word vectors of each word vector data are arranged in the order of the positions before and after of the corresponding tokens in the corresponding initial statements. Those skilled in the art know that any word vector model in the prior art falls within the protection scope of the present invention, and will not be elaborated here.

[0043] As described above, through the tokenization operation, the initial statements are segmented into meaningful word units, which provides basic word units for subsequent word vector conversion, ensures that the order information of the words is retained, is conducive to the classification model to learn the semantics and grammatical structures of the statements, and according to the target input length corresponding to the preset classification model and the preset length of each word vector, obtains the number of word vector inputs corresponding to the preset classification model, ensures that the data input into the preset classification model meets the input requirements of the model, avoids the preset classification model from malfunctioning due to mismatched input lengths, and guarantees the normal operation of the preset classification model.

[0044] In a specific embodiment, the preset classification model is trained through the following steps: S10, obtain a plurality of legal document fragment samples, the corresponding word vector data samples for each legal document fragment sample, and the corresponding reference heat data samples for each legal document fragment sample for each word category.

[0045] S20, input the word vector data samples corresponding to each legal document fragment sample into the initial classification model, and obtain the word heat data samples corresponding to each channel of the initial classification model for each legal document fragment sample.

[0046] S30. Obtain the total model loss corresponding to the initial classification model based on the word heat data samples of each channel of the initial classification model corresponding to each judicial document fragment sample and the reference heat data samples corresponding to each word category for each judicial document fragment sample.

[0047] S40. Update the parameters of the initial classification model according to the total model loss until the total model loss converges, and obtain the trained preset classification model.

[0048] Among them, the judicial document fragment sample refers to the text fragment selected from various judicial documents for model training. In this implementation, a large number of judicial document fragment samples are obtained as the training data of the initial classification model. The judicial document fragment samples cover various types of judicial document contents, such as judgments, indictments of different case types, etc., to ensure that the preset classification model can learn rich semantic features of judicial documents and the correlation patterns of word categories, thereby improving the generalization ability and accuracy of the preset classification model.

[0049] The word vector data sample refers to the data obtained by converting the word segmentation in the judicial document fragment sample into a vector form, which reflects the semantic information and context relationship of the words. The acquisition method of the word vector data sample can refer to the acquisition method of the word vector data corresponding to each judicial document fragment of the target judicial document.

[0050] The reference heat data sample refers to the accurate representation of the correlation between each word segmentation and each word category determined in advance for each judicial document fragment sample. It can be manually labeled and presented in the form of heat values as the target output for the training of the initial classification model, so as to calculate the difference with the actual output to characterize the output accuracy of the initial classification model.

[0051] Based on its own network structure and initial parameters, the initial classification model performs feature extraction and classification prediction on the input word vector data sample, and obtains the word heat data samples of each channel of the initial classification model corresponding to each judicial document fragment sample, which is used to characterize the correlation degree between each word segmentation in each judicial document fragment sample and different word categories.

[0052] Further, the word heat data sample (i.e., the model prediction result) of each channel of the initial classification model corresponding to each judicial document fragment sample is compared with the reference heat data sample (i.e., the true result) corresponding to each judicial document fragment sample for each word category. The difference between the corresponding word heat data sample and the reference heat data sample is calculated through the loss function to measure the gap between the model prediction result and the true result. Thus, the differences of all judicial document fragment samples are aggregated to obtain the total model loss corresponding to the initial classification model, providing a clear direction and goal for model parameter update, that is, adjusting the parameters in the direction of reducing the total model loss. When the total model loss converges, it indicates that the classification model has learned the effective relationship between the input data and the output data, and a trained preset classification model is obtained for more accurate analysis and processing of new judicial document fragments.

[0053] As described above, through training with a large number of sample data and continuous optimization of parameters, the preset classification model can deeply learn the complex relationship between word segmentation and each word category in judicial documents, enabling the preset classification model to be more accurate in semantic analysis and relevance judgment of judicial document fragments in practical applications, thereby improving the accuracy of word heat data.

[0054] In a specific embodiment, S10 includes the following steps: S101, obtain the word segmentation set combination sample corresponding to each judicial document fragment sample, several central word samples corresponding to each word segmentation set combination sample, and the word category corresponding to each central word sample.

[0055] S102, for any central word sample in any word segmentation set combination sample, determine the first function parameter corresponding to the current central word sample according to the sequence number of the current central word sample in the current word segmentation set combination sample, where the first function parameter is used to determine the position of the corresponding distribution function sample on the number axis.

[0056] S103, according to the preset second function parameter and the first function parameter corresponding to the current central word sample, obtain the distribution function sample corresponding to the current central word sample, where the distribution function sample takes the sequence number corresponding to the word segmentation in the current word segmentation set combination sample as the independent variable and the correlation between the word segmentation in the current word segmentation set combination sample and the word category corresponding to the current central word sample as the dependent variable, and the preset second function parameter is used to determine the width and height of the corresponding distribution function sample.

[0057] S104, according to the distribution function sample corresponding to the current central word sample, obtain the heat sub-value sample of each word segmentation in the current word segmentation set combination sample for the word category corresponding to the current central word sample.

[0058] S105. Traverse all the central word samples in the current judicial document fragment sample, and obtain several heat sub-value samples for each word segmentation in the current word segmentation set combination sample for each word category.

[0059] S106. According to the several heat sub-value samples for each word segmentation in the current word segmentation set combination sample for each word category, obtain the heat total value sample for each word segmentation in the current word segmentation set combination sample for each word category.

[0060] S107. For any word category, according to the heat total value sample of each word segmentation in the current word segmentation set combination sample for the current word category, obtain the reference heat data sample of the corresponding judicial document fragment of the current word segmentation set combination sample for the current word category.

[0061] S108. Traverse all the word segmentation set combination samples and all the word categories, and obtain the reference heat data sample of each judicial document fragment for each word category.

[0062] Among them, the acquisition method of the word segmentation set combination sample corresponding to each judicial document fragment sample can refer to the acquisition method of the word segmentation set combination corresponding to each initial statement combination in the target judicial document.

[0063] The central word sample refers to the word with key semantics in the word segmentation set combination sample, representing a specific semantic category, serving as the core for subsequent analysis of the relevance between word segmentation and word category, and playing a key role in the entire text analysis process.

[0064] The sequential number refers to the number assigned to each word segmentation in the word segmentation set combination sample according to its front-back order in the word segmentation set combination sample.

[0065] The position of the central word in the text will affect the degree of its semantic influence on the surrounding word segmentations. The sequential number can reflect the position information and serve as the central position parameter of the Gaussian distribution function, that is, the first function parameter μ, representing the position of the central word in the word segmentation set.

[0066] Combine the preset second function parameter σ and the first function parameter μ corresponding to the current central word sample to construct a Gaussian distribution function with the sequential number corresponding to the word segmentation in the current word segmentation set combination sample as the independent variable and the relevance between the word segmentation in the current word segmentation set combination sample and the word category corresponding to the current central word sample as the dependent variable, that is, the distribution function sample, which is used to describe the attenuation law of the relevance between the word segmentation and the word category corresponding to the central word sample with the distance (measured by the sequential number of the word segmentation). Specifically, μ determines the position of the corresponding distribution function sample on the number axis, and σ determines the width and height of the corresponding Gaussian distribution function, reflecting the influence range of the central word sample.

[0067] Further, substitute the sequence number of each word segment in the combined sample of the current word segment set into the distribution function sample corresponding to the current central word sample, and obtain a heat sub-value sample of each word segment for the word category corresponding to the current central word sample according to the positional relationship between the word segment and the current central word sample, which characterizes the correlation between each word segment and the word category corresponding to the current central word sample.

[0068] A word segment may be relevant to different word categories corresponding to multiple central words. By traversing all central word samples to comprehensively consider various correlations, and summarizing several heat sub-value samples of each word segment for each word category, such as summing or weighted summing, a heat total value sample of each word segment for each word category is obtained, which characterizes the correlation between each word segment and each word category, so as to obtain a more comprehensive and accurate quantitative result of the correlation between word segments and word categories.

[0069] Further, for any word category, combine the heat total value samples of each word segment in the combined sample of the current word segment set for this word category to obtain a reference heat data sample of the judicial document segment corresponding to the combined sample of the current word segment set for the word category corresponding to this word category, that is, the overall correlation information of this judicial document segment in this word category, as the target output for training the initial classification model, so as to calculate the difference with the actual output to characterize the output accuracy of the initial classification model.

[0070] As described above, by segmenting the judicial document segment sample, the central word sample and the corresponding word category are obtained. Using the Gaussian distribution function, based on the position information of the central word (i.e., the first function parameter μ) and the preset distribution range parameter (i.e., the second function parameter σ), the correlation heat value between each word segment and different word categories is calculated to accurately capture the semantic correlation between the word segments and different word categories in the judicial document. Finally, a reference heat data sample of each judicial document segment sample for each word category is obtained, which improves the accuracy of the reference heat data sample. Furthermore, when the reference heat data sample is used as the target output for training the initial classification model, the difference is calculated with the actual output to characterize the output accuracy of the initial classification model, thereby improving the accuracy of the preset classification model.

[0071] In a specific embodiment, S30 includes the following steps: S301, for any channel of the initial classification model and any judicial document segment sample, obtain the model sub-loss of the initial classification model for the current channel and the current judicial document segment sample according to the word heat data sample of the current judicial document segment sample corresponding to the current channel and the reference heat data sample of the current judicial document segment sample corresponding to the word category corresponding to the current channel.

[0072] S302, traverse all channels of the initial classification model and all judicial document fragment samples, and obtain the sub-loss of the model for each channel and each judicial document fragment sample of the initial classification model.

[0073] S303, determine the total loss of the model corresponding to the initial classification model as the sum of the sub-losses of the model for all channels and all judicial document fragment samples of the initial classification model.

[0074] Among them, calculating the sub-loss of the model for each channel and each judicial document fragment sample of the initial classification model characterizes the difference between the word heat data sample predicted by the model on a single channel and a single sample and the reference heat data sample. Then, by traversing, all sub-losses of the model are obtained, and the total loss of the model of the initial classification model is obtained by accumulation, reflecting the overall performance of the classification model on the entire training sample set.

[0075] S2, for any word heat data, determine several segmentation words corresponding to several heat centers of the current word heat data in the judicial document fragment corresponding to the current word heat data as several central words corresponding to the judicial document fragment corresponding to the current word heat data. Among them, the heat center refers to the position where the segmentation word corresponding to the heat value greater than the preset heat value threshold in the word heat data is located in the corresponding judicial document fragment.

[0076] In a specific embodiment, S2 includes the following steps: S21, for any word heat data, obtain several heat centers corresponding to the current word heat data according to the heat value corresponding to each segmentation word in the current word heat data and the preset heat value threshold.

[0077] S22, for any heat center corresponding to the current word heat data, determine the segmentation word corresponding to the current heat center in the judicial document fragment corresponding to the current word heat data as the central word corresponding to the judicial document fragment corresponding to the current word heat data for the current heat center.

[0078] S23, traverse all heat centers corresponding to the current word heat data, and obtain all central words corresponding to the judicial document fragment corresponding to the current word heat data.

[0079] Among them, the preset heat value threshold is a numerical standard set in advance, used to judge whether the heat value of the segmentation word is high enough to determine whether the position corresponding to the segmentation word is a heat center. Specifically, when the heat value of the segmentation word is greater than the preset heat value threshold, the position where the corresponding segmentation word is located can be regarded as a heat center, and then the segmentation word corresponding to the heat center in the judicial document fragment is determined as the corresponding central word to reflect the main theme and key information of the target judicial document.

[0080] In a specific embodiment, the heat value ranges from 0 to 1. The closer the value is to 1, the stronger the correlation. The closer the value is to 0, the weaker the correlation. The specific value of the preset heat value threshold can be set by the implementer according to the actual situation. For example, the heat value threshold can be the maximum value of the dependent variable corresponding to the distribution function. For a Gaussian distribution function, the maximum value of the dependent variable is determined by the second function parameter σ. Specifically, the heat value threshold Z = 1 / (σ×(2×π) 1 / 2 ). When the amount of judicial document data to be processed is large, in order to avoid screening out too many irrelevant heat centers, the heat value threshold can be appropriately increased to ensure that the extracted central words have high representativeness and significance. Conversely, for a small-scale judicial document data, the heat value threshold can be relatively reduced to make full use of the limited data information. Or, if the accuracy requirement for information extraction is high and it is desired to obtain only the core information highly relevant to a specific word category, the heat value threshold should be set higher to ensure that the screened central words have high relevance and accuracy. Conversely, if subsequent tasks require more comprehensive and rich information, such as text summary generation or knowledge graph construction, etc., the heat value threshold is appropriately reduced to obtain more potential central words and provide a more sufficient information basis for subsequent processing.

[0081] As described above, by screening to obtain the heat centers and determining the central words, the key information highly relevant to a specific word category can be accurately extracted from the judicial document fragments, thus facilitating the focus of the analysis on the key central words, avoiding meaningless searches and analyses in a large amount of irrelevant texts, simplifying the process of text analysis, helping to grasp the core content and key viewpoints of the documents more quickly, and thus improving the accuracy and efficiency of information extraction.

[0082] S3. Extract the central sentences corresponding to each central word from the judicial document fragment corresponding to the current word heat data.

[0083] In a specific embodiment, S3 includes the following steps: S31. Segment each judicial document fragment according to the preset punctuation marks to obtain a number of segmented sentences corresponding to each judicial document fragment.

[0084] S32. For any central word corresponding to the judicial document fragment corresponding to the current word heat data, determine the segmented sentence corresponding to the current central word in the current judicial document fragment as the central sentence corresponding to the current central word.

[0085] S33. Traverse each central word corresponding to the judicial document fragment corresponding to the current word heat data to obtain all the central sentences corresponding to the current word heat data.

[0086] Among them, using preset punctuation marks, such as full stops, question marks, exclamation marks, semicolons, etc. as delimiter marks, each judicial document segment is cut and processed, divided into several relatively independent segmented sentences, providing a basic text unit for subsequently determining the central sentence corresponding to the central word, enabling more precise positioning and extraction of sentences related to the central word, and improving the accuracy and efficiency of text processing.

[0087] The central word is a word with important semantics in the judicial document segment, and the sentence where it is located usually contains key information related to this central word. Therefore, by associating the central word with the segmented sentence where it is located, the text content closely related to the central word can be extracted, which helps to further understand the meaning and role of the central word in the context.

[0088] As mentioned above, by extracting the central sentence corresponding to the central word, the analysis focus can be concentrated on the text content related to important semantics, avoiding ineffective searches in a large amount of irrelevant text information, improving the efficiency of information extraction and the efficiency of subsequent quality analysis. Moreover, the central sentence contains specific descriptions and context information of the central word, which helps to more deeply understand the meaning of the central word and its role in the judicial document, thereby better grasping the semantic structure and logical relationship of the entire judicial document, and improving the accuracy of subsequent quality analysis.

[0089] S4. Input each central sentence corresponding to the current word heat data and the word category corresponding to the current word heat data into a preset large language model, and obtain several keywords in each central sentence corresponding to the current word heat data and the word category corresponding to each keyword.

[0090] S5. Traverse all word heat data to obtain all keywords corresponding to the target judicial document and the word category corresponding to each keyword.

[0091] Among them, the preset large language model refers to a pre-trained model with powerful language understanding and generation capabilities. Based on models such as Chat-GPT and BERT, after fine-tuning or training with judicial domain data, it can be used for tasks such as keyword extraction and quality analysis report generation of judicial documents. Specifically, the preset large language model has been trained with a large amount of judicial domain text data and has the ability to semantically understand and analyze natural language text in the judicial domain. By inputting the central sentence and the word category, the preset large language model can identify important words related to this word category as keywords according to the learned language patterns and semantic knowledge, so as to accurately reflect the core content and semantic focus of the central sentence, providing key information for subsequent quality analysis of judicial documents.

[0092] Furthermore, based on the above-mentioned text data in the judicial field, the extracted keywords, and the word categories, a large number of existing high-quality judicial document quality analysis reports are used as reference data. The judicial document quality analysis reports cover the analysis and evaluation of judicial documents of different types and complexities, including multiple evaluation dimensions and specific analysis contents such as whether the factual description in the document is accurate and complete, whether the legal application is correct and reasonable, whether the recorded content is balanced, and whether the logic is reasonable, as well as specific problem points and improvement suggestions.

[0093] During training, it is preset that the large language model will learn the structure and language expression methods of the quality analysis reports, and understand how to analyze and evaluate various aspects of judicial documents starting from keywords and word categories. For example, learn how to judge whether the description of the parties' information in the judicial document is complete and accurate based on the keywords in the "parties' information" category; how to evaluate the correctness of the legal application in the document based on the keywords in the "legal basis" category.

[0094] At the same time, according to the differences between the generated quality analysis reports and the reference quality analysis reports, the parameters of the preset large language model can be continuously adjusted to improve the accuracy, logic, and readability of the generated quality analysis reports. For example, adjust the parameters of the preset large language model in each aspect through the similarity between the analysis results of each aspect in the generated quality analysis report and the corresponding parts of the reference quality analysis report, so that the preset large language model can comprehensively and deeply analyze the judicial documents from multiple dimensions according to the learned language patterns and semantic knowledge, based on all the keywords corresponding to the input target judicial document and the word category corresponding to each keyword, and generate a rich, accurate, and instructive quality analysis report to provide strong support for the quality assessment and improvement of judicial documents.

[0095] Different word categories have different semantic characteristics and context requirements. Inputting the central sentences corresponding to each word category into the preset large language model respectively, the preset large language model can conduct more targeted analysis and understanding for specific categories. For example, for the "parties' information" category, when the preset large language model processes the relevant central sentences, it will be more focused on identifying words related to the identity of the person as keywords, without being interfered by information from other categories, thereby improving the accuracy of keyword extraction. At the same time, it reduces the types of different semantic information that the preset large language model needs to process simultaneously, thus reducing the difficulty of the large language model's understanding and processing. At the same time, when the central sentences corresponding to each word category are input into the preset large language model respectively, the obtained keywords will also be naturally classified according to the word categories, which is convenient for summarizing keywords according to different word categories, understanding the distribution and importance of different categories in the document, and providing a data basis for the subsequent generated quality analysis report.

[0096] By traversing all the word heat data, it is ensured that keyword extraction is performed on every part of the target judicial document, comprehensively covering the content of the document and avoiding omission of important information.

[0097] As described above, the complete keyword set and the corresponding word category information of the target judicial document are obtained through the large language model, providing a rich and comprehensive data basis for subsequent quality analysis. It can more accurately evaluate the content and quality of the judicial document, improve the accuracy of keyword extraction. Moreover, the large language model only performs keyword extraction on the central sentences, avoiding ineffective word extraction in a large amount of irrelevant text information, reducing the resources occupied in the calculation and reasoning process, and improving the efficiency of keyword extraction.

[0098] S6. Input all the keywords corresponding to the target judicial document and the word category corresponding to each keyword into the preset large language model to obtain the quality analysis report corresponding to the target judicial document.

[0099] Among them, the preset large language model analyzes and evaluates the content of the judicial document according to all the keywords corresponding to the target judicial document and the word category corresponding to each keyword. Keywords and word categories reflect the core content and semantic structure of the judicial document. By analyzing the keywords and word categories, the quality of the document in various aspects can be evaluated. For example, the accuracy and completeness of keywords can reflect whether the document clearly and accurately elaborates on the facts and legal points. The distribution of keywords under different word categories can reflect the balance and logic of the document content. Specifically, accurate keywords and corresponding reasonable word categories can accurately extract the key information in the target judicial document. If the facts and legal points in the target judicial document are clearly and accurately elaborated, the extracted keywords will accurately correspond to the corresponding word categories and can cover all the key aspects involved in the target judicial document, indicating that the target judicial document clearly and accurately elaborates on the facts and legal points. For example, in a judicial document of a contract dispute, the keyword "contract breach clause" belongs to the "legal basis" category, the keyword "contract signing date" belongs to the "case facts" category, and the keywords "plaintiff", "defendant", and "third party" belong to the "party information" category. If the above keywords are accurately extracted and the categories are correct, it indicates that the judicial document clearly and accurately elaborates on the legal provisions on which the contract dispute is based and the key facts of the case. If keywords belonging to the "case facts" category such as "description of breach of contract" and "amount of loss" are also extracted, it further reflects the completeness of the fact elaboration in the judicial document.

[0100] The distribution of keywords under different word categories can reflect the degree of emphasis and correlation of the content of the target judicial document in different aspects. A reasonable distribution means that the target judicial document pays appropriate attention to each key area, without excessive verbosity or omission of a certain type of content, reflecting the balance of the document content. At the same time, the logical relationship between keywords can be reflected through the word categories to which the keywords belong. For example, keywords such as "filing a case", "hearing", and "judgment" belonging to the "judicial procedure" category appear in the order of the judicial process, reflecting the logic in describing the judicial procedure in the content of the target judicial document.

[0101] According to the analysis results in various aspects, a quality analysis report corresponding to the target judicial document is generated. The quality analysis report can include evaluations of the integrity, accuracy, logic, etc. of the judicial document content, as well as specific problem pointing and improvement suggestions, to help judicial personnel or relevant personnel quickly understand the quality status of the document, discover existing problems, and make targeted improvements, thereby improving the quality and standardization of judicial documents.

[0102] As described above, starting from the word vector data, the preset classification model is used to output the word heat data, and the correlation between the word segmentation and the word categories is quantitatively displayed in the form of heat values, so as to accurately locate the word segmentation related to each word category in the target judicial document. Then, through a step-by-step manner, the central word, central sentence, and keywords are determined, providing a rich and comprehensive data basis for subsequent quality analysis, improving the accuracy of the quality analysis report, and effectively filtering out irrelevant text information in the target judicial document during the keyword extraction process, enabling subsequent data extraction and analysis to focus on the core content of the target judicial document, reducing the resources occupied by the large language model during the calculation and reasoning process, and improving the efficiency of the quality analysis process.

[0103] Embodiment 2 Embodiment 2 provides a big data-based judicial document analysis system, as Figure 2 shown. The big data-based judicial document analysis system includes: A word classification module 21, configured to input the word vector data corresponding to each judicial document segment of the target judicial document into a preset classification model, and obtain the word heat data corresponding to each channel of the preset classification model for each judicial document segment. Wherein, each channel of the preset classification model corresponds to a word category, and each word heat data is used to represent the correlation between each word segmentation in the corresponding judicial document segment and the corresponding word category through a heat value.

[0104] The central word screening module 22 is used to determine, for any word heat data, several words corresponding to several heat centers in the judicial document segment corresponding to the current word heat data as several central words corresponding to the judicial document segment corresponding to the current word heat data.

[0105] The central sentence extraction module 23 is used to extract the central sentence corresponding to each central word from the judicial document segment corresponding to the current word heat data.

[0106] The first keyword extraction module 24 is used to input each central sentence corresponding to the current word heat data and the word category corresponding to the current word heat data into a preset large language model to obtain several keywords in each central sentence corresponding to the current word heat data and the word category corresponding to each keyword.

[0107] The second keyword extraction module 25 is used to traverse all word heat data to obtain all keywords corresponding to the target judicial document and the word category corresponding to each keyword.

[0108] The quality analysis module 26 is used to input all keywords corresponding to the target judicial document and the word category corresponding to each keyword into a preset large language model to obtain a quality analysis report corresponding to the target judicial document.

[0109] In a specific embodiment, the word classification module 21 includes: The word segmentation set acquisition sub-module is used to perform word segmentation on each initial sentence in the target judicial document to obtain a word segmentation set corresponding to each initial sentence, where the words in each word segmentation set are arranged in the front-back position order in the corresponding initial sentence.

[0110] The quantity calculation sub-module is used to obtain the word vector input quantity corresponding to the preset classification model according to the target input length corresponding to the preset classification model and the preset length of each word vector.

[0111] The combination acquisition sub-module is used to group all initial sentences according to the word vector input quantity and the number of words in the word segmentation set corresponding to each initial sentence to obtain several initial sentence combinations and a word segmentation set combination corresponding to each initial sentence combination, where the total number of words corresponding to each initial sentence combination is less than or equal to the word vector input quantity.

[0112] The judicial document segment determination sub-module is used to determine each initial sentence combination as a judicial document segment.

[0113] The word vector data acquisition sub-module is used to perform word vector conversion on each word in the combined word segmentation sets corresponding to each judicial document segment, and obtain the word vector data corresponding to each judicial document segment.

[0114] In a specific embodiment, the word classification module 21 further includes: The sample data acquisition sub-module is used to obtain a number of judicial document segment samples, the word vector data samples corresponding to each judicial document segment sample, and the reference heat data samples corresponding to each judicial document segment sample for each word category.

[0115] The sample word classification sub-module is used to input the word vector data samples corresponding to each judicial document segment sample into the initial classification model, and obtain the word heat data samples corresponding to each channel of the initial classification model for each judicial document segment sample.

[0116] The loss calculation sub-module is used to obtain the total model loss corresponding to the initial classification model according to the word heat data samples corresponding to each channel of the initial classification model for each judicial document segment sample and the reference heat data samples corresponding to each judicial document segment sample for each word category.

[0117] The model training sub-module is used to update the parameters of the initial classification model according to the total model loss until the total model loss converges, and obtain the trained preset classification model.

[0118] In a specific embodiment, the sample data acquisition sub-module includes: The sample data acquisition unit is used to obtain the combined word segmentation set samples corresponding to each judicial document segment sample, a number of central word samples corresponding to each combined word segmentation set sample, and the word categories corresponding to each central word sample.

[0119] The first function parameter acquisition unit is used to determine the first function parameter corresponding to a current central word sample in any combined word segmentation set sample according to the sequence number of the current central word sample in the current combined word segmentation set sample, where the first function parameter is used to determine the position of the corresponding distribution function sample on the number axis.

[0120] The distribution function acquisition unit is used to obtain the distribution function sample corresponding to the current central word sample according to the preset second function parameter and the first function parameter corresponding to the current central word sample, where the distribution function sample takes the sequence number corresponding to the word in the current combined word segmentation set sample as the independent variable, and the correlation between the word in the current combined word segmentation set sample and the word category corresponding to the current central word sample as the dependent variable, where the preset second function parameter is used to determine the width and height of the corresponding distribution function sample.

[0121] The first heat sub-value acquisition unit is configured to obtain, according to the distribution function sample corresponding to the current central word sample, the heat sub-value sample of each word segmentation in the current word segmentation set combination sample for the word category corresponding to the current central word sample.

[0122] The second heat sub-value acquisition unit is configured to traverse all the central word samples in the current judicial document segment sample, and obtain a plurality of heat sub-value samples of each word segmentation in the current word segmentation set combination sample for each word category.

[0123] The total heat value acquisition unit is configured to obtain, according to a plurality of heat sub-value samples of each word segmentation in the current word segmentation set combination sample for each word category, the total heat value sample of each word segmentation in the current word segmentation set combination sample for each word category.

[0124] The first reference data acquisition unit is configured to, for any word category, obtain, according to the total heat value sample of each word segmentation in the current word segmentation set combination sample for the current word category, the reference heat data sample of the judicial document segment corresponding to the current word segmentation set combination sample for the current word category.

[0125] The second reference data acquisition unit is configured to traverse all the word segmentation set combination samples and all the word categories, and obtain the reference heat data sample of each judicial document segment for each word category.

[0126] In a specific embodiment, the loss calculation sub-module includes: The first model sub-loss calculation unit is configured to, for any channel and any judicial document segment sample of the initial classification model, obtain the model sub-loss of the initial classification model for the current channel and the current judicial document segment sample according to the word heat data sample corresponding to the current judicial document segment sample for the current channel and the reference heat data sample of the word category corresponding to the current channel for the current judicial document segment sample.

[0127] The second model sub-loss calculation unit is configured to traverse all the channels and all the judicial document segment samples of the initial classification model, and obtain the model sub-loss of the initial classification model for each channel and each judicial document segment sample.

[0128] The model total loss calculation unit is configured to determine the total sum of the model sub-losses of the initial classification model for all channels and all judicial document segment samples as the model total loss corresponding to the initial classification model.

[0129] In a specific embodiment, the central word screening module 22 includes: A heat center acquisition sub-module, which is used for any word heat data, and according to the heat value corresponding to each word segmentation in the current word heat data and a preset heat value threshold, obtains several heat centers corresponding to the current word heat data.

[0130] A first central word acquisition sub-module, which is used for any heat center corresponding to the current word heat data, and determines the word segmentation corresponding to the current heat center in the judicial document segment corresponding to the current word heat data as the central word corresponding to the judicial document segment corresponding to the current word heat data for the current heat center.

[0131] A second central word acquisition sub-module, which is used to traverse all heat centers corresponding to the current word heat data, and obtains all central words corresponding to the judicial document segment corresponding to the current word heat data.

[0132] In a specific embodiment, the central sentence extraction module 23 includes: A sentence splitting sub-module, which is used to split each judicial document segment according to preset punctuation marks, and obtains several split sentences corresponding to each judicial document segment.

[0133] A first central sentence acquisition sub-module, which is used for any central word corresponding to the judicial document segment corresponding to the current word heat data, and determines the split sentence corresponding to the current central word in the current judicial document segment as the central sentence corresponding to the current central word.

[0134] A second central sentence acquisition sub-module, which is used to traverse each central word corresponding to the judicial document segment corresponding to the current word heat data, and obtains all central sentences corresponding to the current word heat data.

[0135] It should be noted that for the information interaction, execution process, etc. between the above modules, since they are based on the same concept as the method embodiment of the present invention, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0136] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to obtain equivalent embodiments with equivalent changes, but as long as it does not depart from the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A judicial document analysis method based on big data, characterized in that: The judicial document analysis method based on big data includes the following steps: S1, inputting the word vector data corresponding to each judicial document segment corresponding to the target judicial document into the preset classification model, and obtaining the word thermal data of each judicial document segment corresponding to each channel of the preset classification model, wherein each channel of the preset classification model corresponds to a word category, and each word thermal data is used to characterize the correlation between each word segment in the corresponding judicial document segment and the corresponding word category through the thermal value; S2, for any word thermal data, determine the several word segments corresponding to the several thermal centers corresponding to the current word thermal data in the judicial document segment corresponding to the current word thermal data as the several central words corresponding to the judicial document segment corresponding to the current word thermal data, wherein the thermal center refers to the location of the word segment corresponding to the thermal value greater than the preset thermal value threshold in the word thermal data in the corresponding judicial document segment; S3, extracting the central sentence corresponding to each central word from the judicial document fragment corresponding to the current word thermal data; S4, inputting each central sentence corresponding to the current word thermal data and the word category corresponding to the current word thermal data into a preset large language model, and obtaining a number of keywords in each central sentence corresponding to the current word thermal data and the word category corresponding to each keyword; S5, traversing all word thermal data to obtain all keywords corresponding to the target judicial document and the word category corresponding to each keyword; S6, inputting all keywords corresponding to the target judicial document and the word category corresponding to each keyword into the preset large language model, and obtaining a quality analysis report corresponding to the target judicial document.

2. The judicial document analysis method based on big data according to claim 1 is characterized in that: S1 includes the following steps: S11, segmenting each initial sentence in the target judicial document to obtain a segmentation set corresponding to each initial sentence, wherein the segmentations in each segmentation set are arranged in the order of their front and back positions in the corresponding initial sentence; S12, obtaining the number of word vector inputs corresponding to the preset classification model according to the target input length corresponding to the preset classification model and the preset length of each word vector; S13, grouping all the initial sentences according to the number of word vector inputs and the number of segmentations in the segmentation set corresponding to each initial sentence, obtaining a plurality of initial sentence combinations and a segmentation set combination corresponding to each initial sentence combination, wherein the total number of segmentations corresponding to each initial sentence combination is less than or equal to the number of word vector inputs; S14, each initial sentence combination is identified as a judicial document fragment; S15, performing word vector conversion on each word in the word set combination corresponding to each judicial document fragment, and obtaining word vector data corresponding to each judicial document fragment.

3. The judicial document analysis method based on big data according to claim 1 is characterized in that: The preset classification model is trained by the following steps: S10, obtaining a number of judicial document fragment samples, a word vector data sample corresponding to each judicial document fragment sample, and a reference thermal data sample corresponding to each word category of each judicial document fragment sample; S20, inputting the word vector data sample corresponding to each judicial document fragment sample into the initial classification model, and obtaining the word thermal data sample of each channel of the initial classification model corresponding to each judicial document fragment sample; S30, obtaining a total model loss corresponding to the initial classification model according to a word thermal data sample of each channel of the initial classification model corresponding to each judicial document fragment sample and a reference thermal data sample corresponding to each word category of each judicial document fragment sample; S40, updating the parameters of the initial classification model according to the total loss of the model until the total loss of the model converges, thereby obtaining a trained preset classification model.

4. The judicial document analysis method based on big data according to claim 3 is characterized in that: S10 includes the following steps: S101, obtaining a word segmentation set combination sample corresponding to each judicial document fragment sample, a number of core word samples corresponding to each word segmentation set combination sample, and a word category corresponding to each core word sample; S102, for any central word sample in any word segmentation set combination sample, determine a first function parameter corresponding to the current central word sample according to the sequence number of the current central word sample in the current word segmentation set combination sample, wherein the first function parameter is used to determine the position of the corresponding distribution function sample on the number axis; S103, according to the preset second function parameter and the first function parameter corresponding to the current center word sample, obtain the distribution function sample corresponding to the current center word sample, wherein the distribution function sample uses the sequence number corresponding to the segmentation in the current segmentation set combination sample as an independent variable, and uses the correlation between the segmentation in the current segmentation set combination sample and the word category corresponding to the current center word sample as a dependent variable, wherein the preset second function parameter is used to determine the width and height of the corresponding distribution function sample; S104, according to the distribution function sample corresponding to the current central word sample, obtaining a thermal sub-value sample of each word in the current word segmentation set combination sample for the word category corresponding to the current central word sample; S105, traversing all the central word samples in the current judicial document fragment sample, and obtaining a number of thermal sub-value samples for each word category of each segmentation in the current segmentation set combination sample; S106, obtaining a total thermal value sample of each segmentation in the current segmentation set combined sample for each word category according to a plurality of thermal sub-value samples of each segmentation in the current segmentation set combined sample for each word category; S107, for any word category, according to the total thermal value sample of each word in the current word segmentation set combination sample for the current word category, obtain the reference thermal data sample corresponding to the judicial document fragment corresponding to the current word segmentation set combination sample for the current word category; S108, traverse all word segmentation set combination samples and all word categories to obtain reference thermal data samples corresponding to each word category for each judicial document segment.

5. The judicial document analysis method based on big data according to claim 3 is characterized in that: S30 includes the following steps: S301, for any channel and any judicial document fragment sample of the initial classification model, according to the word thermal data sample corresponding to the current channel of the current judicial document fragment sample and the reference thermal data sample corresponding to the word category corresponding to the current channel of the current judicial document fragment sample, obtain the model sub-loss of the initial classification model for the current channel and the current judicial document fragment sample; S302, traversing all channels and all judicial document fragment samples of the initial classification model, and obtaining the model sub-loss of the initial classification model for each channel and each judicial document fragment sample; S303: Determine the sum of the model sub-losses of the initial classification model for all channels and all judicial document fragment samples as the total model loss corresponding to the initial classification model.

6. The judicial document analysis method based on big data according to claim 1 is characterized in that: S2 includes the following steps: S21, for any word thermal data, according to the thermal value corresponding to each word segment in the current word thermal data and the preset thermal value threshold, obtain a number of thermal centers corresponding to the current word thermal data; S22, for any thermal center corresponding to the thermal data of the current word, determine the segment corresponding to the current thermal center in the judicial document segment corresponding to the thermal data of the current word as the central word corresponding to the current thermal center in the judicial document segment corresponding to the thermal data of the current word; S23, traverse all thermal centers corresponding to the thermal data of the current word, and obtain all central words corresponding to the judicial document fragments corresponding to the thermal data of the current word.

7. The judicial document analysis method based on big data according to claim 1 is characterized in that: S3 includes the following steps: S31, segmenting each judicial document segment according to preset punctuation marks, and obtaining a plurality of segmentation sentences corresponding to each judicial document segment; S32, for any central word corresponding to the judicial document segment corresponding to the current word thermal data, determining the segmentation sentence corresponding to the current central word in the current judicial document segment as the central sentence corresponding to the current central word; S33, traverse each central word corresponding to the judicial document fragment corresponding to the current word thermal data, and obtain all central sentences corresponding to the current word thermal data.

8. A judicial document analysis system based on big data, characterized in that: The judicial document analysis system based on big data includes: A word classification module, used to input the word vector data corresponding to each judicial document segment corresponding to the target judicial document into a preset classification model, and obtain the word thermal data of each judicial document segment corresponding to each channel of the preset classification model, wherein each channel of the preset classification model corresponds to a word category, and each word thermal data is used to characterize the correlation between each word segment in the corresponding judicial document segment and the corresponding word category through a thermal value; The central word screening module is used to determine, for any word thermal data, several participles corresponding to several thermal centers corresponding to the current word thermal data in the judicial document fragment corresponding to the current word thermal data as several central words corresponding to the judicial document fragment corresponding to the current word thermal data, wherein the thermal center refers to the location of the participle corresponding to the thermal value greater than the preset thermal value threshold in the word thermal data in the corresponding judicial document fragment; The core sentence extraction module is used to extract the core sentence corresponding to each core word from the judicial document fragment corresponding to the current word thermal data; The first keyword extraction module is used to input each central sentence corresponding to the current word thermal data and the word category corresponding to the current word thermal data into a preset large language model, and obtain a number of keywords in each central sentence corresponding to the current word thermal data and the word category corresponding to each keyword; The second keyword extraction module is used to traverse all the word thermal data to obtain all the keywords corresponding to the target judicial document and the word category corresponding to each keyword; The quality analysis module is used to input all keywords corresponding to the target judicial document and the word category corresponding to each keyword into the preset large language model to obtain a quality analysis report corresponding to the target judicial document.

9. The judicial document analysis system based on big data according to claim 8 is characterized in that: The word classification module includes: A word segmentation set acquisition submodule is used to segment each initial sentence in the target judicial document to obtain a word segmentation set corresponding to each initial sentence, wherein the word segmentations in each word segmentation set are arranged in the order of their front and back positions in the corresponding initial sentence; A quantity calculation submodule, used to obtain the number of word vector inputs corresponding to the preset classification model according to the target input length corresponding to the preset classification model and the preset length of each word vector; A combination acquisition submodule is used to group all the initial sentences according to the number of word vector inputs and the number of segmentations in the segmentation set corresponding to each initial sentence, and obtain a plurality of initial sentence combinations and a segmentation set combination corresponding to each initial sentence combination, wherein the total number of segmentations corresponding to each initial sentence combination is less than or equal to the number of word vector inputs; A judicial document segment determination submodule, used for determining each initial sentence combination as a judicial document segment; The word vector data acquisition submodule is used to perform word vector conversion on each word in the word set combination corresponding to each judicial document fragment, and obtain the word vector data corresponding to each judicial document fragment.

10. The judicial document analysis system based on big data according to claim 8 is characterized in that: The central word screening module includes: The thermal center acquisition submodule is used to obtain a number of thermal centers corresponding to the thermal data of the current word according to the thermal value corresponding to each word in the thermal data of the current word and the preset thermal value threshold for any thermal data of the word; The first central word acquisition submodule is used to determine, for any thermal center corresponding to the thermal data of the current word, the corresponding word segment of the current thermal center in the judicial document segment corresponding to the thermal data of the current word as the central word corresponding to the current thermal center in the judicial document segment corresponding to the thermal data of the current word; The second center word acquisition submodule is used to traverse all thermal centers corresponding to the thermal data of the current word, and obtain all center words corresponding to the judicial document fragments corresponding to the thermal data of the current word.

Citation Information

Patent Citations

  • Judgment document retrieval method based on semantic matching and server

    CN106502996A

  • Structural analysis method and system for judicial documents

    CN111145052A

  • Imbalanced judicial judgment document data-oriented law article recommendation method and system

    CN114610891A

  • Method and apparatus for realizing element recognition in judicial document

    WO2020114373A1