Batch right protection case intelligent filing system and method based on multi-source data analysis
Through multi-source data analysis technology and intelligent case filing system, batch rights protection case information is automatically extracted and reviewed, which solves the manpower consumption and error problems of traditional case filing methods and realizes an efficient and accurate case filing process.
Patent Information
- Application Number
- CN202510654472.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional method of filing batch rights protection cases relies on manual review, which consumes a lot of manpower and time and is prone to errors. In addition, the online and offline information is inconsistent, which increases the difficulty and workload of the review.
Using multi-source data analysis technology, key information is extracted through natural language processing and convolutional neural networks. Combined with Bayesian theorem and knowledge graphs, case information is automatically reviewed, and application documents for filing are generated and submitted.
It has realized a fully automated case filing process, reduced the time for manual input and review, improved the accuracy of information correlation, reduced the risk of misjudgment, and ensured the timeliness and accuracy of the basis for case filing.
Smart Images

Figure CN120598730A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent processing of court business, and in particular to a system and method for intelligent filing of batch rights protection cases based on multi-source data analysis. Background Art
[0002] In today's era of rapid digitalization and informatization, with the widespread application of Internet technology and the massive growth of various types of data, people's lifestyles and social operation models have undergone profound changes. In the judicial field, this trend has also brought new opportunities and challenges, especially in the handling of batch rights protection cases. The traditional case filing method has gradually exposed many limitations, giving rise to an urgent need for an intelligent case filing system.
[0003] Batch rights protection cases often involve a large number of parties, and the amount of evidence materials and demand information data submitted by them is huge and complicated. The traditional case filing method mainly relies on manual review and entry of paper or electronic documents one by one, which not only consumes a lot of manpower and time, but is also prone to human errors. The information submitted by the parties through the online platform may be different from the content of offline paper materials. There is a lack of effective connection and coordination mechanism between the data, which makes it difficult for case filing personnel to obtain the required information comprehensively and accurately when reviewing the case, increasing the difficulty and workload of the review. Summary of the Invention
[0004] In response to the technical problems existing in the prior art, the present invention provides an intelligent filing system and method for batch rights protection cases based on multi-source data analysis.
[0005] The present invention solves the above-mentioned technical problems with the following technical solutions: an intelligent case filing system for batch rights protection cases based on multi-source data analysis, comprising:
[0006] Data acquisition and analysis module: collects rights protection case data from multiple sources, classifies the text using natural language processing technology, annotates key information to build a data set, and trains a model using a convolutional neural network to extract relevant information from the text;
[0007] Case information integration module: Assigns unique identifiers to entities and associates data to form case profiles, builds knowledge graphs based on laws, regulations, and judicial cases, and determines case types by calculating probabilities using Bayesian theorem;
[0008] Intelligent case filing module: This module collects the rules and legal provisions for filing various rights protection cases, categorizes and organizes them for automatic case review, combines historical data to train models to predict the likelihood of winning and risks, and integrates case information, rule judgments, and prediction results to provide a basis for generating case filing strategies.
[0009] Batch filing application generation module: integrates case information that meets the filing conditions, automatically fills in the document template to generate the filing application document, and after content verification, submits the online filing application to the court filing system through the API interface.
[0010] In a preferred embodiment, the data acquisition and analysis module collects data related to rights protection cases from multiple sources, converts audio and video data into text data through audio-video conversion, performs preliminary cleaning on the collected data, removes noise from the text data collected from multiple channels and unifies the text encoding format, uses a word segmentation tool to segment the text into individual words, uses natural language processing technology to perform text classification, defines text classification categories based on the type and needs of the rights protection case, extracts features from the segmented text by calculating the word frequency in the text, annotates the names of the parties, company names, locations, and dates in the text, constructs a training data set, and uses the annotated training data to perform model training through a convolutional neural network model. The specific steps of model training are as follows:
[0011] S1. Divide the labeled data into training set, validation set and test set, and divide the data proportionally;
[0012] S2. Determine the length of the word frequency vector and pass the word frequency vector as input to the convolutional neural network;
[0013] S3. Perform a convolution operation on the input word frequency vector, sliding the convolution kernel over the input data to identify keywords related to rights protection cases.
[0014] S4, downsample the output of the convolutional layer through the pooling layer to reduce the dimension of the data and retain important feature information, thereby highlighting the important features in the text;
[0015] S5. Flatten the output of the pooling layer into a one-dimensional vector and connect it to the fully connected layer. The fully connected layer combines and transforms the identified features and learns the complex relationships between the features.
[0016] S6, the output layer uses the softmax activation function to output the probability distribution of each labeled category;
[0017] S7. Use the cross-entropy loss function to measure the difference between the model's prediction results and the true annotations. Use stochastic gradient descent to update the model's parameters according to the gradient of the loss function, so that the model's prediction results gradually approach the true annotations. Define evaluation metrics to help evaluate the performance of the model during training.
[0018] S8. Input the training set data into the convolutional neural network model and perform multiple iterations of training. Each iteration is called an epoch. In each epoch, the model performs forward propagation and backward propagation on all samples in the training set to update the model parameters. After each epoch, the model is evaluated using the validation set data to monitor the performance changes of the model. Based on the evaluation results of the validation set, the model hyperparameters can be adjusted.
[0019] S9. Evaluate the trained model based on the test set data, calculating the model's accuracy, recall, and F1 value on the test set. Deploy the trained convolutional neural network model to the actual system. During system operation, input text data from information rights protection cases into the model. The model can automatically extract the names of the parties, company names, locations, and dates from the text to support the case filing process.
[0020] The extracted information is standardized and stored in a relational database. An index is created by case ID, and the name, company, location, and date field tables are associated. Business information and administrative division change data are synchronized regularly, and the company name and location fields are updated.
[0021] In a preferred embodiment, the case information integration module assigns a unique identifier to each entity, uses the identity card number of the party as the unique identifier, and uses the social credit code of the enterprise as the unique identifier, generates a unique case number for each case as the unique identifier, uses the unique identifier of the party as a bridge, establishes an association between the party information table and the case information table, and sets a field in the case information table to store the unique identifier of the party. Through this field, the party and the case are matched, and the associated data are integrated into a unified case information model to form a case portrait. A knowledge map of rights protection cases is constructed based on laws, regulations and judicial cases, and legal and regulatory texts are collected. The collected legal and regulatory texts are stored in folders, named and classified according to the rules, and the collected legal and regulatory texts are sorted. Preprocess the legal texts, select the BERT-CRF model based on deep learning for named entity recognition, and prepare a legal text dataset for fine-tuning. The dataset is divided into training set, validation set and test set. Professional legal personnel annotate the legal texts in the training set and validation set, annotate the legal provision entities and legal concept entities, input the annotated training data into the BERT-CRF model, and fine-tune the model parameters through the back propagation algorithm to adapt it to the named entity recognition task of legal texts. Collect judicial case texts from the judicial judgment document network and the official website of the court, classify and number the case texts, establish a case database, use the text classification algorithm to classify the case texts, and determine the case type. Suppose the text D consists of feature words ω1, ω2, ..., ω n Composition, for each case type C j , according to Bayesian theorem, the text D belongs to type C j The probability P(C j |D), the specific calculation formula is as follows:
[0022]
[0023] Among them, P(C j ) indicates case type C j The probability of P(D|C j ) indicates that the known case type is C j The probability of text D appearing under the condition of , to determine the type of case.
[0024] In a preferred embodiment, the intelligent case filing module collects case filing rules and legal provisions related to various rights protection cases, and automatically reviews case information through rule judgment. When the rules are met, the case filing conditions are met. When the rules are not met, the case filing conditions are not met. The collected content is classified and sorted according to the case type and the scope of application of the rules. When making rule judgments, if the rules need to meet multiple conditions at the same time, in order to achieve accurate logical judgment, a double-condition compound judgment formula is used. The double-condition compound judgment formula is: meeting the conditions = A∧B. Further extended to the case of multiple conditions (A, B, ..., N), the following formula is used for judgment:
[0025]
[0026] Among them, C i The Boolean value of the i-th condition, f d Indicates that the condition is met. The value of A that meets the condition is 1, and the value of B that meets the condition is 0. When there is a need for fuzzy matching between case information and rule conditions, cosine similarity is used to calculate text similarity. Suppose the vector obtained by quantizing the rule description text is The vector of the case facts after the same quantization process is The similarity calculation formula is as follows:
[0027]
[0028] Among them, Similarity represents similarity, r i Represents a rule description text vector The i-th component of i Text vector representing case facts The i-th component of Represents a rule description text vector The norm of Text vector representing case facts norm, set a threshold θ, and when the calculated similarity ≥ θ, it is determined that the case facts and rule conditions are successfully matched. Historical rights protection case data are collected, and the trained model is used to predict the possibility of winning and the degree of risk of the case. Based on the prediction results, suggestions are provided for the formulation of the case filing strategy. The case information is integrated with the results of the rule judgment and the prediction results of the machine learning model to provide basic data for the generation of the case filing strategy.
[0029] In a preferred embodiment, the batch filing application generation module integrates case information that meets the filing conditions, automatically fills in the document template based on the case information, generates the filing application document, verifies the content of the generated document, and transmits the generated application data to the court's filing system through the API interface to realize online filing application submission.
[0030] The embodiment of the present invention further provides a method for intelligently filing batch rights protection cases based on multi-source data analysis, comprising the following steps:
[0031] S101. Collect rights protection case data from multiple sources, classify the text using natural language processing technology, annotate key information to build a data set, and train a model using a convolutional neural network to extract relevant information from the text;
[0032] S102. Assign unique identifiers to entities and associate data to form case profiles. Build a knowledge graph based on laws, regulations, and judicial cases. Calculate probabilities using Bayesian theorem to determine case types.
[0033] S103. Collect the rules and legal provisions for filing various rights protection cases, classify and organize them for automatic case review, combine historical data to train models to predict the likelihood of winning and risks, and integrate case information, rule judgments, and prediction results to provide a basis for generating case filing strategies;
[0034] S104. Integrate case information that meets the filing conditions, automatically fill in the document template to generate the filing application document, and after content verification, submit the online filing application to the court filing system through the API interface.
[0035] The beneficial effects of the present invention are: the present invention covers all types of evidence such as contracts, audio and video, chat records, etc., directly connects to government systems and social platforms through third-party interfaces, combines paper scanning and audio and video text conversion technology to ensure the integrity of the evidence chain, avoids the obstruction of case filing due to missing evidence, and automates the entire process from data cleaning, text processing to model training and document generation, reducing the time spent on manual entry and review. NLP technology is combined with the BERT-CRF model to achieve high-precision entity recognition of legal texts, and cooperates with the Bayesian algorithm to perform intelligent classification of case types, reducing the risk of misjudgment due to text ambiguity or annotation errors, integrating laws and regulations with judicial cases to build a knowledge graph, updating the correlation between legal provisions in real time, and assisting the rule engine to quickly locate applicable clauses, ensuring the timeliness and accuracy of the basis for case filing, combining multi-conditional logic formulas to quantify the feasibility of case filing, avoiding subjective experience bias, and providing parties with data-driven case filing strategy recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flow chart of the present invention;
[0037] Figure 2 This is a system block diagram of the present invention. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0039] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of this application, "plurality" means two or more, unless otherwise specifically specified.
[0040] In the description of this application, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art will recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.
[0041] Example 1
[0042] like Figure 1 This embodiment provides: a method for intelligently filing batch rights protection cases based on multi-source data analysis, comprising the following steps:
[0043] S101. Collect rights protection case data from multiple sources, classify the text using natural language processing technology, annotate key information to build a data set, and train a model using a convolutional neural network to extract relevant information from the text;
[0044] S102. Assign unique identifiers to entities and associate data to form case profiles. Build a knowledge graph based on laws, regulations, and judicial cases. Calculate probabilities using Bayesian theorem to determine case types.
[0045] S103. Collect the rules and legal provisions for filing various rights protection cases, classify and organize them for automatic case review, combine historical data to train models to predict the likelihood of winning and risks, and integrate case information, rule judgments, and prediction results to provide a basis for generating case filing strategies;
[0046] S104. Integrate case information that meets the filing conditions, automatically fill in the document template to generate the filing application document, and after content verification, submit the online filing application to the court filing system through the API interface.
[0047] Example 2
[0048] like Figure 2 This embodiment provides: an intelligent case filing system for batch rights protection cases based on multi-source data analysis, including:
[0049] Data acquisition and analysis module: collects rights protection case data from multiple sources, classifies the text using natural language processing technology, annotates key information to build a data set, and trains a model using a convolutional neural network to extract relevant information from the text;
[0050] In this embodiment, the data acquisition and parsing module needs to be specifically explained. The data acquisition and parsing module collects data related to rights protection cases from multiple sources, obtains electronic documents such as contracts, invoices, chat records, and video and audio evidence by establishing a third-party data interface, establishes telecommunication links with government systems, electronic platforms, and social media, scans paper documents, obtains scanned copies of paper documents, converts video and audio data into text data through audio-video conversion, performs preliminary cleaning on the collected data, removes duplicate, erroneous, and invalid information, removes noise from the text data collected from multiple channels, and unifies the text encoding format, uses word segmentation tools to segment the text into individual words, uses natural language processing technology to classify the text, defines text classification categories based on the type and needs of the rights protection case, such as intellectual property infringement, consumer rights disputes, labor disputes, etc., extracts features from the segmented text by calculating the word frequency in the text, annotates the names of the parties, company names, locations, and dates in the text, constructs a training data set, and trains the model using the annotated training data through a convolutional neural network model. The specific steps of model training are as follows:
[0051] S1. Divide the labeled data into training set, validation set and test set, and divide the data proportionally;
[0052] S2. Determine the length of the word frequency vector and pass the word frequency vector as input to the convolutional neural network;
[0053] S3. Perform a convolution operation on the input word frequency vector, sliding the convolution kernel over the input data to identify keywords related to rights protection cases.
[0054] S4, downsample the output of the convolutional layer through the pooling layer to reduce the dimension of the data and retain important feature information, thereby highlighting the important features in the text;
[0055] S5. Flatten the output of the pooling layer into a one-dimensional vector and connect it to the fully connected layer. The fully connected layer combines and transforms the identified features and learns the complex relationships between the features.
[0056] S6, the output layer uses the softmax activation function to output the probability distribution of each labeled category;
[0057] S7. Use the cross-entropy loss function to measure the difference between the model's prediction results and the true annotations. Use stochastic gradient descent to update the model's parameters according to the gradient of the loss function, so that the model's prediction results gradually approach the true annotations. Define evaluation metrics to help evaluate the performance of the model during training.
[0058] S8. Input the training set data into the convolutional neural network model and perform multiple iterations of training. Each iteration is called an epoch. In each epoch, the model performs forward propagation and backward propagation on all samples in the training set to update the model parameters. After each epoch, the model is evaluated using the validation set data to monitor the performance changes of the model. Based on the evaluation results of the validation set, the model hyperparameters can be adjusted.
[0059] S9. Evaluate the trained model based on the test set data, calculating the model's accuracy, recall, and F1 value on the test set. Deploy the trained convolutional neural network model to the actual system. During system operation, input text data from information rights protection cases into the model. The model can automatically extract the names of the parties, company names, locations, and dates from the text to support the case filing process.
[0060] The extracted information is standardized, including format unification and outlier processing. The standardized data is stored in a relational database, indexed by case ID, and associated with the name, company, location, and date field tables. Business information and administrative division change data are synchronized regularly, and the company name and location fields are updated.
[0061] It should be noted that audio and video conversion includes audio-to-text and video-to-text. Video-to-text uses digital signal processing technology to extract audio tracks from videos, adopts audio coding conversion algorithms to convert multi-format video files into standard audio formats, and uses neural network models to perform acoustic feature analysis on audio to achieve speech-to-text conversion. Optical character recognition is performed on embedded subtitles in the video, and timeline information is combined to generate timestamped text data. Audio-to-text uses open source models to recognize text in audio, and voiceprint feature analysis to achieve speech separation and identity labeling in multi-speaker scenarios.
[0062] Case information integration module: Assigns unique identifiers to entities and associates data to form case profiles, builds knowledge graphs based on laws, regulations, and judicial cases, and determines case types by calculating probabilities using Bayesian theorem;
[0063] In this embodiment, the case information integration module needs to be specifically explained. The case information integration module assigns a unique identifier to each entity. For the parties, their ID card numbers are used as the unique identifiers, and for the enterprises, their social credit codes are used as the unique identifiers. A unique case number is generated for each case as the unique identifier. The unique identifiers of the parties are used as a bridge to establish an association between the party information table and the case information table. In the case information table, a field is set to store the unique identifiers of the parties. Through this field, the parties are matched with the cases, and the associated data is integrated into a unified case information model to form a case profile.
[0064] Build a knowledge graph of rights protection cases based on laws, regulations and judicial cases, collect legal and regulatory texts, store the collected legal and regulatory texts in folders, and name and classify them according to rules, such as organizing them according to legal departments and promulgation times, and preprocess the collected legal and regulatory texts, including text cleaning, encoding format unification and sentence processing. Select the BERT-CRF model based on deep learning for named entity recognition, load the pre-trained BERT-CRF model, and prepare a legal and regulatory text dataset for fine-tuning. Divide the dataset into training set, validation set and test set. Have professional legal personnel annotate the legal and regulatory texts in the training set and validation set, annotate the legal provision entities and legal concept entities, input the annotated training data into the BERT-CRF model, and fine-tune the model parameters through the back-propagation algorithm to adapt it to the named entity recognition task of legal and regulatory texts. During the fine-tuning process, use the validation set to evaluate the performance of the model, adjust the model's hyperparameters based on the evaluation results to improve the accuracy of the model, and use the test set to test the fine-tuned model. The model is evaluated and its accuracy, recall rate and F1 value are calculated. When the model performance meets the requirements, it is applied to all collected legal and regulatory texts to identify the legal provisions and legal concept entities. Based on the results of named entity recognition, a training dataset for relationship extraction is constructed to extract the relationship between legal provisions and legal concepts from the annotated text and mark the type of relationship. The dataset is also divided into training set, validation set and test set. A convolutional neural network is selected for relationship extraction. The training dataset is input into the model and trained through supervised learning to enable the model to learn the relationship pattern between legal provisions and legal concepts. During the training process, the validation set is used to evaluate and tune the model. The trained relationship extraction model is applied to the text after named entity recognition to identify the specific relationship between legal provisions and legal concepts. For example, the relationship between "Article 8 of the Consumer Protection Law of the People's Republic of China stipulates the consumer's right to know" and "the consumer's right to know" is determined.
[0065] Collect judicial case texts from the judicial judgment document network and the official website of the court, classify and number the case texts, establish a case database, use text classification algorithms to classify case texts, determine the case type, and assume that the text D consists of feature words ω1, ω2, ..., ω n Composition, for each case type C j , according to Bayesian theorem, the text D belongs to type C j The probability P(C j |D), the specific calculation formula is as follows:
[0066]
[0067] Among them, P(C j ) indicates case type C j The probability of P(D|C j ) indicates that the known case type is C j The probability of text D appearing under the condition of , to determine the type of case.
[0068] Intelligent case filing module: This module collects the rules and legal provisions for filing various rights protection cases, categorizes and organizes them for automatic case review, combines historical data to train models to predict the likelihood of winning and risks, and integrates case information, rule judgments, and prediction results to provide a basis for generating case filing strategies.
[0069] In this embodiment, it is necessary to explain the intelligent case filing module. The intelligent case filing module collects case filing rules and legal provisions related to various rights protection cases, and automatically reviews case information through rule judgment. When the rules are met, the case filing conditions are met. When the rules are not met, the case filing conditions are not met. The collected content is classified and sorted according to the case type and the scope of application of the rules. When making rule judgments, if the rules need to meet multiple conditions at the same time, in order to achieve accurate logical judgment, a double-condition compound judgment formula is used. The double-condition compound judgment formula is: meet the condition = A∧B. Further extended to the case of multiple conditions (A, B, ..., N), the following formula is used for judgment:
[0070]
[0071] Among them, C i The Boolean value of the i-th condition, f d Indicates that the condition is met. The value of A that meets the condition is 1, and the value of B that meets the condition is 0. When there is a need for fuzzy matching between case information and rule conditions, cosine similarity is used to calculate text similarity. Suppose the vector obtained by quantizing the rule description text is The vector of the case facts after the same quantization process is The similarity calculation formula is as follows:
[0072]
[0073] Among them, Similarity represents similarity, r i Represents a rule description text vector The i-th component of i Text vector representing case facts The i-th component of Represents a rule description text vector The norm of Text vector representing case facts norm, set a threshold θ, and when the calculated similarity ≥ θ, it is determined that the case facts and rule conditions are successfully matched. Historical rights protection case data are collected, and the trained model is used to predict the possibility of winning and the degree of risk of the case. Based on the prediction results, suggestions are provided for the formulation of the case filing strategy. The case information is integrated with the results of the rule judgment and the prediction results of the machine learning model to provide basic data for the generation of the case filing strategy.
[0074] It should be noted that when making rule judgments, the following assumptions are introduced: condition A indicates whether the party is a Chinese citizen, and condition B indicates whether the case occurred within the territory of China. The compound judgment formula for this dual-condition combination is: Condition A = A∧B. Further expanding to the case of multiple conditions (A, B, ..., N), the following formula is used for judgment:
[0075]
[0076] Through this logical AND operation, it is possible to comprehensively and accurately determine whether the case meets all the set conditions.
[0077] Batch filing application generation module: integrates case information that meets the filing conditions, automatically fills in the document template to generate the filing application document, and after content verification, submits the online filing application to the court filing system through the API interface.
[0078] In this embodiment, what needs to be specifically explained is the batch filing application generation module, which integrates case information that meets the filing conditions, automatically fills in the document template based on the case information, generates a filing application document, verifies the content of the generated document, and transmits the generated application data to the court's filing system through the API interface to realize online filing application submission.
[0079] It should be noted that the model for predicting the possibility of winning and the degree of risk collects a large amount of historical rights protection case data, including basic information of the case (such as information of the parties involved, case type, filing time, etc.), description of case facts, evidence materials, judgment results, etc., and pre-processes the collected data, including data cleaning (removing noise and erroneous data), data conversion (such as converting text data into numerical features), data normalization, etc., to improve the quality and availability of the data, and extract relevant features from the pre-processed data. These features should be able to reflect the key information of the case and the factors affecting the possibility of winning and the degree of risk. For example, for intellectual property infringement cases, the features may include the nature of the infringement, the length of the infringement, the degree of loss of the infringed party, etc. The feature selection algorithm is used to select the most representative and important features, reduce the feature dimension, and improve the model. Training efficiency and performance: According to the characteristics of case data and the requirements of prediction tasks, select appropriate machine learning algorithms, divide the preprocessed data into training set, validation set and test set, use the training set to train the model, adjust the model parameters to make the model achieve optimal performance on the validation set, use the test set to evaluate the trained model, calculate the model's performance indicators, such as accuracy, recall rate, F1 value, mean square error, etc., analyze the advantages and disadvantages of the model based on the evaluation results, optimize the model, apply the optimized model to new rights protection case data, predict the possibility of winning the case and the degree of risk, and provide suggestions for the formulation of the case filing strategy based on the prediction results. For example, if the predicted possibility of winning is high and the risk is low, the case filing can be actively promoted. If the predicted possibility of winning is low or the risk is high, further evaluation and adjustment of the case filing strategy are needed.
[0080] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0081] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0082] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0083] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0084] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0085] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0086] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. An intelligent case filing system for batch rights protection cases based on multi-source data analysis, characterized by: include: Data acquisition and analysis module: collects rights protection case data from multiple sources, classifies the text using natural language processing technology, annotates key information to build a data set, and trains a model using a convolutional neural network to extract relevant information from the text; Case information integration module: Assigns unique identifiers to entities and associates data to form case profiles, builds knowledge graphs based on laws, regulations, and judicial cases, and determines case types by calculating probabilities using Bayesian theorem; Intelligent case filing module: This module collects the rules and legal provisions for filing various rights protection cases, categorizes and organizes them for automatic case review, combines historical data to train models to predict the likelihood of winning and risks, and integrates case information, rule judgments, and prediction results to provide a basis for generating case filing strategies. Batch filing application generation module: integrates case information that meets the filing conditions, automatically fills in the document template to generate the filing application document, and after content verification, submits the online filing application to the court filing system through the API interface.
2. The intelligent filing system for batch rights protection cases based on multi-source data analysis according to claim 1 is characterized in that: The data acquisition and analysis module collects data related to rights protection cases from multiple sources, converts audio and video data into text data through audio-video conversion, performs preliminary cleaning on the collected data, removes noise from the text data collected from multiple channels and unifies the text encoding format, uses word segmentation tools to divide the text into individual words, uses natural language processing technology to classify the text, defines the categories of text classification according to the type and needs of the rights protection case, extracts features from the segmented text by calculating the word frequency in the text, annotates the names of the parties, company names, locations and dates in the text, constructs a training data set, and uses the annotated training data to train the model through a convolutional neural network model.
3. The intelligent filing system for batch rights protection cases based on multi-source data analysis according to claim 2 is characterized in that: The specific steps of model training are as follows: S1. Divide the labeled data into training set, validation set and test set, and divide the data proportionally; S2. Determine the length of the word frequency vector and pass the word frequency vector as input to the convolutional neural network; S3. Perform a convolution operation on the input word frequency vector, sliding the convolution kernel over the input data to identify keywords related to rights protection cases. S4, downsampling the output of the convolutional layer through the pooling layer, reducing the dimension of the data and retaining important feature information, thereby highlighting the important features in the text; S5. Flatten the output of the pooling layer into a one-dimensional vector and connect it to the fully connected layer. The fully connected layer combines and transforms the identified features and learns the complex relationships between the features. S6, the output layer uses the softmax activation function to output the probability distribution of each labeled category; S7. Use the cross-entropy loss function to measure the difference between the model's prediction results and the true annotations. Use stochastic gradient descent to update the model's parameters according to the gradient of the loss function, so that the model's prediction results gradually approach the true annotations. Define evaluation metrics to help evaluate the performance of the model during training. S8. Input the training set data into the convolutional neural network model and perform multiple iterations of training. Each iteration is called an epoch. In each epoch, the model performs forward propagation and backward propagation on all samples in the training set to update the model parameters. After each epoch, the model is evaluated using the validation set data to monitor the performance changes of the model. Based on the evaluation results of the validation set, the model hyperparameters can be adjusted. S9. Evaluate the trained model based on the test set data, calculating the model's accuracy, recall, and F1 value on the test set. Deploy the trained convolutional neural network model to the actual system. During system operation, input text data from information rights protection cases into the model. The model can automatically extract the names of the parties, company names, locations, and dates from the text to support the case filing process. The extracted information is standardized and stored in a relational database. An index is created by case ID, and the name, company, location, and date field tables are associated. Business information and administrative division change data are synchronized regularly, and the company name and location fields are updated.
4. The intelligent filing system for batch rights protection cases based on multi-source data analysis according to claim 1 is characterized in that: The case information integration module assigns a unique identifier to each entity. For the parties, their ID numbers are used as the unique identifiers, and for the enterprises, their social credit codes are used as the unique identifiers. A unique case number is generated for each case as the unique identifier. The unique identifiers of the parties are used as a bridge to establish an association between the party information table and the case information table. In the case information table, a field is set to store the unique identifiers of the parties. The parties are matched with the cases through this field. The associated data are integrated into a unified case information model to form a case portrait. A knowledge graph of rights protection cases is constructed based on laws, regulations and judicial cases. The texts of laws and regulations are collected and named and classified according to the rules. The BERT-CRF model based on deep learning is selected for named entity recognition. Professional legal personnel annotate the legal texts in the training set and the validation set, annotate the legal provision entities and legal concept entities, and input the annotated training data into the BERT-CRF model. The parameters of the model are fine-tuned through the back-propagation algorithm to make it suitable for the named entity recognition task of legal texts. Judicial case texts are collected from the judicial judgment document network and the official website of the court, and the case texts are classified and numbered to establish a case database.
5. The intelligent filing system for batch rights protection cases based on multi-source data analysis according to claim 4 is characterized in that: The case text is classified to determine the case type. Suppose the text D consists of feature words ω1, ω2, ..., ω n Composition, for each case type C j , according to Bayesian theorem, the text D belongs to type C j The probability P(C j |D), the specific calculation formula is as follows: Among them, P(C j ) indicates case type C j The probability of P(D|C j ) indicates that the known case type is C j The probability of text D appearing under the condition of , to determine the type of case.
6. The intelligent filing system for batch rights protection cases based on multi-source data analysis according to claim 1 is characterized in that: The intelligent case filing module collects case filing rules and legal provisions related to various rights protection cases, and automatically reviews case information through rule judgment. When the rules are met, the case filing conditions are met. When the rules are not met, the case filing conditions are not met.
7. The intelligent filing system for batch rights protection cases based on multi-source data analysis according to claim 6 is characterized in that: The judgment formula of the rule judgment is: meet the condition = A∧B, further extended to the case of multiple conditions (A, B, ..., N), the following formula is used for judgment: Among them, C i The Boolean value of the i-th condition, f d Indicates that the condition is met. The value of A that meets the condition is 1, and the value of B that meets the condition is 0. When there is a need for fuzzy matching between case information and rule conditions, cosine similarity is used to calculate text similarity. Suppose the vector obtained by quantizing the rule description text is The vector of the case facts after the same quantization process is The similarity calculation formula is as follows: Among them, Similarity represents similarity, r i Represents a rule description text vector The i-th component of i Text vector representing case facts The i-th component of Represents a rule description text vector The norm of Text vector representing case facts norm, set a threshold θ, and when the calculated similarity ≥ θ, it is determined that the case facts and rule conditions are successfully matched. Historical rights protection case data are collected, and the trained model is used to predict the possibility of winning and the degree of risk of the case. Based on the prediction results, suggestions are provided for the formulation of the case filing strategy. The case information is integrated with the results of the rule judgment and the prediction results of the machine learning model to provide basic data for the generation of the case filing strategy.
8. The intelligent filing system for batch rights protection cases based on multi-source data analysis according to claim 1 is characterized in that: The batch filing application generation module integrates case information that meets the filing conditions, automatically fills in the document template based on the case information, generates the filing application document, verifies the content of the generated document, and transmits the generated application data to the court's filing system through the API interface to realize online filing application submission.
9. A method for intelligent filing of batch rights protection cases based on multi-source data analysis is applied to a system for intelligent filing of batch rights protection cases based on multi-source data analysis as described in any one of claims 1 to 8, characterized in that: The following steps are involved: S101. Collect rights protection case data from multiple sources, classify the text using natural language processing technology, annotate key information to build a data set, and train a model using a convolutional neural network to extract relevant information from the text; S102. Assign unique identifiers to entities and associate data to form case profiles. Build a knowledge graph based on laws, regulations, and judicial cases. Calculate probabilities using Bayesian theorem to determine case types. S103. Collect the rules and legal provisions for filing various rights protection cases, classify and organize them for automatic case review, combine historical data to train models to predict the likelihood of winning and risks, and integrate case information, rule judgments, and prediction results to provide a basis for generating case filing strategies; S104. Integrate case information that meets the filing conditions, automatically fill in the document template to generate the filing application document, and after content verification, submit the online filing application to the court filing system through the API interface.