Machine learning-based civil and commercial case classification method and system
By adopting the dual-interactive self-attention network model and legal common sense reasoning in the classification of civil and commercial cases and combining optimization algorithms for hyperparameter optimization, the problem of traditional methods being difficult to understand case logic and lacking global search capabilities is solved, and a more efficient and accurate case classification is achieved.
Patent Information
- Application Number
- CN202510069104.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional civil and commercial case classification method is difficult to deeply explore the relationship between complex logic and legal provisions in the case text, lacks an understanding of the legal reasoning behind the case, and is difficult to introduce global search capabilities and diversity, affecting the classification effect.
A dual-interactive self-attention network model is used to classify civil and commercial cases, and a comprehensive legal knowledge reasoning is generated to deeply understand the legal background and article details in the case, and the model hyperparameter optimization is carried out through optimization algorithms. Hyperbolic tangent chaos mapping and Cauchy mutation strategy are used to improve global search capabilities and diversity.
It effectively improves the accuracy and efficiency of case classification, can better understand the legal logic and article relationship in the case, improves the classification effect, and avoids premature convergence in the search process.
Smart Images

Figure CN119939352A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of civil and commercial case data processing, and specifically to a civil and commercial case classification method and system based on machine learning. Background Art
[0002] Intelligent civil and commercial case classification is the process of using artificial intelligence technology to automatically classify and analyze civil and commercial cases. It improves the efficiency and accuracy of case classification through technologies such as machine learning and natural language processing. The application of this technology can greatly improve judicial efficiency, optimize the allocation of legal resources, and help legal workers quickly locate the focus of case disputes. At the same time, it promotes judicial transparency and standardization and reduces human errors.
[0003] However, the traditional civil and commercial case classification method is limited to the processing of surface information of the case text, and it is difficult to deeply explore the complex logic in the case text and the relationship between the legal provisions, and lacks the understanding of the legal reasoning behind the case. The traditional civil and commercial case classification method has the technical problem of difficulty in introducing global search capabilities and diversity in the optimization process, and it is impossible to obtain the ideal hyperparameter combination when facing complex and nonlinear problems, which affects the final classification effect. Summary of the invention
[0004] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a civil and commercial case classification method and system based on machine learning. In view of the technical problems that the traditional civil and commercial case classification method is limited to the processing of surface information of the case text, it is difficult to deeply explore the complex logic in the case text and the relationship between the legal provisions, and lacks the understanding of the legal reasoning behind the case, this scheme creatively adopts a dual interactive self-attention network model to classify civil and commercial cases. By combining the generated legal common sense reasoning, it deeply understands the legal background and provision details in the case, supplements the limitations of relying solely on the case text, and effectively improves the classification effect; in view of the technical problems that the traditional civil and commercial case classification method is difficult to introduce global search capabilities and diversity in the optimization process, and cannot obtain ideal hyperparameter combinations when facing complex and nonlinear problems, which affects the final classification effect, this scheme creatively adopts an optimization algorithm to optimize the model hyperparameters, combines hyperbolic tangent chaos mapping and Cauchy mutation strategy, avoids premature convergence in the search process, improves the global search capability, and enhances the diversity of the search process.
[0005] The technical solution adopted by the present invention is as follows: The civil and commercial case classification method based on machine learning provided by the present invention comprises the following steps:
[0006] Step S1: data collection;
[0007] Step S2: data preprocessing;
[0008] Step S3: Case classification model construction;
[0009] Step S4: hyperparameter optimization;
[0010] Step S5: Classification of civil and commercial cases.
[0011] Furthermore, in step S1, the data collection is used to collect data required for classifying civil and commercial cases, specifically, to obtain an original case data set through data collection;
[0012] The original case data set specifically includes a historical case data set and a current case data set. The historical case data set specifically includes historical case text data, historical legal document data and historical case label data. The current case data set specifically includes current case text data and current legal document data. The case text data specifically includes judgment data and trial record data. The legal document data specifically includes relevant legal provisions data. The historical case label data specifically includes label data based on the types of civil and commercial cases.
[0013] Furthermore, in step S2, the data preprocessing is used to preprocess the collected original case data, and specifically includes the following steps:
[0014] Step S21: data cleaning, for cleaning the original data, specifically removing abnormal values, duplicate values, special symbols, extra spaces and stop words in the historical case data set and the current case data set, to obtain a historical preliminary processing data set and a current preliminary processing data set;
[0015] Step S22: segmentation, which is used to segment the preliminary processed data into text segments, specifically segmenting the text data in the historical preliminary processed data set and the current preliminary processed data set into segments in units of sentences to obtain a historical data set and a case set to be classified;
[0016] Step S23: data set segmentation, which is used to segment the data set, specifically, segmenting the historical data set to obtain a classification training set and a classification test set;
[0017] Step S24: performing preprocessing, specifically, preprocessing the current case data set through the data cleaning and the paragraph segmentation to obtain a case set to be classified, and preprocessing the historical case data set through the data cleaning, the paragraph segmentation and the data set segmentation to obtain a classification training set and a classification test set.
[0018] Further, in step S3, the case classification model is constructed to construct a model required for classifying civil and commercial cases, specifically by constructing a dual interactive self-attention network model as a civil and commercial case classification model, and the dual interactive self-attention network model specifically includes a word embedding module, a feature extraction module, a self-attention module, an interactive attention module and an output module;
[0019] The case classification model construction specifically includes the following steps:
[0020] Step S31: word embedding module construction, which is used to construct a word embedding module, specifically, combining the word frequency-inverse word frequency method to construct a word embedding module, and the steps include:
[0021] Step S311: Calculate the word frequency-inverse word frequency weight, using the following formula:
[0022] ;
[0023] In the formula, represents the word frequency-inverse word frequency weight, Indicates the number of times word c appears in text b, Indicates the number of times word a appears in text b. Indicates the number of word types in text b, represents the number of times word d appears in the text set, represents the number of times word c appears in the text set, and D represents the number of word types in the text set;
[0024] Step S312: Obtain a word embedding vector, specifically obtaining the word embedding vector of the input data through the GloVe model, and the formula used is as follows:
[0025] ;
[0026] In the formula, represents the input data word embedding vector, Represents the GloVe model processing function, Int represents the input data;
[0027] Step S313: Obtaining the embedding vector of legal common sense reasoning words, specifically, using the legal document data as the input of the common sense transformer model, and further obtaining the embedding vector of legal common sense reasoning words, the formula used is as follows:
[0028] ;
[0029] In the formula, represents the embedding vector of legal common sense reasoning words, represents the common sense transformer processing function, Indicates legal document input data;
[0030] Step S314: Calculate the weighted word embedding vector, which is used to calculate the weighted word embedding vector output by the word embedding module. The formula used is as follows:
[0031] ;
[0032] In the formula, represents the weighted word embedding vector, represents the normalized word frequency-inverse word frequency weight;
[0033] Step S32: constructing a feature extraction module, specifically constructing a feature extraction module based on a gated recurrent unit, the steps comprising:
[0034] Step S321: feature extraction, the formula used is as follows:
[0035] ;
[0036] In the formula, Represents the output features of the feature extraction module corresponding to the input data, It represents the output features of the feature extraction module corresponding to the legal document input data. Represents the gated recurrent unit processing function;
[0037] Step S322: constructing a position embedding vector, specifically constructing a position embedding vector based on the input data and the legal document input data, and the formula used is as follows:
[0038] ;
[0039] In the formula, represents the input data position embedding vector, Represents the embedding vector of the legal document input data position, represents the position embedding function;
[0040] Step S323: construct fusion features, the formula used is as follows:
[0041] ;
[0042] In the formula, represents the fusion features corresponding to the input data, Indicates the fusion features corresponding to the legal document input data, Indicates a connection operation;
[0043] Step S33: constructing a self-attention module, specifically constructing a self-attention module based on a multi-head self-attention mechanism, the formula used is as follows:
[0044] ;
[0045] In the formula, represents the query of the hth head, represents the key of the hth head, Represents the value of the hth head, represents the query transformation matrix of the hth head, represents the key transformation matrix of the hth head, represents the value transformation matrix of the hth head, represents the self-attention output feature of the h-th head, represents the softmax function, represents the dimension of the key of the hth head, Represents the connection operation function, represents the self-attention output feature of the first head, represents the self-attention output feature of the second head, represents the self-attention linear transformation weight matrix, represents the multi-head self-attention output feature, and T represents the transposition operation;
[0046] Step S34: constructing an interactive attention module, specifically constructing an interactive attention module for interaction between features, the steps comprising:
[0047] Step S341: construct an interaction matrix, the formula used is as follows:
[0048] ;
[0049] Where Im represents the interaction matrix;
[0050] Step S342: Obtain the self-interaction feature, the formula used is as follows:
[0051] ;
[0052] Where Si represents the self-interaction score, represents the hyperbolic tangent function, represents the learnable matrix used to calculate the self-interaction score, Sia represents the self-interaction attention weight, represents the learnable matrix used to calculate the self-interaction attention weights, represents the self-interaction feature;
[0053] Step S343: Obtain common sense interaction features, the formula used is as follows:
[0054] ;
[0055] Where Ci represents the common sense interaction score, represents the learnable matrix used to calculate the common sense interaction score, Cia represents the common sense interaction attention weight, represents a learnable matrix for computing commonsense interaction attention weights, Represents common sense interaction features;
[0056] Step S35: constructing an output module, specifically constructing an output module for fusing interactive features and obtaining model output. The formula used is as follows:
[0057] ;
[0058] In the formula, Represents the classification result output by the model, represents the model output weight, represents the fully connected layer function, which is used to map features to the output space. Represents the model output bias term;
[0059] Step S36: Model construction and training, specifically, constructing a dual interactive self-attention network model through the word embedding module construction, the feature extraction module construction, the self-attention module construction, the interactive attention module construction and the output module construction, and training the model based on the classification training set, verifying the model performance based on the classification test set, selecting the cross entropy loss function as the model loss function, obtaining the dual interactive self-attention network model, and using it as a civil and commercial case classification model.
[0060] Further, in step S4, the hyperparameter optimization is specifically to optimize the hyperparameters of the civil and commercial case classification model using an optimization algorithm to obtain an optimized civil and commercial case classification model;
[0061] The hyperparameter optimization specifically includes the following steps:
[0062] Step S41: Generate an initial solution, specifically generate an initial solution of the optimization algorithm based on chaos mapping, the formula used is as follows:
[0063] ;
[0064] In the formula, represents the initial value of the chaotic map, r represents the chaotic control parameter, represents the nonlinear degree adjustment parameter, represents the chaotic value of the j+1th dimension, represents the chaos value of the jth dimension, represents the position of the i-th individual unit in the j-th dimension, represents the upper limit of the j-th dimension, represents the lower limit of the j-th dimension, and the individual unit is used to represent the hyperparameter combination of the model to be optimized;
[0065] Step S42: determining a fitness value, specifically selecting the classification accuracy of the civil and commercial case classification model as the fitness value of the individual unit;
[0066] Step S43: Individual unit position update, the steps include:
[0067] Step S431: Update the evolution factor, the formula used is as follows:
[0068] ;
[0069] In the formula, represents the evolution factor at the dtth iteration, represents a random number in the range [-1,1], dt represents the current number of iterations, Indicates the maximum number of iterations;
[0070] Step S432: Calculate the percentage difference between the current position and the optimal position, using the following formula:
[0071] ;
[0072] In the formula, Indicates the percentage difference between the current position and the best position. represents the dimension mean of the i-th individual unit, J represents the total number of individual unit dimensions, represents the position of the jth dimension of the best position, represents the development control factor;
[0073] Step S433: Calculate the reduction factor using the following formula:
[0074] ;
[0075] In the formula, represents the reduction factor of the position of the i-th individual unit in the j-th dimension, represents the position of a random individual unit in the jth dimension;
[0076] Step S434: Location update, the formula used is as follows:
[0077] ;
[0078] In the formula, represents the position of the i-th individual unit in the j-th dimension at the dt+1-th iteration, represents the exploration control factor, represents a random number in the range [0,1], represents a random number in the range [0,1], represents a random number in the range [0,1], Represents a random number in the range [0,1];
[0079] Step S435: position mutation, the formula used is as follows:
[0080] ;
[0081] In the formula, represents the position of the i-th individual unit in the j-th dimension at the dt+1-th iteration after the mutation, and Cr represents a random number that obeys the standard Cauchy distribution;
[0082] Step S44: Iterative update, specifically iterative update of the optimization algorithm until the iterative termination condition of the optimization algorithm is reached, so as to obtain the final global optimal position, and then obtain the optimized civil and commercial case classification model. The final global optimal position is specifically the optimal model hyperparameter combination, and the iterative termination condition of the optimization algorithm is specifically reaching the maximum number of iterations or the individual unit fitness function value is greater than the preset threshold.
[0083] Furthermore, in step S5, the civil and commercial case classification is specifically to process the case set to be classified using the optimized civil and commercial case classification model to obtain reference data on the types of civil and commercial cases, and classify the civil and commercial cases based on the reference data on the types of civil and commercial cases.
[0084] The civil and commercial case classification system based on machine learning provided by the present invention includes a data acquisition module, a data preprocessing module, a case classification model building module, a hyperparameter optimization module and a civil and commercial case classification module;
[0085] The data acquisition module is used for data acquisition, obtains the original case data set through data acquisition, and sends the original case data set to the data preprocessing module;
[0086] The data preprocessing module is used for data preprocessing. Through data preprocessing, a case set to be classified, a classification training set and a classification test set are obtained, and the case set to be classified is sent to the civil and commercial case classification module, and the classification training set and the classification test set are sent to the case classification model construction module;
[0087] The case classification model construction module is used for case classification model construction, and obtains a civil and commercial case classification model by constructing a dual interactive self-attention network model, and sends the civil and commercial case classification model to a hyperparameter optimization module;
[0088] The hyperparameter optimization module is used for hyperparameter optimization, and optimizes the model hyperparameters by using an optimization algorithm to obtain an optimized civil and commercial case classification model, and sends the optimized civil and commercial case classification model to the civil and commercial case classification module;
[0089] The civil and commercial case classification module is used for classifying civil and commercial cases. Civil and commercial cases are classified by adopting the optimized civil and commercial case classification model to obtain reference data on the types of civil and commercial cases.
[0090] The beneficial effects achieved by the present invention using the above scheme are as follows:
[0091] (1) In view of the technical problems that traditional civil and commercial case classification methods are limited to processing surface information in case texts, it is difficult to deeply explore the complex logic in case texts and the relationship between legal provisions, and lack an understanding of the legal reasoning behind the cases, this scheme creatively adopts a dual interactive self-attention network model to classify civil and commercial cases. By combining the generated legal common sense reasoning, it deeply understands the legal background and provision details in the case, supplements the limitations of relying solely on case texts, and effectively improves the classification effect.
[0092] (2) In view of the technical problems that traditional civil and commercial case classification methods have difficulty in introducing global search capabilities and diversity in the optimization process, and cannot obtain ideal hyperparameter combinations when facing complex and nonlinear problems, which affects the final classification effect, this solution creatively adopts the optimization algorithm to optimize the model hyperparameters, combined with the hyperbolic tangent chaos mapping and Cauchy mutation strategy, to avoid premature convergence in the search process, improve the global search capability, and enhance the diversity of the search process. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 A schematic diagram of the process flow of the civil and commercial case classification method based on machine learning provided by the present invention;
[0094] Figure 2 A schematic diagram of the modules of the civil and commercial case classification system based on machine learning provided by the present invention;
[0095] Figure 3 This is a schematic diagram of the process of data preprocessing in step S2;
[0096] Figure 4 Schematic diagram of the process of building a case classification model for step S3;
[0097] Figure 5 Schematic diagram of the process of hyperparameter optimization in step S4.
[0098] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0099] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0100] In the description of the present invention, it should be understood that terms such as “upper”, “lower”, “front”, “back”, “left”, “right”, “top”, “bottom”, “inside” and “outside” indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as limiting the present invention.
[0101] Example 1, see Figure 1 The technical solution adopted by the present invention is as follows: The civil and commercial case classification method based on machine learning provided by the present invention comprises the following steps:
[0102] Step S1: data collection;
[0103] Step S2: data preprocessing;
[0104] Step S3: Case classification model construction;
[0105] Step S4: hyperparameter optimization;
[0106] Step S5: Classification of civil and commercial cases.
[0107] Example 2, see Figure 1 and Figure 2 In step S1, the data collection is used to collect the data required for classifying civil and commercial cases, specifically, to obtain the original case data set through data collection;
[0108] The original case data set specifically includes a historical case data set and a current case data set. The historical case data set specifically includes historical case text data, historical legal document data and historical case label data. The current case data set specifically includes current case text data and current legal document data. The case text data specifically includes judgment data and trial record data. The legal document data specifically includes relevant legal provisions data. The historical case label data specifically includes label data based on the types of civil and commercial cases.
[0109] Example 3, see Figure 1 , Figure 2and Figure 3 This embodiment is based on the above embodiment. In step S2, the data preprocessing is used to preprocess the collected original case data, and specifically includes the following steps:
[0110] Step S21: data cleaning, for cleaning the original data, specifically removing abnormal values, duplicate values, special symbols, extra spaces and stop words in the historical case data set and the current case data set, to obtain a historical preliminary processing data set and a current preliminary processing data set;
[0111] Step S22: segmentation, which is used to segment the preliminary processed data into text segments, specifically segmenting the text data in the historical preliminary processed data set and the current preliminary processed data set into segments in units of sentences to obtain a historical data set and a case set to be classified;
[0112] Step S23: data set segmentation, which is used to segment the data set, specifically, segmenting the historical data set to obtain a classification training set and a classification test set;
[0113] Step S24: performing preprocessing, specifically, preprocessing the current case data set through the data cleaning and the paragraph segmentation to obtain a case set to be classified, and preprocessing the historical case data set through the data cleaning, the paragraph segmentation and the data set segmentation to obtain a classification training set and a classification test set.
[0114] Example 4, see Figure 1 , Figure 2 and Figure 4 This embodiment is based on the above embodiment. In step S3, the case classification model is constructed to construct a model required for classifying civil and commercial cases. Specifically, a dual interactive self-attention network model is constructed as a civil and commercial case classification model. The dual interactive self-attention network model specifically includes a word embedding module, a feature extraction module, a self-attention module, an interactive attention module and an output module.
[0115] The case classification model construction specifically includes the following steps:
[0116] Step S31: word embedding module construction, which is used to construct a word embedding module, specifically, combining the word frequency-inverse word frequency method to construct a word embedding module, and the steps include:
[0117] Step S311: Calculate the word frequency-inverse word frequency weight, using the following formula:
[0118] ;
[0119] In the formula, represents the word frequency-inverse word frequency weight, Indicates the number of times word c appears in text b, Indicates the number of times word a appears in text b. Indicates the number of word types in text b, represents the number of times word d appears in the text set, represents the number of times word c appears in the text set, and D represents the number of word types in the text set;
[0120] Step S312: Obtain a word embedding vector, specifically obtaining the word embedding vector of the input data through the GloVe model, and the formula used is as follows:
[0121] ;
[0122] In the formula, represents the input data word embedding vector, Represents the GloVe model processing function, Int represents the input data;
[0123] Step S313: Obtaining the embedding vector of legal common sense reasoning words, specifically, using the legal document data as the input of the common sense transformer model, and further obtaining the embedding vector of legal common sense reasoning words, the formula used is as follows:
[0124] ;
[0125] In the formula, represents the embedding vector of legal common sense reasoning words, represents the common sense transformer processing function, Indicates legal document input data;
[0126] Step S314: Calculate the weighted word embedding vector, which is used to calculate the weighted word embedding vector output by the word embedding module. The formula used is as follows:
[0127] ;
[0128] In the formula, represents the weighted word embedding vector, represents the normalized word frequency-inverse word frequency weight;
[0129] Step S32: constructing a feature extraction module, specifically constructing a feature extraction module based on a gated recurrent unit, the steps comprising:
[0130] Step S321: feature extraction, the formula used is as follows:
[0131] ;
[0132] In the formula, Represents the output features of the feature extraction module corresponding to the input data, It represents the output features of the feature extraction module corresponding to the legal document input data. Represents the gated recurrent unit processing function;
[0133] Step S322: constructing a position embedding vector, specifically constructing a position embedding vector based on the input data and the legal document input data, and the formula used is as follows:
[0134] ;
[0135] In the formula, represents the input data position embedding vector, Represents the position embedding vector of the legal document input data, represents the position embedding function;
[0136] Step S323: construct fusion features, the formula used is as follows:
[0137] ;
[0138] In the formula, represents the fusion features corresponding to the input data, Indicates the fusion features corresponding to the legal document input data, Indicates a connection operation;
[0139] Step S33: constructing a self-attention module, specifically constructing a self-attention module based on a multi-head self-attention mechanism, the formula used is as follows:
[0140] ;
[0141] In the formula, represents the query of the hth head, represents the key of the hth head, Represents the value of the hth head, represents the query transformation matrix of the hth head, represents the key transformation matrix of the hth head, represents the value transformation matrix of the hth head, represents the self-attention output feature of the h-th head, represents the softmax function, represents the dimension of the key of the hth head, Represents the connection operation function, represents the self-attention output feature of the first head, represents the self-attention output feature of the second head, represents the self-attention linear transformation weight matrix, represents the multi-head self-attention output feature, and T represents the transposition operation;
[0142] Step S34: constructing an interactive attention module, specifically constructing an interactive attention module for interaction between features, the steps comprising:
[0143] Step S341: construct an interaction matrix, the formula used is as follows:
[0144] ;
[0145] Where Im represents the interaction matrix;
[0146] Step S342: Obtain the self-interaction feature, the formula used is as follows:
[0147] ;
[0148] Where Si represents the self-interaction score, represents the hyperbolic tangent function, represents the learnable matrix used to calculate the self-interaction score, Sia represents the self-interaction attention weight, represents the learnable matrix used to calculate the self-interaction attention weights, represents the self-interaction feature;
[0149] Step S343: Obtain common sense interaction features, the formula used is as follows:
[0150] ;
[0151] Where Ci represents the common sense interaction score, represents the learnable matrix used to calculate the common sense interaction score, Cia represents the common sense interaction attention weight, represents a learnable matrix for computing commonsense interaction attention weights, Represents common sense interaction features;
[0152] Step S35: constructing an output module, specifically constructing an output module for fusing interactive features and obtaining model output. The formula used is as follows:
[0153] ;
[0154] In the formula, Represents the classification result output by the model, represents the model output weight, represents the fully connected layer function, which is used to map features to the output space. Represents the model output bias term;
[0155] Step S36: Model construction and training, specifically, constructing a dual interactive self-attention network model through the word embedding module construction, the feature extraction module construction, the self-attention module construction, the interactive attention module construction and the output module construction, and training the model based on the classification training set, verifying the model performance based on the classification test set, selecting the cross entropy loss function as the model loss function, obtaining the dual interactive self-attention network model, and using it as a civil and commercial case classification model.
[0156] By performing the above operations, this solution creatively uses a dual interactive self-attention network model to classify civil and commercial cases. By combining the generated legal common sense reasoning, it deeply understands the legal background and clause details in the case, supplements the limitations of relying solely on case texts, and effectively improves the classification effect.
[0157] Example 5, see Figure 1 , Figure 2 and Figure 5 This embodiment is based on the above embodiment. In step S4, the hyperparameter optimization is specifically to optimize the hyperparameters of the civil and commercial case classification model using an optimization algorithm to obtain an optimized civil and commercial case classification model;
[0158] The hyperparameter optimization specifically includes the following steps:
[0159] Step S41: Generate an initial solution, specifically generate an initial solution of the optimization algorithm based on chaos mapping, the formula used is as follows:
[0160] ;
[0161] In the formula, represents the initial value of the chaotic map, r represents the chaotic control parameter, represents the nonlinear degree adjustment parameter, represents the chaotic value of the j+1th dimension, represents the chaos value of the jth dimension, represents the position of the i-th individual unit in the j-th dimension, represents the upper limit of the j-th dimension, represents the lower limit of the j-th dimension, and the individual unit is used to represent the hyperparameter combination of the model to be optimized;
[0162] Step S42: determining a fitness value, specifically selecting the classification accuracy of the civil and commercial case classification model as the fitness value of the individual unit;
[0163] Step S43: Individual unit position update, the steps include:
[0164] Step S431: Update the evolution factor, the formula used is as follows:
[0165] ;
[0166] In the formula, represents the evolution factor at the dtth iteration, represents a random number in the range [-1,1], dt represents the current number of iterations, Indicates the maximum number of iterations;
[0167] Step S432: Calculate the percentage difference between the current position and the optimal position, using the following formula:
[0168] ;
[0169] In the formula, Indicates the percentage difference between the current position and the best position. represents the dimension mean of the i-th individual unit, J represents the total number of individual unit dimensions, represents the position of the jth dimension of the best position, represents the development control factor;
[0170] Step S433: Calculate the reduction factor using the following formula:
[0171] ;
[0172] In the formula, represents the reduction factor of the position of the i-th individual unit in the j-th dimension, represents the position of a random individual unit in the jth dimension;
[0173] Step S434: Location update, the formula used is as follows:
[0174] ;
[0175] In the formula, represents the position of the i-th individual unit in the j-th dimension at the dt+1-th iteration, represents the exploration control factor, represents a random number in the range [0,1], represents a random number in the range [0,1], represents a random number in the range [0,1], Represents a random number in the range [0,1];
[0176] Step S435: position mutation, the formula used is as follows:
[0177] ;
[0178] In the formula, represents the position of the ith individual unit in the jth dimension at the dt+1th iteration after the mutation, and Cr represents a random number that obeys the standard Cauchy distribution;
[0179] Step S44: Iterative update, specifically iterative update of the optimization algorithm until the iterative termination condition of the optimization algorithm is reached, so as to obtain the final global optimal position, and then obtain the optimized civil and commercial case classification model. The final global optimal position is specifically the optimal model hyperparameter combination, and the iterative termination condition of the optimization algorithm is specifically reaching the maximum number of iterations or the individual unit fitness function value is greater than the preset threshold.
[0180] By performing the above operations, this solution creatively uses optimization algorithms to optimize model hyperparameters, combining hyperbolic tangent chaos mapping and Cauchy mutation strategy to avoid premature convergence in the search process, improve global search capabilities, and enhance diversity in the search process, in order to address the technical problems that traditional civil and commercial case classification methods have difficulty introducing global search capabilities and diversity in the optimization process, and cannot obtain ideal hyperparameter combinations when facing complex and nonlinear problems, which affects the final classification effect.
[0181] Example 6, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. In step S5, the civil and commercial case classification is specifically to use the optimized civil and commercial case classification model to process the case set to be classified, obtain civil and commercial case type reference data, and classify the civil and commercial cases based on the civil and commercial case type reference data.
[0182] Embodiment 7, see Figure 1 and Figure 2 , this embodiment is based on the above embodiment, and the civil and commercial case classification system based on machine learning provided by the present invention includes a data acquisition module, a data preprocessing module, a case classification model construction module, a hyperparameter optimization module and a civil and commercial case classification module;
[0183] The data acquisition module is used for data acquisition, obtains the original case data set through data acquisition, and sends the original case data set to the data preprocessing module;
[0184] The data preprocessing module is used for data preprocessing. Through data preprocessing, a case set to be classified, a classification training set and a classification test set are obtained, and the case set to be classified is sent to the civil and commercial case classification module, and the classification training set and the classification test set are sent to the case classification model construction module;
[0185] The case classification model construction module is used for case classification model construction, and obtains a civil and commercial case classification model by constructing a dual interactive self-attention network model, and sends the civil and commercial case classification model to a hyperparameter optimization module;
[0186] The hyperparameter optimization module is used for hyperparameter optimization, and optimizes the model hyperparameters by using an optimization algorithm to obtain an optimized civil and commercial case classification model, and sends the optimized civil and commercial case classification model to the civil and commercial case classification module;
[0187] The civil and commercial case classification module is used for classifying civil and commercial cases. Civil and commercial cases are classified by adopting the optimized civil and commercial case classification model to obtain reference data on the types of civil and commercial cases.
[0188] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0189] While the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that many changes, modifications, substitutions and variations can be made to the embodiments without departing from the principles and spirit of the invention.
[0190] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. A civil and commercial case classification method based on machine learning, characterized by: The method comprises the following steps: Step S1: data collection, through which an original case data set is obtained, the original case data set specifically includes a historical case data set and a current case data set; Step S2: Data preprocessing, preprocessing the collected raw data to obtain a case set to be classified, a classification training set, and a classification test set; Step S3: Case classification model construction, which is used to construct the model required for classifying civil and commercial cases, specifically, by constructing a dual interactive self-attention network model to obtain a civil and commercial case classification model; Step S4: hyperparameter optimization, which is used to optimize the model hyperparameters, specifically, using an optimization algorithm to optimize the hyperparameters of the civil and commercial case classification model to obtain an optimized civil and commercial case classification model; Step S5: Classification of civil and commercial cases, specifically, classifying civil and commercial cases through the optimized case classification model to obtain reference data on the types of civil and commercial cases.
2. The method for classifying civil and commercial cases based on machine learning according to claim 1 is characterized by: In step S3, the case classification model is constructed to construct a model required for classifying civil and commercial cases, specifically by constructing a dual interactive self-attention network model as a civil and commercial case classification model, wherein the dual interactive self-attention network model specifically includes a word embedding module, a feature extraction module, a self-attention module, an interactive attention module, and an output module; The case classification model construction specifically includes the following steps: Step S31: word embedding module construction, which is used to construct a word embedding module, specifically, combining the word frequency-inverse word frequency method to construct a word embedding module, and the steps include: Step S311: Calculate the word frequency-inverse word frequency weight, using the following formula: ; In the formula, represents the word frequency-inverse word frequency weight, Indicates the number of times word c appears in text b, Indicates the number of times word a appears in text b. Indicates the number of word types in text b, represents the number of times word d appears in the text set, represents the number of times word c appears in the text set, and D represents the number of word types in the text set; Step S312: Obtain a word embedding vector, specifically obtaining the word embedding vector of the input data through the GloVe model, and the formula used is as follows: ; In the formula, represents the input data word embedding vector, Represents the GloVe model processing function, Int represents the input data; Step S313: Obtain the embedding vector of legal common sense reasoning words. Specifically, the legal document data is used as the input of the common sense transformer model, and the embedding vector of legal common sense reasoning words is further obtained. The formula used is as follows: ; In the formula, represents the embedding vector of legal common sense reasoning words, represents the common sense transformer processing function, Indicates legal document input data; Step S314: Calculate the weighted word embedding vector, which is used to calculate the weighted word embedding vector output by the word embedding module. The formula used is as follows: ; In the formula, represents the weighted word embedding vector, represents the normalized word frequency-inverse word frequency weight; Step S32: constructing a feature extraction module, specifically constructing a feature extraction module based on a gated recurrent unit, the steps comprising: Step S321: feature extraction, the formula used is as follows: ; In the formula, Represents the output features of the feature extraction module corresponding to the input data, It represents the output features of the feature extraction module corresponding to the legal document input data. Represents the gated recurrent unit processing function; Step S322: constructing a position embedding vector, specifically constructing a position embedding vector based on the input data and the legal document input data, and the formula used is as follows: ; In the formula, represents the input data position embedding vector, Represents the embedding vector of the legal document input data position, represents the position embedding function; Step S323: construct fusion features, the formula used is as follows: ; In the formula, represents the fusion features corresponding to the input data, Indicates the fusion features corresponding to the legal document input data, Indicates a connection operation; Step S33: constructing a self-attention module, specifically constructing a self-attention module based on a multi-head self-attention mechanism, the formula used is as follows: ; In the formula, represents the query of the hth head, represents the key of the hth head, Represents the value of the hth head, represents the query transformation matrix of the hth head, represents the key transformation matrix of the hth head, represents the value transformation matrix of the hth head, represents the self-attention output feature of the h-th head, represents the softmax function, represents the dimension of the key of the hth head, Represents the connection operation function, represents the self-attention output feature of the first head, represents the self-attention output feature of the second head, represents the self-attention linear transformation weight matrix, represents the multi-head self-attention output feature, and T represents the transposition operation; Step S34: constructing an interactive attention module, specifically constructing an interactive attention module for interaction between features, the steps comprising: Step S341: construct an interaction matrix, the formula used is as follows: ; Where Im represents the interaction matrix; Step S342: Obtain the self-interaction feature, the formula used is as follows: ; Where Si represents the self-interaction score, represents the hyperbolic tangent function, represents the learnable matrix used to calculate the self-interaction score, Sia represents the self-interaction attention weight, represents the learnable matrix used to calculate the self-interaction attention weights, represents the self-interaction feature; Step S343: Obtain common sense interaction features, the formula used is as follows: ; Where Ci represents the common sense interaction score, represents the learnable matrix used to calculate the common sense interaction score, Cia represents the common sense interaction attention weight, represents a learnable matrix for computing commonsense interaction attention weights, Represents common sense interaction features; Step S35: constructing an output module, specifically constructing an output module for fusing interactive features and obtaining model output. The formula used is as follows: ; In the formula, Represents the classification result output by the model, represents the model output weight, represents the fully connected layer function, which is used to map features to the output space. Represents the model output bias term; Step S36: Model construction and training, specifically, constructing a dual interactive self-attention network model through the word embedding module construction, the feature extraction module construction, the self-attention module construction, the interactive attention module construction and the output module construction, and training the model based on the classification training set, verifying the model performance based on the classification test set, selecting the cross entropy loss function as the model loss function, obtaining the dual interactive self-attention network model, and using it as a civil and commercial case classification model.
3. The method for classifying civil and commercial cases based on machine learning according to claim 1 is characterized by: In step S4, the hyperparameter optimization is specifically to optimize the hyperparameters of the civil and commercial case classification model using an optimization algorithm to obtain an optimized civil and commercial case classification model; The hyperparameter optimization specifically includes the following steps: Step S41: Generate an initial solution, specifically generate an initial solution of the optimization algorithm based on chaos mapping, the formula used is as follows: ; In the formula, represents the initial value of the chaotic map, r represents the chaotic control parameter, represents the nonlinear degree adjustment parameter, represents the chaotic value of the j+1th dimension, represents the chaos value of the jth dimension, represents the position of the i-th individual unit in the j-th dimension, represents the upper limit of the j-th dimension, represents the lower limit of the j-th dimension, and the individual unit is used to represent the hyperparameter combination of the model to be optimized; Step S42: determining a fitness value, specifically selecting the classification accuracy of the civil and commercial case classification model as the fitness value of the individual unit; Step S43: Individual unit position update, the steps include: Step S431: Update the evolution factor, the formula used is as follows: ; In the formula, represents the evolution factor at the dtth iteration, represents a random number in the range [-1,1], dt represents the current number of iterations, Indicates the maximum number of iterations; Step S432: Calculate the percentage difference between the current position and the optimal position, using the following formula: ; In the formula, Indicates the percentage difference between the current position and the best position. represents the dimension mean of the i-th individual unit, J represents the total number of individual unit dimensions, represents the position of the jth dimension of the best position, represents the development control factor; Step S433: Calculate the reduction factor using the following formula: ; In the formula, represents the reduction factor of the position of the i-th individual unit in the j-th dimension, represents the position of a random individual unit in the jth dimension; Step S434: Location update, the formula used is as follows: ; In the formula, represents the position of the i-th individual unit in the j-th dimension at the dt+1-th iteration, represents the exploration control factor, represents a random number in the range [0,1], represents a random number in the range [0,1], represents a random number in the range [0,1], Represents a random number in the range [0,1]; Step S435: position mutation, the formula used is as follows: ; In the formula, represents the position of the ith individual unit in the jth dimension at the dt+1th iteration after the mutation, and Cr represents a random number that obeys the standard Cauchy distribution; Step S44: Iterative update, specifically iterative update of the optimization algorithm until the iterative termination condition of the optimization algorithm is reached, so as to obtain the final global optimal position, and then obtain the optimized civil and commercial case classification model. The final global optimal position is specifically the optimal model hyperparameter combination, and the iterative termination condition of the optimization algorithm is specifically reaching the maximum number of iterations or the individual unit fitness function value is greater than the preset threshold.
4. The method for classifying civil and commercial cases based on machine learning according to claim 1 is characterized by: In step S1, the data collection is used to collect the data required for classifying civil and commercial cases, specifically, to obtain the original case data set through data collection; The original case data set specifically includes a historical case data set and a current case data set. The historical case data set specifically includes historical case text data, historical legal document data and historical case label data. The current case data set specifically includes current case text data and current legal document data. The case text data specifically includes judgment data and trial record data. The legal document data specifically includes relevant legal provisions data. The historical case label data specifically includes label data based on the types of civil and commercial cases.
5. The method for classifying civil and commercial cases based on machine learning according to claim 1 is characterized by: In step S2, the data preprocessing is used to preprocess the collected original case data, and specifically includes the following steps: Step S21: data cleaning, for cleaning the original data, specifically removing abnormal values, duplicate values, special symbols, extra spaces and stop words in the historical case data set and the current case data set, to obtain a historical preliminary processing data set and a current preliminary processing data set; Step S22: segmentation, which is used to segment the preliminary processed data into text segments, specifically segmenting the text data in the historical preliminary processed data set and the current preliminary processed data set into segments in units of sentences to obtain a historical data set and a case set to be classified; Step S23: data set segmentation, which is used to segment the data set, specifically, segmenting the historical data set to obtain a classification training set and a classification test set; Step S24: performing preprocessing, specifically, preprocessing the current case data set through the data cleaning and the paragraph segmentation to obtain a case set to be classified, and preprocessing the historical case data set through the data cleaning, the paragraph segmentation and the data set segmentation to obtain a classification training set and a classification test set.
6. The method for classifying civil and commercial cases based on machine learning according to claim 1 is characterized by: In step S5, the civil and commercial case classification is specifically to use the optimized civil and commercial case classification model to process the case set to be classified, obtain civil and commercial case type reference data, and classify the civil and commercial cases based on the civil and commercial case type reference data.
7. A civil and commercial case classification system based on machine learning, used to implement the civil and commercial case classification method based on machine learning as described in any one of claims 1 to 6, characterized in that: It includes data collection module, data preprocessing module, case classification model building module, hyperparameter optimization module and civil and commercial case classification module.
8. The civil and commercial case classification system based on machine learning according to claim 7 is characterized by: The data acquisition module is used for data acquisition, obtains the original case data set through data acquisition, and sends the original case data set to the data preprocessing module; The data preprocessing module is used for data preprocessing. Through data preprocessing, a case set to be classified, a classification training set and a classification test set are obtained, and the case set to be classified is sent to the civil and commercial case classification module, and the classification training set and the classification test set are sent to the case classification model construction module; The case classification model construction module is used for case classification model construction, and obtains a civil and commercial case classification model by constructing a dual interactive self-attention network model, and sends the civil and commercial case classification model to a hyperparameter optimization module; The hyperparameter optimization module is used for hyperparameter optimization, and optimizes the model hyperparameters by using an optimization algorithm to obtain an optimized civil and commercial case classification model, and sends the optimized civil and commercial case classification model to the civil and commercial case classification module; The civil and commercial case classification module is used for classifying civil and commercial cases. Civil and commercial cases are classified by adopting the optimized civil and commercial case classification model to obtain reference data on the types of civil and commercial cases.