Case association identification method and system based on artificial intelligence

Through artificial intelligence-based methods, including data collection, feature extraction and convolutional neural network model construction, the problem of inefficient case association recognition in the existing technology is solved, more efficient and accurate case association recognition is achieved, and work quality and informatization level are improved.

CN119939426APending Publication Date: 2025-05-06SHANGHAI JUYIN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510021055.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology is inefficient in case association identification, easily misses important clues, and is difficult to deal with large-scale and complex data.

Method used

Using an artificial intelligence-based method, a case association model is constructed through data collection and preprocessing, feature extraction and representation, and convolutional neural networks to identify potential association patterns and laws between cases.

Benefits of technology

It improves the efficiency and accuracy of case association identification, reduces labor costs and time costs, and improves the informatization level and work quality of judicial, law enforcement and other work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939426A_ABST
    Figure CN119939426A_ABST
Patent Text Reader

Abstract

The invention discloses a case association identification method and system based on artificial intelligence, and belongs to the technical field of artificial intelligence, and the method comprises the following specific steps: 1, data collection and preprocessing: collecting case related data from a plurality of data sources, comprising case text information, involved person information, material evidence information and time and place information, then the collected data is cleaned, noise data, repeated data and incomplete data are removed, and then standardization processing is carried out, so that the format of the data is unified, and subsequent analysis is facilitated. According to the method, the case association model is constructed by using the convolutional neural network, so that the efficiency and accuracy of case association identification can be effectively improved, the labor cost and the time cost are reduced, the informatization level and the work quality of judicial and law enforcement work and the like are improved, and the method has wide application prospects and practical values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a case association identification method and system based on artificial intelligence. Background Art

[0002] In judicial, law enforcement and various investigations, the number of cases is numerous and complex. Traditional case association identification mainly relies on manual experience and simple data retrieval, which is inefficient, easy to miss important clues, and difficult to deal with large-scale and complex data. With the rapid development of information technology, artificial intelligence technology has shown great potential in data processing and analysis, but there is currently no efficient, accurate and practical case association identification solution based on artificial intelligence to meet the urgent needs of actual work. Summary of the invention

[0003] The technical problem to be solved by the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a case association identification method and system based on artificial intelligence.

[0004] The technical solution adopted to solve the above technical problems is: a case association identification method based on artificial intelligence, including the following specific steps:

[0005] Step 1: Data collection and preprocessing: Collect case-related data from multiple data sources, including case text information, information about people involved in the case, physical evidence information, time and location information, and then clean the collected data to remove noise data, duplicate data, and incomplete data. Then standardize the data to make it in a unified format for subsequent analysis.

[0006] Step 2: Extract features and represent them. Use natural language processing technology to extract features from case text information, extract keywords, key phrases, and semantic vectors to represent the key content of the case. For information about people involved in the case, extract their personal attribute features and social relationship features. For physical evidence information, extract its type, features, and source. Structuralize time and place information to extract time series features and spatial location features.

[0007] Step 3: Use convolutional neural networks to build a case association model. Take the data after feature extraction and representation as the input of the model. Train it with a large amount of labeled and unlabeled case data to enable the model to learn the potential association patterns and rules between cases. During the training process, optimize the model parameters to improve the accuracy and generalization ability of the model.

[0008] Step 4: When new case data is input, it is preprocessed and feature extracted, and then input into the trained case association model. The model outputs the association scores between cases, and determines whether there is an association relationship between cases based on the preset threshold. If the association score is higher than the threshold, it is identified as a related case, and the specific type and degree of association are further analyzed;

[0009] Step 5: Present the results of case association identification to the user in an intuitive and visual manner.

[0010] Furthermore, it includes six subsystems: data acquisition module, data preprocessing module, feature extraction module, case association model construction module, case association identification and analysis module and result visualization module. The data acquisition module is responsible for connecting multiple data sources, collecting case-related data according to predetermined rules and frequencies, and transmitting the collected data to the data preprocessing module; the data preprocessing module is responsible for cleaning and standardizing the data transmitted from the data acquisition module to ensure the quality and consistency of the data, and provide a reliable data foundation for subsequent feature extraction and model training; the feature extraction module is responsible for using natural language processing technology to extract features and quantify the preprocessed data, convert different types of case data into feature vectors that can be processed by the model, and pass them to the case association model construction module.

[0011] Furthermore, the case association model construction module is responsible for building a case association model based on a deep learning framework, and using training data to train and optimize the model so that the model has the ability to accurately identify case associations. After the model training is completed, it waits for the input of new case data for association identification.

[0012] Furthermore, the case association identification and analysis module is responsible for receiving new case data, calling the data preprocessing module and feature extraction module to process it, and then inputting it into the trained case association model to obtain the association score output by the model, and judging the case association based on the threshold, conducting detailed association analysis on related cases, determining the type and degree of association, and transmitting the analysis results to the result visualization module.

[0013] Furthermore, the result visualization module is responsible for visualizing the results transmitted by the case association identification and analysis module, generating an intuitive graphical display interface, displaying the association network and key information between cases, facilitating user viewing and operation, while providing interactive functions, so that users can further query detailed case information and association details as needed.

[0014] Furthermore, the case association model construction module adopts convolutional neural network to construct the model, and the specific expression formula is as follows:.

[0015] Convolutional Layer:

[0016] The case text is converted into an input feature map, whose dimension is Among them, H in is the height of the input feature map, W in is the width, C in is the number of input channels;

[0017] There are N convolution kernels, each with a dimension of (k is the size of the convolution kernel, which is set to a square convolution kernel), for the output feature map A position (i, j, n) in (n represents the nth output channel), its value The calculation of is as follows:

[0018]

[0019] Among them, b n is the bias term corresponding to the nth output channel, x i+h,j+w,c is the value at position (i+h, j+w, c) in the input feature map X, is the nth convolution kernel K n The value at position (h, w, c) in the middle;

[0020] The size of the output feature map is calculated as:

[0021]

[0022] Among them, padding is the number of pixels filled around the input feature map, ranging from 0 to 2, and stride is the step size of the convolution kernel sliding on the input feature map;

[0023] Activation Function

[0024] After the convolution layer, an activation function is usually applied to introduce nonlinear characteristics and enhance the expressiveness of the model. The formula is:

[0025] y=max(0,x)

[0026] Among them, x is the feature value output by the convolutional layer, and y is the output value after activation;

[0027] Pooling Layer

[0028] The input feature map is The pooling window size is p×p and the step size is s; for the output feature map For a position (i, j, c) in , its value is calculated as:

[0029] y ijc =max 0≤h<p,0≤w<p x i×s+h,j×s+w,c

[0030] The size of the output feature map is:

[0031]

[0032] The role of the pooling layer is to downsample the feature map, reduce the size of the feature map, while retaining the main feature information, reducing the amount of calculation and preventing overfitting.

[0033] Furthermore, after multiple convolution, activation and pooling layers, the obtained feature map is flattened into a one-dimensional vector and then connected to the fully connected layer. Let the flattened feature vector be X∈R m , the weight matrix W∈R of the fully connected layer m×n The bias vector is b∈R n , then the output vector y∈R of the fully connected layer n The calculation formula is:

[0034] y=Wx+b

[0035] In case association recognition, the fully connected layer can map the previously extracted features to the final output space. If case association is defined as a binary classification problem, then n = 2. The output is converted into a probability distribution through the softmax(y) function to represent the probability of two cases being associated or unassociated:

[0036]

[0037] Among them, y i is the i-th element of the output vector y, softmax(y i )The vector obtained represents the probability distribution of each category, and the category with the highest probability is the case association result predicted by the model.

[0038] Furthermore, the convolutional neural network is optimized using a batch gradient descent method, and each time a parameter is updated, the gradients of all training samples are used to calculate the parameter update value.

[0039] The beneficial effects of the present invention are as follows: By adopting a convolutional neural network to construct a case association model, the present invention can effectively improve the efficiency and accuracy of case association identification, reduce labor costs and time costs, and improve the informatization level and work quality of judicial and law enforcement work, and has broad application prospects and practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a system framework diagram of the present invention. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0042] like Figure 1 As shown, a case association identification method and system based on artificial intelligence in this embodiment includes the following specific steps:

[0043] Step 1: Data collection and preprocessing: Collect case-related data from multiple data sources, including case text information, information about people involved in the case, physical evidence information, time and location information, and then clean the collected data to remove noise data, duplicate data, and incomplete data. Then standardize the data to make it in a unified format for subsequent analysis.

[0044] Step 2: Extract features and represent them. Use natural language processing technology to extract features from case text information, extract keywords, key phrases, and semantic vectors to represent the key content of the case. For information about people involved in the case, extract their personal attribute features and social relationship features. For physical evidence information, extract its type, features, and source. Structuralize time and place information to extract time series features and spatial location features.

[0045] Step 3: Use convolutional neural networks to build a case association model. Take the data after feature extraction and representation as the input of the model. Train it with a large amount of labeled and unlabeled case data to enable the model to learn the potential association patterns and rules between cases. During the training process, optimize the model parameters to improve the accuracy and generalization ability of the model.

[0046] Step 4: When new case data is input, it is preprocessed and feature extracted, and then input into the trained case association model. The model outputs the association scores between cases, and determines whether there is an association relationship between cases based on the preset threshold. If the association score is higher than the threshold, it is identified as a related case, and the specific type and degree of association are further analyzed;

[0047] Step 5: Present the results of case association identification to the user in an intuitive and visual manner.

[0048] It includes six subsystems: data acquisition module, data preprocessing module, feature extraction module, case association model construction module, case association identification and analysis module, and result visualization module. The data acquisition module is responsible for connecting multiple data sources, collecting case-related data according to predetermined rules and frequencies, and transmitting the collected data to the data preprocessing module; the data preprocessing module is responsible for cleaning and standardizing the data transmitted from the data acquisition module to ensure the quality and consistency of the data, and provide a reliable data foundation for subsequent feature extraction and model training; the feature extraction module is responsible for using natural language processing technology to extract features and quantify the preprocessed data, convert different types of case data into feature vectors that can be processed by the model, and pass them to the case association model construction module.

[0049] The case association model construction module is responsible for building a case association model based on a deep learning framework, and using training data to train and optimize the model so that the model has the ability to accurately identify case associations. After the model training is completed, it waits for the input of new case data for association identification.

[0050] The case association identification and analysis module is responsible for receiving new case data, calling the data preprocessing module and feature extraction module to process it, and then inputting it into the trained case association model to obtain the association score output by the model, and judging the case association based on the threshold, conducting detailed association analysis on related cases, determining the type and degree of association, and transmitting the analysis results to the result visualization module.

[0051] The result visualization module is responsible for visualizing the results transmitted by the case association identification and analysis module, generating an intuitive graphical display interface, showing the association network and key information between cases, making it convenient for users to view and operate. It also provides interactive functions, allowing users to further query detailed case information and association details as needed.

[0052] The case association model construction module adopts convolutional neural network to construct the model, and the specific expression formula is as follows:.

[0053] Convolutional Layer:

[0054] The case text is converted into an input feature map, whose dimension is Among them, H in is the height of the input feature map, W in is the width, C in is the number of input channels;

[0055] There are N convolution kernels, each with a dimension of (k is the size of the convolution kernel, which is set to a square convolution kernel), for the output feature map A position (i, j, n) in (n represents the nth output channel), its value The calculation of is as follows:

[0056]

[0057] Among them, b n is the bias term corresponding to the nth output channel, x i+h,j+w,c is the value at position (i+h, j+w, c) in the input feature map X, is the nth convolution kernel K n The value at position (h, w, c) in the middle;

[0058] The size of the output feature map is calculated as:

[0059]

[0060] Among them, padding is the number of pixels filled around the input feature map, ranging from 0 to 2, and stride is the step size of the convolution kernel sliding on the input feature map;

[0061] Activation Function

[0062] After the convolution layer, an activation function is usually applied to introduce nonlinear characteristics and enhance the expressiveness of the model. The formula is:

[0063] y=max(0,x)

[0064] Among them, x is the feature value output by the convolutional layer, and y is the output value after activation;

[0065] Pooling Layer

[0066] The input feature map is The pooling window size is p×p and the step size is s; for the output feature map For a position (i, j, c) in , its value is calculated as:

[0067] y ijc =max 0≤h<p,0≤w<p x i×s+h,j×s+w,c

[0068] The size of the output feature map is:

[0069]

[0070] The role of the pooling layer is to downsample the feature map, reduce the size of the feature map, while retaining the main feature information, reducing the amount of calculation and preventing overfitting.

[0071] After multiple convolution, activation and pooling layers, the obtained feature map is flattened into a one-dimensional vector and then connected to the fully connected layer. Let the flattened feature vector be X∈R m , the weight matrix W∈R of the fully connected layer m×n The bias vector is b∈R n , then the output vector y∈R of the fully connected layer n The calculation formula is:

[0072] y=Wx+b

[0073] In case association recognition, the fully connected layer can map the previously extracted features to the final output space. If case association is defined as a binary classification problem, then n = 2. The output is converted into a probability distribution through the softmax(y) function to represent the probability of two cases being associated or unassociated:

[0074]

[0075] Among them, y i is the i-th element of the output vector y, softmax(y i )The vector obtained represents the probability distribution of each category. The category with the highest probability is the case association result predicted by the model. The convolutional neural network is optimized using the batch gradient descent method. Each time the parameters are updated, the gradients of all training samples are used to calculate the parameter update value.

[0076] During the data collection and preprocessing stage, connections are established with various data sources through the data collection module, and the text description of the case and the basic information table of the persons involved are obtained from the judicial database through database query statements. Detailed records of physical evidence and time and location information of on-site investigations are obtained from the law enforcement record system.

[0077] The collected data is cleaned and standardized by the data preprocessing module to remove garbled characters, incorrect formats, and obviously illogical data, unify the date format to "YYYY-MM-DD", and convert full-width characters in the text into half-width characters.

[0078] In the feature extraction and representation stage, for case texts, natural language processing tools are used to segment the text into words or phrases and convert them into corresponding word vectors. For information on persons involved in the case, their age, gender and other attributes are numerically encoded, and for social relationships, a relationship matrix is ​​constructed to represent them. Physical evidence information is classified and coded and features are quantified according to its type and characteristics.

[0079] When constructing the case association model, a convolutional neural network is used to build the model, and a large amount of historical case data is used for training. The weight parameters of the model are continuously adjusted through the back propagation algorithm to minimize the error between the predicted results and the actual association situation.

[0080] During the case association identification and analysis stage, when new case data arrives, the system automatically starts the data preprocessing and feature extraction process, and then inputs the feature data into the trained case association model. The model quickly calculates the association score between the case and other cases.

[0081] For a new case, if the model finds that it has a high degree of correlation with previous cases in terms of modus operandi, proximity of time and place, and the correlation score exceeds the preset threshold, it will determine that these cases are related, and further analyze the types of personnel connections that may be committed by the same criminal gang, as well as detailed information such as the degree of spatial correlation at the crime scene.

[0082] Finally, the result visualization module presents this association information to the user in a visual way. The user can see that each case is distributed in the form of nodes on the graphical interface, and the related cases are connected by lines. The thickness and color of the lines intuitively reflect the strength of the association.

[0083] Users can also click on nodes to view detailed information about the case, such as case details, a list of persons involved in the case, a list of physical evidence, etc., to help users quickly grasp the overall situation and key clues of the case, and provide strong support for the investigation and trial of the case.

[0084] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. A case association identification method based on artificial intelligence, characterized in that: The specific steps include: Step 1: Data collection and preprocessing: Collect case-related data from multiple data sources, including case text information, information about people involved in the case, physical evidence information, time and location information, and then clean the collected data to remove noise data, duplicate data, and incomplete data. Then standardize the data to make it in a unified format for subsequent analysis. Step 2: Extract features and represent them. Use natural language processing technology to extract features from case text information, extract keywords, key phrases, and semantic vectors to represent the key content of the case. For information about people involved in the case, extract their personal attribute features and social relationship features. For physical evidence information, extract its type, features, and source. Structuralize time and place information to extract time series features and spatial location features. Step 3: Use convolutional neural networks to build a case association model. Take the data after feature extraction and representation as the input of the model. Train it with a large amount of labeled and unlabeled case data to enable the model to learn the potential association patterns and rules between cases. During the training process, optimize the model parameters to improve the accuracy and generalization ability of the model. Step 4: When new case data is input, it is preprocessed and feature extracted, and then input into the trained case association model. The model outputs the association scores between cases, and determines whether there is an association relationship between cases based on the preset threshold. If the association score is higher than the threshold, it is identified as a related case, and the specific type and degree of association are further analyzed; Step 5: Present the results of case association identification to the user in an intuitive and visual manner.

2. The case association identification system based on artificial intelligence according to claim 1 is characterized in that: It includes six subsystems: data collection module, data preprocessing module, feature extraction module, case association model building module, case association identification and analysis module, and result visualization module. The data collection module is responsible for connecting multiple data sources, collecting case-related data according to predetermined rules and frequencies, and transmitting the collected data to the data preprocessing module; The data preprocessing module is responsible for cleaning and standardizing the data sent by the data acquisition module to ensure the quality and consistency of the data and provide a reliable data basis for subsequent feature extraction and model training; The extraction module is responsible for using natural language processing technology to perform feature extraction and quantitative representation on the preprocessed data, converting different types of case data into feature vectors that can be processed by the model, and passing them to the case association model construction module.

3. The case association identification system based on artificial intelligence according to claim 2 is characterized in that: The case association model construction module is responsible for building a case association model based on a deep learning framework, and using training data to train and optimize the model so that the model has the ability to accurately identify case associations. After the model training is completed, it waits for the input of new case data for association identification.

4. The case association identification system based on artificial intelligence according to claim 3 is characterized in that: The case association identification and analysis module is responsible for receiving new case data, calling the data preprocessing module and feature extraction module to process it, and then inputting it into the trained case association model to obtain the association score output by the model, and judging the case association based on the threshold, conducting detailed association analysis on related cases, determining the type and degree of association, and transmitting the analysis results to the result visualization module.

5. The case association identification system based on artificial intelligence according to claim 4 is characterized in that: The result visualization module is responsible for visualizing the results transmitted by the case association identification and analysis module, generating an intuitive graphical display interface, showing the association network and key information between cases, making it convenient for users to view and operate. It also provides interactive functions, allowing users to further query detailed case information and association details as needed.

6. The case association identification system based on artificial intelligence according to claim 5 is characterized in that: The case association model construction module adopts convolutional neural network to construct the model, and the specific expression formula is as follows:. Convolutional Layer: The case text is converted into an input feature map, whose dimension is Among them, H in is the height of the input feature map, W in is the width, C in is the number of input channels; There are N convolution kernels, each with a dimension of (k is the size of the convolution kernel, which is set to a square convolution kernel), for the output feature map A position (i, j, n) in (n represents the nth output channel), its value The calculation of is as follows: Among them, b n is the bias term corresponding to the nth output channel, x i+h,j+w,c is the value at position (i+h, j+w, c) in the input feature map X, is the nth convolution kernel K n The value at position (h, w, c) in the middle; The size of the output feature map is calculated as: Among them, padding is the number of pixels filled around the input feature map, ranging from 0 to 2, and stride is the step size of the convolution kernel sliding on the input feature map; Activation Function After the convolution layer, an activation function is usually applied to introduce nonlinear characteristics and enhance the expressiveness of the model. The formula is: y=max(0,x) Among them, x is the feature value output by the convolutional layer, and y is the output value after activation; Pooling Layer The input feature map is The pooling window size is p×p and the step size is s; for the output feature map For a position (i, j, c) in , its value is calculated as: and ijc =max 0≤h<p,0≤w<p x i×s+h,j×s+w,c The size of the output feature map is: The role of the pooling layer is to downsample the feature map, reduce the size of the feature map, while retaining the main feature information, reducing the amount of calculation and preventing overfitting.

7. The case association identification system based on artificial intelligence according to claim 6 is characterized in that: After multiple convolution, activation and pooling layers, the obtained feature map is flattened into a one-dimensional vector and then connected to the fully connected layer. Let the flattened feature vector be X∈R m , the weight matrix W∈R of the fully connected layer m×n The bias vector is b∈R n , then the output vector y∈R of the fully connected layer n The calculation formula is: y=Wx+b In case association recognition, the fully connected layer maps the previously extracted features to the final output space. Case association is defined as a binary classification problem, then n = 2. The output is converted into a probability distribution through the softmax(y) function to represent the probability of two cases being associated or unassociated: Among them, y i is the i-th element of the output vector y, softmax(y i )The vector obtained represents the probability distribution of each category, and the category with the highest probability is the case association result predicted by the model.

8. The case association identification system based on artificial intelligence according to claim 7 is characterized in that: The convolutional neural network is optimized using a batch gradient descent method, and each time a parameter is updated, the gradients of all training samples are used to calculate the parameter update value.