Case analysis method and device and storage medium
By preprocessing the judicial documents and extracting features of large language models, combined with complex networks and statistical analysis, the problem of difficulty in taking into account the flexibility, economy and accuracy of case extraction and feature analysis in the existing technology is solved, and a more in-depth and efficient case analysis is achieved.
Patent Information
- Application Number
- CN202510132471.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-09
AI Technical Summary
The existing methods of extracting case content and feature analysis of judicial documents cannot take into account flexibility, economy and accuracy, and the case feature analysis is not in-depth enough.
By pre-processing the judgment documents of the target case type, a large language model is used to extract case feature information, including case personnel characteristics and case process characteristics, and complex network analysis and statistical analysis are carried out based on these characteristics.
It realizes the rapid, accurate and flexible extraction of the characteristics of the case personnel and the characteristics of the case process from the judgment documents, improves the depth and accuracy of the analysis, reduces the training cost, and is more economical and efficient than manual methods.
Smart Images

Figure CN119962520A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence and big data technology, and in particular to a case analysis method, device and storage medium. Background Art
[0002] In the field of case analysis (e.g., crime analysis) based on cases (e.g., judicial documents), researchers have long analyzed cases from the perspectives of phenomena, causes, and measures, lacking certain quantitative analysis methods. Due to the diverse content of the judicial documents corpus, the complexity of many cases, and the differences in regulations and language habits in various places, the structured degree of judicial documents is poor. Case analysis based on cases can be briefly divided into two parts: case content extraction and case feature analysis.
[0003] At present, the case content extraction and case feature analysis of judicial documents mainly rely on three methods: integrating and extracting key fields based on the judicial document payment platform, manually traversing judicial documents for selection, and extracting and analyzing based on deep learning models. However, the payment platform method has poor flexibility, the manual method has high time and economic costs, and the method based on deep learning models has poor accuracy and efficiency on less structured documents, and training new models requires high initial costs. In summary, the existing technical solutions cannot take into account the flexibility, economy and accuracy of extraction, and the case feature analysis is not in-depth enough, and it is in urgent need of improvement and optimization. Summary of the invention
[0004] In view of this, the present disclosure proposes a case analysis method, device and storage medium.
[0005] According to one aspect of the present disclosure, a case analysis method is provided. The method comprises:
[0006] Preprocessing the text of at least one case of the target case type to obtain a preprocessed text, wherein the text of the case includes a judgment document;
[0007] Input the preprocessed text into the large language model to obtain case feature information, which includes case personnel features and case process features;
[0008] The analysis is performed based on the case characteristic information to obtain the analysis results.
[0009] In a possible implementation, preprocessing the text of at least one case of the target case type to obtain the preprocessed text includes:
[0010] Regularized extraction of case personnel information and case process information in the text of at least one case of the target case type to obtain a regularized extracted text;
[0011] Determine the prompt text based on the target case type, and use the prompt text to guide the large language model to extract case feature information;
[0012] The regularized extracted text and the prompt text are combined to obtain the preprocessed text.
[0013] In a possible implementation, regularized extraction is performed on case personnel information and case process information in the text of at least one case of the target case type to obtain regularized extracted text, including:
[0014] The preset regularized expressions are used to match the text including the case personnel information and the text including the case process information in the case text respectively, and the text including the case personnel information and the text including the case process information are used as the regularized extracted text.
[0015] In a possible implementation, the preprocessed text is input into a large language model to obtain case feature information, including:
[0016] The preprocessed text is input into the large language model, so that the large language model performs feature extraction on the regularized extracted text based on the prompt text in the preprocessed text to obtain case feature information;
[0017] The prompt text includes any one or more of text indicating a feature extraction type, text indicating a feature extraction range, and text indicating a feature extraction format.
[0018] In one possible implementation, case personnel characteristics and case process characteristics are determined based on the target case type;
[0019] Among them, the case person characteristics include any one or more of the case person's age, education level, occupation, and address;
[0020] The case process characteristics include any one or more of the geographical feature information involved in the case process, the behavioral feature information involved in the case process, the item information involved in the case process, and the division of labor information of each case personnel in the case process.
[0021] In a possible implementation, the analysis results include complex network analysis results, which are analyzed based on case feature information to obtain analysis results including:
[0022] Based on the geographical feature information involved in the case process characteristics, a complex network of the target case type is constructed. The nodes of the complex network represent the locations involved in the case process, and the edges of the complex network represent the migration path between two locations in the case process.
[0023] Analyze the complex network of at least one case type to obtain complex network analysis results.
[0024] In a possible implementation, the complex network analysis results include key locations and key paths in the case process. The complex network of at least one case type is analyzed to obtain the complex network analysis results, including:
[0025] Calculate the node centrality of each node in the complex network of the target case type, and take the location corresponding to the node whose node centrality is greater than a first preset threshold as the key location;
[0026] The edge centrality is calculated for each edge in the complex network of the target case type, and the path corresponding to the edge whose edge centrality is greater than a second preset threshold is taken as the key path.
[0027] In one possible implementation, the node centrality is determined based on the number of edges that exist between the node and other adjacent nodes;
[0028] Edge centrality is determined based on the proportion of the shortest paths between all node pairs in a complex network that pass through the edge.
[0029] In a possible implementation, the analysis result includes a statistical analysis result, which is obtained based on the case feature information and includes:
[0030] Statistical analysis is performed based on case personnel characteristics and case process characteristics to obtain statistical analysis results, which include any one or more of statistical distribution results and correlation analysis results.
[0031] According to another aspect of the present disclosure, a case analysis device is provided. The device comprises:
[0032] A preprocessing module, used for preprocessing the text of at least one case of the target case type to obtain a preprocessed text, wherein the text of the case includes a judgment document;
[0033] A determination module is used to input the preprocessed text into a large language model to obtain case feature information, which includes case personnel features and case process features;
[0034] The analysis module is used to perform analysis based on case feature information and obtain analysis results.
[0035] In a possible implementation, the preprocessing module is used to:
[0036] Regularized extraction of case personnel information and case process information in the text of at least one case of the target case type to obtain a regularized extracted text;
[0037] Determine the prompt text based on the target case type, and use the prompt text to guide the large language model to extract case feature information;
[0038] The regularized extracted text and the prompt text are combined to obtain the preprocessed text.
[0039] In a possible implementation, regularized extraction is performed on case personnel information and case process information in the text of at least one case of the target case type to obtain regularized extracted text, including:
[0040] The preset regularized expressions are used to match the text including the case personnel information and the text including the case process information in the case text respectively, and the text including the case personnel information and the text including the case process information are used as the regularized extracted text.
[0041] In a possible implementation, a module is determined to:
[0042] The preprocessed text is input into the large language model, so that the large language model performs feature extraction on the regularized extracted text based on the prompt text in the preprocessed text to obtain case feature information;
[0043] The prompt text includes any one or more of text indicating a feature extraction type, text indicating a feature extraction range, and text indicating a feature extraction format.
[0044] In one possible implementation, case personnel characteristics and case process characteristics are determined based on the target case type;
[0045] Among them, the case person characteristics include any one or more of the case person's age, education level, occupation, and address;
[0046] The case process characteristics include any one or more of the geographical feature information involved in the case process, the behavioral feature information involved in the case process, the item information involved in the case process, and the division of labor information of each case personnel in the case process.
[0047] In a possible implementation, the analysis result includes a complex network analysis result, and the analysis module is used to:
[0048] Based on the geographical feature information involved in the case process characteristics, a complex network of the target case type is constructed. The nodes of the complex network represent the locations involved in the case process, and the edges of the complex network represent the migration path between two locations in the case process.
[0049] Analyze the complex network of at least one case type to obtain complex network analysis results.
[0050] In a possible implementation, the complex network analysis results include key locations and key paths in the case process. The complex network of at least one case type is analyzed to obtain the complex network analysis results, including:
[0051] Calculate the node centrality of each node in the complex network of the target case type, and take the location corresponding to the node whose node centrality is greater than a first preset threshold as the key location;
[0052] The edge centrality is calculated for each edge in the complex network of the target case type, and the path corresponding to the edge whose edge centrality is greater than a second preset threshold is taken as the key path.
[0053] In one possible implementation, the node centrality is determined based on the number of edges that exist between the node and other adjacent nodes;
[0054] Edge centrality is determined based on the proportion of the shortest paths between all node pairs in a complex network that pass through the edge.
[0055] In a possible implementation, the analysis result includes a statistical analysis result, and the analysis module is used to:
[0056] Statistical analysis is performed based on case personnel characteristics and case process characteristics to obtain statistical analysis results, which include any one or more of statistical distribution results and correlation analysis results.
[0057] According to another aspect of the present disclosure, a case analysis device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0058] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0059] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0060] According to the embodiment of the present disclosure, by preprocessing the text of at least one case of the target case type including the judgment document, a preprocessed text is obtained, the preprocessed text is input into the large language model, and the case feature information including the case personnel features and the case process features is obtained, and the case feature information is analyzed based on the case feature information to obtain the analysis result, so that the large language model can be used to quickly, accurately and flexibly extract the case personnel features and the case process features from the judgment document, so that the scheme of the present disclosure has higher flexibility and extraction accuracy, and is more economical and efficient than the manual method. At the same time, based on the more accurate and rich features extracted, a more in-depth analysis result can be obtained, which is more accurate and efficient than the method based on the deep learning model, and the training cost is low. The analysis results of the embodiment of the present disclosure can be used to determine the targeted prevention and control measures related to the case personnel and the case process or as reference information for case detection, so as to promote the improvement of case analysis and the progress of related prevention and control governance.
[0061] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0063] Figure 1 A schematic diagram showing an application scenario according to an embodiment of the present disclosure.
[0064] Figure 2 A flowchart of a case analysis method according to an embodiment of the present disclosure is shown.
[0065] Figure 3 A schematic diagram showing a complex network according to an embodiment of the present disclosure.
[0066] Figure 4 A schematic diagram showing statistical analysis results according to an embodiment of the present disclosure.
[0067] Figure 5 The overall flow chart of the case analysis method according to the embodiment of the present disclosure is shown.
[0068] Figure 6 A structural diagram of a case analysis device according to an embodiment of the present disclosure is shown.
[0069] Figure 7 It is a block diagram of a device 1900 for case analysis according to an exemplary embodiment. DETAILED DESCRIPTION
[0070] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0071] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0072] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.
[0073] In the field of case analysis (e.g., crime analysis) based on cases (e.g., judicial documents), researchers have long analyzed cases from the perspectives of phenomena, causes, and measures, lacking certain quantitative analysis methods. Due to the diverse content of the judicial documents corpus, the complexity of many cases, and the differences in regulations and language habits in various places, the structured degree of judicial documents is poor. Case analysis based on cases can be briefly divided into two parts: case content extraction and case feature analysis.
[0074] At present, the case content extraction and case feature analysis of judicial documents mainly rely on three methods: integrating and extracting key fields based on the judicial document payment platform, manually traversing judicial documents for selection, and extracting and analyzing based on deep learning models. However, the payment platform method has poor flexibility, the manual method has high time and economic costs, and the method based on deep learning models has poor accuracy and efficiency on less structured documents, and training new models requires high initial costs. In summary, the existing technical solutions cannot take into account the flexibility, economy and accuracy of extraction, and the case feature analysis is not in-depth enough, and it is in urgent need of improvement and optimization.
[0075] In view of this, the present disclosure provides a case analysis method, device and storage medium. The method of the embodiment of the present disclosure preprocesses the text of at least one case including a judgment document of the target case type to obtain a preprocessed text, inputs the preprocessed text into a large language model, obtains case feature information including case personnel features and case process features, analyzes based on the case feature information, and obtains analysis results, which can realize the use of a large language model to quickly, accurately and flexibly extract case personnel features and case process features from judgment documents, so that the scheme of the present disclosure has higher flexibility and extraction accuracy, and is more economical and efficient than manual methods. At the same time, based on the more accurate and rich features extracted, a more in-depth analysis result can be obtained, which is more accurate and efficient than the method based on the deep learning model, and the training cost is low. The analysis results of the embodiment of the present disclosure can be used to determine targeted prevention and control measures related to case personnel and case processes or as reference information for case detection, thereby promoting the improvement of case analysis and the progress of related prevention and control governance.
[0076] Figure 1 A schematic diagram of an application scenario according to an embodiment of the present disclosure is shown. The case analysis system of the embodiment of the present disclosure can be used to perform batch analysis on judicial documents, and generate corresponding analysis results by automatically extracting features and performing feature analysis. The output analysis results can be used by relevant personnel to understand the trends, laws and potential risks in the case, and provide a scientific basis for relevant decision-making, so as to assist relevant personnel in formulating targeted prevention and control measures related to case personnel and case processes.
[0077] like Figure 1 As shown, the case analysis system of the embodiment of the present disclosure may include a data preprocessing module, a large language model module, a complex network analysis module and a statistical analysis module. Among them, the data preprocessing module can be used to obtain the text of at least one case for preprocessing and output the preprocessed text; the large language model module can be used to extract features based on the preprocessed text and output case feature information; the complex network module and the statistical analysis module can be used to perform analysis based on the case feature information to obtain corresponding analysis results.
[0078] The case analysis system of the disclosed embodiment can be used for a terminal device or a server, and the terminal device can be any one or more of a mobile phone, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, and a vehicle-mounted device. The disclosed embodiment does not impose any special restrictions on the specific type of the terminal device, and it can have a wired or wireless communication function. The server can be located locally or in the cloud, and can be a physical device or a virtual device, such as a virtual machine, a container, etc., with a wireless communication function, wherein the wireless communication function can be set in the chip (system) or other parts or components of the server. The wireless communication function can be realized, for example, by mobile communication technologies such as 2G / 3G / 4G / 5G, as well as Wi-Fi, Bluetooth, frequency modulation (FM), digital radio, satellite communication, etc. Communication can also be carried out by wired connection to achieve interaction with other devices.
[0079] Figure 2 A flowchart of a case analysis method according to an embodiment of the present disclosure is shown. The method can be used in the above-mentioned case analysis system, such as Figure 2 As shown, the method may include:
[0080] Step S201, preprocessing the text of at least one case of the target case type to obtain a preprocessed text.
[0081] The cases in the disclosed embodiments may refer to various types of cases in the field of crime. Different case types may correspond to different types of crimes, such as property crimes, financial crimes, contraband crimes, etc. The target case type may correspond to any of the above crime types, such as contraband crimes. When classifying case types, it is also possible to further refine them in combination with regional factors, such as classifying the target case type into contraband crime cases in a certain province or city, so as to more accurately meet the analysis needs of specific regions or scenarios.
[0082] For the target case type, the text of at least one related case can be obtained. For example, for the contraband crime in X Province, X City, the text of the case related to the contraband crime in X Province, X City can be obtained. The text of the case can be uploaded by the user in the form of a text document or obtained through other means, such as through a public legal database or platform. The text of the case can include a judgment document, and can also include other legal documents or judicial documents, etc. This disclosure does not limit this, as long as the text of the case includes information related to the persons involved (such as criminals, etc.) and the process.
[0083] The preprocessing process in the embodiment of the present disclosure can be performed by Figure 1 The data preprocessing module in the implementation can be used to disassemble at least one case text segment by segment through a regularized extraction algorithm, remove information irrelevant to the case personnel and process, and extract key information fields (such as case personnel and case process). According to the crime type and extraction requirements, the preset prompt text is combined with the key information field to output structured preprocessed data. In step S201, it is possible to:
[0084] Regularized extraction is performed on case personnel information and case process information in the text of at least one case of the target case type to obtain regularized extracted text; prompt text is determined based on the target case type; and the regularized extracted text and the prompt text are combined to obtain preprocessed text.
[0085] The case personnel information in the case text may represent relevant information of the persons involved in the case, such as name, gender, age, date of birth, occupation, education, place of residence, etc.; the case process information in the case text may be a description of the crime process and criminal facts in the case text.
[0086] In the process of performing regularized extraction on case personnel information and case process information in the text of at least one case of the target case type to obtain the regularized extracted text, the following can be done:
[0087] The preset regularized expressions are used to match the text including the case personnel information and the text including the case process information in the case text respectively, and the text including the case personnel information and the text including the case process information are used as the regularized extracted text.
[0088] For the target case type, if the texts of multiple cases are obtained, for example, if 4,000 judicial documents are obtained, the case personnel information and case process information can be regularized and extracted from the text of each case separately using the preset regularization expression to obtain the regularized extracted text.
[0089] For example, the following preset regular expression can be used to extract the case personnel information corresponding to some case personnel in the judgment document:
[0090] pattern1=r"Case personnel [\u4e00-\u9fa5]+, male, [\u4e00-\u9fa5]+"
[0091] Among them, pattern1 is used to store the matched case personnel information, "case personnel", ", male," is a fixed text for positioning, and "[\u4e00-\u9fa5]+" is used to match any number of Chinese characters (excluding punctuation marks). By using this regular expression, when "case personnel Zhang San, male" appears in the case text, the regular expression will match "case personnel Zhang San, male".
[0092] The case process information in the judgment documents can be extracted through the following preset regularized expressions:
[0093] pattern2=r"After investigation, it was found that [\u4e00-\u9fa5]+"
[0094] Among them, pattern2 is used to store the matched case process information, and "after trial and investigation," is a fixed text used for positioning.
[0095] For a certain judgment document, an example of text including case personnel information in the judgment document obtained by regularization extraction is:
[0096] “Case Person 1: Name: Yang XX; Gender: Male; Age: 35; Date of Birth: XXXX-XX-XX; Occupation: Unemployed; Education: Junior High School; Place of Residence: XX City, XX Province
[0097] Case Person 2: Name: Li XX; Gender: Male; Age: 33; Date of Birth: XXXX-XX-XX; Occupation: Unemployed; Education: Illiterate; Place of Residence: XX City, XX Province”;
[0098] An example of text containing case process information in a judgment document obtained through regularization extraction is:
[0099] "Case process: The Second Branch of the XX City Procuratorate accused the defendant Yang XX of shipping a package containing XXX from XX City, XX Province to XX City via a long-distance passenger bus (license plate number ×××) on XX / XX / XXXX. At about 7 o'clock the next day, when the bus arrived at a location in XX District, XX City, the defendant Li XX was instructed by others to take the package from the bus driver and was seized on the spot by the police. Upon identification, the XXX content was 64.4%. The above-mentioned XXX has been confiscated."
[0100] The text including the case personnel information and the text including the case process information in the case text may be used as regularized extracted text.
[0101] In order to extract relevant features from the regularized extracted text more accurately and efficiently in the future, the disclosed embodiment also designs prompt texts for different case types.
[0102] The prompt text may be pre-set as needed to guide the large language model to extract case feature information. The prompt text may include any one or more of text indicating the feature extraction type, text indicating the feature extraction range, and text indicating the feature extraction format. For different case types, the content of the text indicating the feature extraction type, the text indicating the feature extraction range, and the text indicating the feature extraction format may be partially or completely different, or may be the same.
[0103] Among them, the text indicating the feature extraction type can be determined based on the type of case feature information to be extracted, and is used to clarify the type of feature to be extracted; the text indicating the feature extraction range can be used to clarify the feature range extracted under the corresponding feature type, that is, the extracted features should belong to at least one feature specified by the feature extraction range; the text indicating the feature extraction format can be used to clarify the format of the feature extraction results output by the large language model. In this way, the results of subsequent feature extraction can be further improved in terms of standardization, accuracy, etc.
[0104] Taking the case type of contraband crime as an example, the preset prompt text is as follows:
[0105] "Below is a passage about a legal case involving contraband. Please extract the following information based on the passage:
[0106] 1. Contraband transportation method (one of the following three: express delivery, body transportation, transportation by vehicle)
[0107] 2. Means of transportation for transporting prohibited items (one of the following four: by express, by bus, by car, by taxi, by motorcycle, by flight, by train)
[0108] 3. Quantity of contraband items involved
[0109] (Output format: contraband xxx, total weight xxx)
[0110] 4. The spatial transportation route of the contraband involved in the case (list the locations where the contraband arrived in order, separated by commas)
[0111] 5. The role of each person in the case (in the form of name + role, role type: contraband transporter, contraband inspector, contraband seller, person who organizes others to transport contraband)
[0112] 6. Place where the prohibited items were sent from
[0113] 7. Destination of prohibited items
[0114] 8. Places where prohibited items pass through
[0115] 9. Location where the contraband was seized
[0116] In the above prompt text, "methods of hiding and transporting prohibited items", "vehicles used to transport prohibited items", "number of prohibited items involved in the case", "spatial transportation path of prohibited items involved in the case", "role assumed by each person in the case", "place where prohibited items are sent from", "destination of prohibited items", "place where prohibited items are passed through", "location where prohibited items are seized", etc. can represent text indicating the type of feature extraction; "output one of the following three: express delivery, body concealment, and concealment by means of transportation", "output one of the following four: with the help of express delivery, with the help of buses, with the help of self-driving cars, with the help of taxis, with the help of motorcycles, with the help of flights, with the help of trains", "role type: contraband transporter, contraband inspector, contraband seller, person who organizes others to transport contraband", etc. can represent text indicating the scope of feature extraction; "output format: prohibited items xxx, total weight xxx", "list the places where the contraband arrives in sequence, separated by commas", "in the form of name + role", etc. can represent text indicating the format of feature extraction.
[0117] The above prompt text is only an example. For other types of cases, the content in the prompt text can be adjusted accordingly according to specific needs. For example, for financial crime cases, the content in the prompt text can be replaced with elements that adapt to the characteristics of financial crimes, such as replacing the text indicating the feature extraction type with "financial crime methods", "amount involved in the case", "criminal fund flow path", "location of the crime", etc. At the same time, some common parts can remain unchanged. For example, the content such as "the role assumed by each case personnel" in the text indicating the feature extraction type does not need to be modified. This flexibility enables the prompt text to be compatible with multiple case types while retaining applicable common fields, improving its adaptability and extensibility.
[0118] The regularized extracted text and the prompt text can be combined to obtain the preprocessed text. For example, the regularized extracted text is placed after the prompt text as a context to guide the large language model to further process, thereby combining to generate a complete preprocessed text.
[0119] Step S202: input the preprocessed text into the large language model to obtain case feature information.
[0120] This process can be done by Figure 1The large language model module in the implementation can be any type of deep learning model that can perform natural language processing tasks, such as ChatGPT-API (Chat Generative Pre-trainedTransformer API), ChatGLM (Chat Generative Language Model), etc. The preprocessed text can be input into the large language model through the corresponding interface of the large language model to extract the case feature information using the large language model. In view of the problems of short length, different formats, and poor structuring in the case text, the embodiment of the present disclosure can quickly, accurately, and flexibly extract case feature information by utilizing the powerful versatility of the large language model.
[0121] The case characteristic information includes case personnel characteristics and case process characteristics. Case personnel characteristics may include any one or more of the case personnel's age, education level, occupation, and address; case process characteristics may include any one or more of the geographical feature information involved in the case process, the behavioral feature information involved in the case process, the case item information involved in the case process, and the division of labor information of each case personnel in the case process.
[0122] The geographical feature information involved in the case process can indicate the geographical location, location migration path and other information involved in the occurrence and development of the case; the behavioral feature information involved in the case process can indicate the criminal acts and reaction behaviors performed by the various case personnel (such as criminals) in the case process; the information on items involved in the case process can indicate the quantity, weight, volume and other information of the items, evidence, tools, etc. involved in the case; the division of labor information of the various case personnel in the case process can indicate the role assumed by each case personnel.
[0123] Case personnel characteristics and case process characteristics can be determined based on the target case type. For different case types, the extracted case personnel characteristics can be different. For example, for some case types, the age of the case personnel may be more concerned than other characteristics. In this case, only the age can be extracted when extracting case personnel characteristics; while for another part of the case types, the education level and occupation may be more concerned than other characteristics. In this case, only the education level and occupation can be extracted when extracting case personnel characteristics. The same is true for the case process characteristics extracted for different case types.
[0124] In step S202, it is possible to:
[0125] The preprocessed text is input into the large language model, so that the large language model performs feature extraction on the regularized extracted text based on the prompt text in the preprocessed text to obtain case feature information.
[0126] The large language model can determine the type, scope and format of feature extraction for the regularized extracted text based on the prompt text, thereby generating corresponding case feature information based on the regularized extracted text.
[0127] For example, examples of case personnel features obtained after feature extraction of the regularized extracted text corresponding to a certain judgment document include: "Age: 35, Education level: junior high school, Occupation: illiterate; Place of residence: XX city, XX province".
[0128] For another example, with respect to the prompt text of the contraband crime given in the example of step S201 above, an example of the case process features that can be obtained after feature extraction of the regularized extracted text corresponding to a certain judgment document is as follows:
[0129] "1. Contraband transportation method: hiding in the body
[0130] 2. Transportation of contraband: Self-driving cars
[0131] 3. Quantity of contraband involved: Net weight of contraband 370 grams
[0132] 4. The spatial transportation route of the contraband involved in the case: No specific route was mentioned
[0133] 5. Role of each case personnel: XXX-contraband transporter
[0134] 6. Place where the prohibited items were sent from: No specific location was mentioned
[0135] 7. Destination of prohibited items: XX City
[0136] 8. Contraband passage: 65 km post on XXX Road
[0137] 9. Location where the contraband was seized: 65 km post on XXX Road
[0138] Among them, “spatial transportation route of contraband involved in the case”, “place where contraband is sent from”, “destination of contraband”, “place where contraband is passed through”, and “location where contraband is seized” may correspond to the geographical feature information involved in the case process; “method of hiding and transporting contraband” and “vehicles used to transport contraband” may correspond to the behavioral feature information involved in the case process; “number of contraband involved in the case” may correspond to the information on items involved in the case process; and “role assumed by each case personnel” may correspond to the division of labor information of each case personnel during the case process.
[0139] By preprocessing and extracting features from the texts of multiple cases under the target case type, multiple case personnel features and case process features can be quickly and accurately obtained. These case personnel features and case process features can be formatted in a unified manner, for example, output in a tabular form or other structured form to facilitate subsequent analysis.
[0140] Step S203: perform analysis based on the case feature information to obtain analysis results.
[0141] According to the embodiment of the present disclosure, by preprocessing the text of at least one case of the target case type including the judgment document, a preprocessed text is obtained, the preprocessed text is input into the large language model, and the case feature information including the case personnel features and the case process features is obtained, and the case feature information is analyzed based on the case feature information to obtain the analysis result, so that the large language model can be used to quickly, accurately and flexibly extract the case personnel features and the case process features from the judgment document, so that the scheme of the present disclosure has higher flexibility and extraction accuracy, and is more economical and efficient than the manual method. At the same time, based on the more accurate and rich features extracted, a more in-depth analysis result can be obtained, which is more accurate and efficient than the method based on the deep learning model, and the training cost is low. The analysis results of the embodiment of the present disclosure can be used to determine the targeted prevention and control measures related to the case personnel and the case process or as reference information for case detection, so as to promote the improvement of case analysis and the progress of related prevention and control governance.
[0142] The disclosed embodiments may utilize either or both of complex network analysis and statistical analysis in the process of analyzing case characteristic information. Thus, the analysis results may include complex network analysis results and statistical analysis results. The analysis results may be used to determine targeted prevention and control measures related to case personnel and case processes or as reference information for case detection. Thus, it is possible to analyze the spatial and temporal evolution characteristics of crimes, so as to better reveal the deep-seated laws of various crimes, and solve the problem that the flexibility, economy, and accuracy of legal judgment extraction cannot be taken into account at the same time, and the crime characteristic analysis is not in-depth enough.
[0143] The process of complex network analysis can be Figure 1 The complex network analysis module in step S203 can be implemented as follows:
[0144] Based on the geographical feature information involved in the case process characteristics, a complex network of the target case type is constructed; the complex network of at least one case type is analyzed to obtain a complex network analysis result.
[0145] Compared with the simple statistical analysis method of sampling, through analysis based on complex networks, we can make full use of the large amount of crime characteristic data extracted by the large language model, analyze the crime process more deeply, explain the geographical correlation in the crime process, and make the analysis results richer and more complete.
[0146] The geographical feature information involved in the case process characteristics refers to the above-extracted geographical location, location migration path and other information. For example, for crimes involving contraband, the geographical feature information involved in the case process characteristics may include "spatial transportation path of contraband involved in the case", "place where contraband was sent from", "destination of contraband", "place where contraband passed through", "location where contraband was seized" and so on.
[0147] Complex network analysis tools such as NetworkX can be used to construct a corresponding complex network based on the geographic feature information involved in the case process characteristics corresponding to the target case type.
[0148] Among them, the nodes of the complex network can represent the locations involved in the case process, such as the geographical locations corresponding to the "place where the contraband was sent from", "destination of the contraband", "place where the contraband passed through", and "location where the contraband was seized" in a contraband case.
[0149] The edges of a complex network represent the path between two locations during the case process, such as the "Spatial Transportation Path of Contraband Involved in the Case" in the contraband case, which shows the transportation route of contraband between locations. A complex network can also be an undirected network, that is, the edges in a complex network can be edges without directions. A complex network can be a directed network, that is, the edges in a complex network can be edges with directions.
[0150] Figure 3 A schematic diagram of a complex network according to an embodiment of the present disclosure is shown. For a target case type in a certain area, a corresponding complex network can be constructed based on the geographical feature information involved in the case process characteristics corresponding to the target case type. Figure 3 As shown in Figure 2, there may or may not be edges between nodes. Figure 3 As an example, the edges between nodes may be directed (the arrows in the figure indicate the direction of the edges).
[0151] Taking the case of contraband as an example, if the geographical feature information involved in a case process feature indicates that the "place where contraband is sent from", "destination of contraband", "place where contraband is passed through", and "location where contraband is seized" are A, C, B, and E respectively, and the "spatial transportation path of the contraband involved in the case" is A-B-C (that is, the transportation path of the contraband passes through A, B, and C respectively), then the complex network may include nodes corresponding to the four locations A, C, B, and E, and A and B, and B and C are connected by an edge respectively. If the complex network is a directed network, the directions of the two edges are from A to B and from B to C respectively, to reflect the transportation path during the transportation of contraband.
[0152] Corresponding complex networks can be generated for different case types. In the process of analyzing complex networks, the embodiments of the present disclosure can analyze the complex networks corresponding to each case type to obtain key locations and key paths in the case process under the case type. The complex network analysis results can include key locations and key paths in the case process. Key locations can be locations corresponding to nodes that have an important impact on the topological structure of the complex network, and key paths can be paths corresponding to edges that have an important impact on the topological structure of the complex network.
[0153] In the process of analyzing the complex network of at least one case type and obtaining the complex network analysis results, you can:
[0154] The node centrality is calculated for each node in the complex network of the target case type, and the locations corresponding to the nodes whose node centrality is greater than the first preset threshold are taken as key locations; the edge centrality is calculated for each edge in the complex network of the target case type, and the paths corresponding to the edges whose edge centrality is greater than the second preset threshold are taken as key paths.
[0155] The first preset threshold and the second preset threshold can be preset based on needs. There can be one or more key locations, and there can be one or more key paths.
[0156] The node centrality can be determined based on the number of edges between the node and other adjacent nodes. For any node in a complex network, the more edges there are between the node and other adjacent nodes, the greater the node centrality corresponding to the node; otherwise, the smaller the node centrality corresponding to the node. An example of a method for calculating the node centrality can also be found in formula (1):
[0157]
[0158] Among them, i and j can represent the identifiers of different nodes in the complex network. i It can represent the i-th node in a complex network.D (N i ) can represent node N i The corresponding node centrality. N can represent the total number of nodes in the complex network. ij For node N i and node N j The corresponding adjacency matrix, if the node N i and node N j There is an edge between them, then A ij =1; otherwise A ij =0.
[0159] Edge centrality can be determined based on the proportion of paths passing through the edge in the shortest paths between all pairs of nodes in the complex network. For any edge in a complex network, if the proportion of paths passing through the edge in the shortest paths between all pairs of nodes in the complex network is larger, the edge centrality corresponding to the edge is larger; otherwise, the edge centrality corresponding to the edge is smaller. An example of a method for calculating edge centrality can also be found in formula (2):
[0160]
[0161] Among them, s, t, and v can represent different nodes in the complex network respectively. e can represent any edge in the complex network, and c B (e) can represent the edge centrality corresponding to edge e. st It can represent the number of all shortest paths from node s to node t, σ st (e) represents the number of paths passing through edge e among all the shortest paths from node s to node t.
[0162] Targeted prevention and control measures can be generated for the key locations and key paths identified in the complex network of the target case type. For example, if the key locations are determined to be City A and City B, and the key path is from City A to City C for the complex network of the target case type, the generated targeted prevention and control measures may include strengthening inspections and investigations related to the target case type in City A and City B, strengthening traffic control on the path from City A to City C, or setting up roadblocks at the intersection of the path from City A to City C, and conducting random inspections of relevant transport vehicles related to the target case type.
[0163] Corresponding complex networks can be generated for different case types. The disclosed embodiments can also perform overall network analysis on the complex networks of each case type. For example, the network density and average network degree of each type of complex network can be determined, and the case types corresponding to the complex networks with network density and / or average network degree greater than a preset threshold can be selected as key case types. The complex network analysis results can also include key case types, and targeted prevention and control measures can be generated based on key case types, such as strengthening inspections and investigations related to key case types.
[0164] The network density can represent the degree of connection between the nodes in the complex network. The greater the network density, the more likely the corresponding complex network is to be a complex network of the key case type we are concerned about. An example of a calculation method for network density can be found in formula (3):
[0165]
[0166] Wherein, G may represent a complex network. d(G) may represent the network density of the complex network G. M may represent the number of edges in the complex network G. N may represent the number of nodes in the complex network G.
[0167] The average network degree can represent the average connection density of a single node in a complex network. The larger the average network degree, the more likely the corresponding complex network is to be a complex network of the key case type we are concerned about. An example of a method for calculating the average network degree can be found in formula (4):
[0168]
[0169] Wherein, G may represent a complex network. k may represent the average network degree of the complex network. M may represent the number of edges in the complex network G. N may represent the number of nodes in the complex network.
[0170] The disclosed embodiment can also perform statistical analysis based on case feature information extracted by the large language model to obtain statistical analysis results. The statistical analysis process can be performed by Figure 1 The statistical analysis module is implemented in .
[0171] In step S203, it is also possible to:
[0172] Statistical analysis is performed based on the characteristics of case personnel and case process to obtain statistical analysis results.
[0173] By conducting statistical analysis on the characteristics of case personnel and case process, we can obtain statistical analysis results such as statistical distribution results and correlation analysis results. We can fully and multi-angle utilize the large amount of case feature information obtained by processing the large language model, making the analysis of the case flexible and the analysis results more complete, which plays an important role in revealing the key development factors of crime.
[0174] Among them, when statistically analyzing the characteristics of case personnel, some or all of the characteristics of case personnel can be selected for statistical analysis as needed, for example, only the age of case personnel can be statistically analyzed, or only the correlation between the age and education level of case personnel can be statistically analyzed. Similarly, when statistically analyzing the characteristics of case process, all or part of the characteristics of case process can be selected for statistical analysis as needed, for example, only the behavioral characteristic information involved in the case process can be statistically analyzed, or only the information of items involved in the case process can be statistically analyzed.
[0175] The statistical analysis results may include any one or more of the statistical distribution results and the correlation analysis results. The embodiments of the present disclosure do not limit the distribution analysis and the correlation analysis methods, which may be implemented based on existing technologies.
[0176] Among them, the statistical distribution results can represent the distribution of any feature in dimensions such as time and quantity. For example, the frequency statistics (such as quantity distribution, etc.) of a single feature in the case personnel characteristics and case process characteristics can be performed to obtain the statistical distribution results. The correlation analysis results can represent the correlation between two or more features. For example, the correlation between the case personnel characteristics and two or more features in the case process characteristics can be analyzed. For example, the correlation between the age of the case personnel and the behavioral characteristic information involved in the case process, the correlation between the behavioral characteristic information involved in the case process and the division of labor information of each case personnel in the case process, etc. can be analyzed. The results obtained by analyzing the correlation between two or more features can be used as correlation analysis results.
[0177] The above statistical distribution results and correlation distribution results can be presented in the form of text descriptions, tables, graphs, etc., and the present disclosure does not limit this. Figure 4 A schematic diagram showing the statistical analysis results according to an embodiment of the present disclosure. Figure 4 As shown in the figure, for the behavior characteristic information involved in the case process of contraband crime in a certain city, such as the means of transport for contraband transportation, statistical analysis can be performed to obtain the quantitative distribution results of various modes of transportation. Among them, it can be seen that the number of self-driving cars as the means of transport for contraband in contraband crimes is the largest.
[0178] Statistical distribution results and correlation analysis results can be used to generate targeted prevention and control measures. For example, frequent features in statistical distribution results (such as Figure 4 Specific prevention and control strategies can be specified for self-driving cars that are used as the most common means of transporting contraband, such as strengthening the inspection and investigation of contraband-related crimes in the control of self-driving cars. Targeted prevention and control measures can also be generated for potential risk points indicated in the correlation analysis results. For example, if the correlation analysis results show that people in a specific profession are highly correlated with the amount of financial crimes, the generated targeted prevention and control measures can strengthen the review or education related to financial crimes for people in specific professions.
[0179] Figure 5 The overall flow chart of the case analysis method according to the embodiment of the present disclosure is shown. Figure 5 As shown, in the case analysis method of the embodiment of the present disclosure, a prompt text can be set for the case type, and the text of the case corresponding to the case type can be obtained and input into the data preprocessing module. The case personnel information and case process information obtained by extracting the text of the case in the data preprocessing module are fused with the prompt text to construct a preprocessed text. The preprocessed text is input into the large language model to extract the features therein to obtain the case personnel features and the case process features. The case process features can be input into the complex network analysis module to construct a complex network, perform complex network analysis, and obtain complex network analysis results; the case personnel features and case process features are input into the statistical analysis module for statistical analysis to obtain statistical analysis results. The complex network analysis results output by the complex network analysis module and the statistical analysis results output by the statistical analysis module are used as the final analysis results.
[0180] Figure 6 FIG. 2 shows a structural diagram of a case analysis device according to an embodiment of the present disclosure. Figure 6 As shown, the device comprises:
[0181] A preprocessing module 601 is used to preprocess the text of at least one case of the target case type to obtain a preprocessed text, where the text of the case includes a judgment document;
[0182] A determination module 602 is used to input the preprocessed text into a large language model to obtain case feature information, where the case feature information includes case personnel features and case process features;
[0183] The analysis module 603 is used to perform analysis based on case feature information to obtain analysis results.
[0184] In a possible implementation, the preprocessing module 601 is used to:
[0185] Regularized extraction of case personnel information and case process information in the text of at least one case of the target case type to obtain a regularized extracted text;
[0186] Determine the prompt text based on the target case type, and use the prompt text to guide the large language model to extract case feature information;
[0187] The regularized extracted text and the prompt text are combined to obtain the preprocessed text.
[0188] In a possible implementation, regularized extraction is performed on case personnel information and case process information in the text of at least one case of the target case type to obtain regularized extracted text, including:
[0189] The preset regularized expressions are used to match the text including the case personnel information and the text including the case process information in the case text respectively, and the text including the case personnel information and the text including the case process information are used as the regularized extracted text.
[0190] In a possible implementation, the determining module 602 is configured to:
[0191] The preprocessed text is input into the large language model, so that the large language model performs feature extraction on the regularized extracted text based on the prompt text in the preprocessed text to obtain case feature information;
[0192] The prompt text includes any one or more of text indicating a feature extraction type, text indicating a feature extraction range, and text indicating a feature extraction format.
[0193] In one possible implementation, case personnel characteristics and case process characteristics are determined based on the target case type;
[0194] Among them, the case person characteristics include any one or more of the case person's age, education level, occupation, and address;
[0195] The case process characteristics include any one or more of the geographical feature information involved in the case process, the behavioral feature information involved in the case process, the item information involved in the case process, and the division of labor information of each case personnel in the case process.
[0196] In a possible implementation, the analysis result includes a complex network analysis result, and the analysis module 603 is used to:
[0197] Based on the geographical feature information involved in the case process characteristics, a complex network of the target case type is constructed. The nodes of the complex network represent the locations involved in the case process, and the edges of the complex network represent the migration path between two locations in the case process.
[0198] Analyze the complex network of at least one case type to obtain complex network analysis results.
[0199] In a possible implementation, the complex network analysis results include key locations and key paths in the case process. The complex network of at least one case type is analyzed to obtain the complex network analysis results, including:
[0200] Calculate the node centrality of each node in the complex network of the target case type, and take the location corresponding to the node whose node centrality is greater than a first preset threshold as the key location;
[0201] The edge centrality is calculated for each edge in the complex network of the target case type, and the path corresponding to the edge whose edge centrality is greater than a second preset threshold is taken as the key path.
[0202] In one possible implementation, the node centrality is determined based on the number of edges that exist between the node and other adjacent nodes;
[0203] Edge centrality is determined based on the proportion of the shortest paths between all node pairs in a complex network that pass through the edge.
[0204] In a possible implementation, the analysis result includes a statistical analysis result, and the analysis module 603 is used to:
[0205] Statistical analysis is performed based on case personnel characteristics and case process characteristics to obtain statistical analysis results, which include any one or more of statistical distribution results and correlation analysis results.
[0206] According to the embodiment of the present disclosure, by preprocessing the text of at least one case of the target case type including the judgment document, a preprocessed text is obtained, the preprocessed text is input into the large language model, and the case feature information including the case personnel features and the case process features is obtained, and the case feature information is analyzed based on the case feature information to obtain the analysis result, so that the large language model can be used to quickly, accurately and flexibly extract the case personnel features and the case process features from the judgment document, so that the scheme of the present disclosure has higher flexibility and extraction accuracy, and is more economical and efficient than the manual method. At the same time, based on the more accurate and rich features extracted, a more in-depth analysis result can be obtained, which is more accurate and efficient than the method based on the deep learning model, and the training cost is low. The analysis results of the embodiment of the present disclosure can be used to determine the targeted prevention and control measures related to the case personnel and the case process or as reference information for case detection, so as to promote the improvement of case analysis and the progress of related prevention and control governance.
[0207] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0208] The embodiment of the present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0209] The disclosed embodiment also proposes a case analysis device, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0210] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0211] Figure 7 1 is a block diagram of a device 1900 for case analysis according to an exemplary embodiment. For example, the device 1900 may be provided as a server or a terminal device. Figure 7 , the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0212] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2000. TM , MacOS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.
[0213] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to perform the above method.
[0214] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0215] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0216] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0217] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0218] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0219] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0220] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0221] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0222] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A case analysis method, characterized in that: The method comprises: Preprocessing the text of at least one case of the target case type to obtain a preprocessed text, wherein the text of the case includes a judgment document; Inputting the preprocessed text into a large language model to obtain case feature information, wherein the case feature information includes case personnel features and case process features; An analysis is performed based on the case characteristic information to obtain an analysis result.
2. The method according to claim 1, characterized in that The preprocessing of the text of at least one case of the target case type to obtain the preprocessed text includes: Regularized extraction of case personnel information and case process information in the text of at least one case of the target case type to obtain a regularized extracted text; Determine a prompt text based on the target case type, wherein the prompt text is used to guide the large language model to extract case feature information; The regularized extracted text and the prompt text are combined to obtain the preprocessed text.
3. The method according to claim 2, characterized in that The step of performing regularized extraction on case personnel information and case process information in the text of at least one case of the target case type to obtain regularized extracted text includes: The preset regularized expressions are used to respectively match the text including the case personnel information and the text including the case process information in the text of the case, and the text including the case personnel information and the text including the case process information are used as the regularized extracted text.
4. The method according to claim 2, characterized in that: The step of inputting the preprocessed text into a large language model to obtain case feature information includes: Inputting the preprocessed text into the large language model, so that the large language model performs feature extraction on the regularized extracted text based on the prompt text in the preprocessed text to obtain case feature information; The prompt text includes any one or more of text indicating a feature extraction type, text indicating a feature extraction range, and text indicating a feature extraction format.
5. The method according to claim 1, characterized in that The case personnel characteristics and the case process characteristics are determined based on the target case type; The case personnel characteristics include any one or more of the case personnel's age, education level, occupation, and address; The case process characteristics include any one or more of the geographical feature information involved in the case process, the behavioral feature information involved in the case process, the item information involved in the case process, and the division of labor information of each case personnel in the case process.
6. The method according to claim 1, characterized in that The analysis result includes a complex network analysis result, and the analysis is performed based on the case feature information to obtain the analysis result, including: Based on the geographical feature information involved in the case process characteristics, a complex network of the target case type is constructed, wherein the nodes of the complex network represent the locations involved in the case process, and the edges of the complex network represent the migration path between two locations in the case process; Analyze the complex network of at least one case type to obtain the complex network analysis result.
7. The method according to claim 6, characterized in that The complex network analysis result includes key locations and key paths in the case process. The complex network analysis result obtained by analyzing the complex network of at least one case type includes: Calculating the node centrality of each node in the complex network of the target case type, and taking the location corresponding to the node whose node centrality is greater than a first preset threshold as the key location; The edge centrality is calculated for each edge in the complex network of the target case type, and the path corresponding to the edge whose edge centrality is greater than a second preset threshold is used as the key path.
8. The method according to claim 7, characterized in that The node centrality is determined based on the number of edges between the node and other adjacent nodes; The edge centrality is determined based on the proportion of the paths passing through the edge in all the shortest paths between all node pairs in the complex network.
9. The method according to claim 1, characterized in that: The analysis result includes a statistical analysis result, and the analysis result obtained by performing the analysis based on the case feature information includes: A statistical analysis is performed based on the case personnel characteristics and the case process characteristics to obtain a statistical analysis result, wherein the statistical analysis result includes any one or more of a statistical distribution result and a correlation analysis result.
10. A case analysis device, characterized in that: The device comprises: A preprocessing module, used to preprocess the text of at least one case of the target case type to obtain a preprocessed text, wherein the text of the case includes a judgment document; A determination module, used for inputting the preprocessed text into a large language model to obtain case feature information, wherein the case feature information includes case personnel features and case process features; The analysis module is used to perform analysis based on the case feature information to obtain analysis results.
11. A case analysis device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 9 when executing the instructions stored in the memory.
12. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.