Multi-modal interactive law enforcement document generation method and device, electronic equipment and storage medium

By building a law enforcement vector database and on-site inquiry model, law enforcement documents are automatically generated, and the problem of low efficiency in document production is solved, and the intelligent generation of documents and law enforcement efficiency is improved.

CN120562382APending Publication Date: 2025-08-29JIANGSU HONGXIN SYST INTEGRATION
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510417705.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The production efficiency of law enforcement documents and the lack of intelligent and digital management have made it difficult to achieve standardization and unified cultural law enforcement cases.

Method used

A multimodal interactive law enforcement document generation method is adopted, and a law enforcement vector database is constructed using the convolutional neural network RNN ​​and Word2Vec model. The on-site inquiry model is automatically generated and the content is reviewed to generate an inquiry document.

Benefits of technology

It improves the efficiency and accuracy of document production, assists law enforcement personnel in completing document generation, and improves law enforcement efficiency and workflow optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120562382A_ABST
    Figure CN120562382A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal interactive law enforcement document generation method and device, electronic equipment and a storage medium, and relates to the field of large model text generation. The method comprises the following steps: collecting a law enforcement structured data set, and extracting a text data set from the law enforcement structured data set; analyzing the text data set, extracting word vectors and grammar structure information from the text data set, integrating the word vectors and the grammar structure information through a weighted average strategy, and constructing a law enforcement vector database; performing fine tuning training on the open-source large model by utilizing the law enforcement vector database to obtain an on-site inquiry model; a cause of action is input into the on-site inquiry model, the on-site inquiry model automatically generates a sniffing outline according to the cause of action, sniffing guidance is provided for law enforcement officers, and inquiry records are generated; and according to the generated inquiry record document, content auditing and checking are carried out, and a preview file is generated after auditing and checking. Complex case information can be efficiently processed, the document quality and the processing speed are improved, the overall law enforcement efficiency is improved, and efficient workflow and system optimization are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large-model text generation, and in particular to a method, device, electronic device and storage medium for generating multimodal interactive law enforcement documents. Background Art

[0002] In today's era of rapid advancements in artificial intelligence, large-scale model technology, a cutting-edge advancement in deep learning, has demonstrated its significant potential and application value across multiple fields. In particular, in natural language processing (NLP), large-scale model technology, through training on large datasets, enables a deep understanding and simulation of human language, enabling efficient and accurate text generation. Models such as the Transformer and BERT, with their billions or even hundreds of billions of parameters, are capable of processing complex linguistic structures while also capturing subtle nuances and deep semantic relationships within language, providing a solid technical foundation for intelligent text generation. Simultaneously, researchers are working to improve the interpretability and real-time update capabilities of large models to ensure their accuracy and transparency in real-world applications. These advances have enabled large models to demonstrate enhanced capabilities and broad application potential in a variety of fields, including automated content generation, intelligent recommendations, language recognition, and data analysis.

[0003] Cultural law enforcement documents are a crucial component of the case process, and their quality directly impacts the authority and effectiveness of cultural law enforcement. From case filing to case closure, administrative penalty cases may involve the completion of numerous types of documents. This process consumes significant energy and time for law enforcement personnel, resulting in low efficiency in document preparation and hindering the standardization and unification of cultural law enforcement cases. Furthermore, the repetitive organization and archiving of various documents lacks intelligent application and digital management tools.

[0004] Big model technology has revolutionized text generation and holds broad application prospects in cultural market management and comprehensive law enforcement. By incorporating big model technology, law enforcement personnel can be assisted in document generation, enabling intelligent analysis of behavior, intelligent citation of enforcement evidence, intelligent assessment of violations, and intelligent summary of case key points. Furthermore, the construction of text vector databases for law enforcement is crucial for training multimodal big models. Summary of the Invention

[0005] Purpose of the invention: To propose a multimodal interactive law enforcement document generation method, device, electronic device and storage medium, which uses AI algorithms to effectively solve the problem of difficulty in producing law enforcement documents, comply with legal norms and actual law enforcement needs, ensure the fairness and transparency of the law enforcement process, and have important practical significance for improving law enforcement efficiency and carrying out smart supervision.

[0006] A first aspect of the present invention provides a method for generating a multimodal interactive law enforcement document, comprising the following steps: Collecting a law enforcement structured data set and extracting a text data set from it; the text data set includes law enforcement text data and basic data; The text data set is parsed using a convolutional neural network (RNN) and a Word2Vec model, word vectors and grammatical structure information are extracted from the text data set, and the word vectors and grammatical structure information are integrated through a weighted average strategy to construct a law enforcement vector database; The law enforcement vector database is used to fine-tune the convolutional neural network (RNN) and Word2Vec model to obtain an on-site inquiry model; Inputting the cause of action into the on-site inquiry model, the on-site inquiry model automatically generates an inquiry outline based on the cause of action, provides inquiry guidance to law enforcement personnel, and generates an inquiry record document; Based on the generated interrogation record document, the content is reviewed and verified, and a preview file is generated after the review and verification.

[0007] In a further embodiment of the first aspect, a convolutional neural network (RNN) is used to extract grammatical structure information of the text data set; The Word2Vec model is used to extract word vectors of the text data set.

[0008] In a further embodiment of the first aspect, the joint loss function of the convolutional neural network RNN ​​and Word2Vec model is :

[0009] in, Maximize the similarity between the text vector and the target word vector; Optimizing similarity for negative sampling; 、 are their respective weights;

[0010]

[0011] Where, and are the word vectors of the target word and the context word respectively; K is the number of negative samples; is a negative sample; is a text vector; is the target word vector.

[0012] In a further embodiment of the first aspect, the grammatical structure information extracted by the convolutional neural network RNN ​​is And the word vector extracted by the Word2Vec model The feature fusion is performed through the weighted average strategy, and the weighted average formula used is:

[0013] Where, It is a hyperparameter that controls the output weight ratio.

[0014] In a further embodiment of the first aspect, the law enforcement vector database is used to perform Lora fine-tuning on the convolutional neural network RNN ​​and Word2Vec model, and the gradient update method in Lora fine-tuning is as follows: , , ,

[0015] Where B is the initialization zero matrix; A is the random Gaussian initialization matrix; Represents the weight matrix of the original pre-trained model; Represents the updated weight; B is a matrix of shape d×r; A is a matrix of shape r×k; d, r, k represent the matrix latitude; represents the rank of the low-rank matrix.

[0016] In a further embodiment of the first aspect, during training, Frozen, the gradient is not updated, and the random Gaussian initialization matrix A and the initialization zero matrix B contain trainable parameters, then the weight update is as follows:

[0017] Where h represents the hidden layer output; x represents the input vector.

[0018] In a further embodiment of the first aspect, intelligent assisted law enforcement inquiry is implemented throughout the entire process before, during, and after the inquiry based on the on-site inquiry model; Before questioning, generate an inquiry outline based on the case to provide on-site law enforcement officers with inquiry ideas; During the interrogation, the voice recognition and transcription function is used to call the on-site interrogation model to provide law enforcement officers with dynamic interrogation guidance and determine the completeness of the case inquiry in real time; After the questioning, the on-site questioning voice transcription data will be saved, and the on-site questioning model will be used to generate the investigation and questioning record into an investigation and questioning transcript.

[0019] A second aspect of the present invention provides a multimodal interactive law enforcement document generation device, comprising: An extraction module is used to collect law enforcement structured data sets and extract text data sets therefrom; the text data sets include law enforcement text data and basic data; A parsing module, composed of a convolutional neural network (RNN) and a Word2Vec model, is used to parse the text data set, extract word vectors and grammatical structure information, and integrate them through a weighted average strategy to construct a law enforcement vector database; A model building module, which uses the law enforcement vector database to fine-tune the convolutional neural network RNN ​​and Word2Vec model to obtain an on-site inquiry model; A document generation module is used to input the case cause into the on-site inquiry model, and the on-site inquiry model automatically generates an inquiry outline based on the case cause, provides inquiry guidance to law enforcement personnel, and generates an inquiry record document; The preview module is used to review and verify the content of the generated interrogation record document, and generate a preview file after review and verification.

[0020] The third aspect of the present invention proposes an electronic device, which includes a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the multimodal interactive law enforcement document generation method as described in the first aspect.

[0021] The fourth aspect of the present invention proposes a computer-readable storage medium, which stores at least one executable instruction. When the executable instruction is run on an electronic device, the electronic device executes the multimodal interactive law enforcement document generation method as described in the first aspect.

[0022] Beneficial Effects: Based on multimodal large-scale model technology, a real-time on-site inquiry assistant is constructed to assist law enforcement officers in conducting on-site inquiries. This invention can efficiently process complex case information, improve document quality and processing speed, enhance overall law enforcement efficiency, and achieve efficient workflow and system optimization. Leveraging the model's text generation and semantic understanding capabilities, it automatically generates different types of law enforcement documents, including case information, case elements, clause basis, and judgment results. This further assists law enforcement officers in improving the efficiency of document processing, promoting the application of smart law enforcement, and improving the quality and effectiveness of law enforcement. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Flowchart of the method for generating multimodal interactive law enforcement documents.

[0024] Figure 2 This is a flowchart of the law enforcement agent process. DETAILED DESCRIPTION

[0025] In the following description, numerous specific details are provided to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced without one or more of these details. In other instances, certain technical features well known in the art have not been described to avoid confusion with the present invention.

[0026] Before describing the following embodiments, some terms that will appear in the text are first explained: Agents: In the field of artificial intelligence, agents are entities capable of intelligently performing tasks. They can perceive their environment, take autonomous actions to achieve goals, and improve their performance through learning or acquiring knowledge. In the context of large models, agents can be understood as intelligent entities capable of autonomously understanding, planning, and executing complex tasks.

[0027] LoRA fine-tuning: This is an efficient model fine-tuning method that adjusts the weights of a pre-trained model by adding a low-rank matrix. This method can achieve customized adaptation of the model while reducing computational and storage overhead, thereby improving the model's performance on specific tasks.

[0028] Prompt engineering involves designing and optimizing the prompts fed into large language models to improve the quality and relevance of the responses generated by the models. By crafting effective question descriptions or instructions, the models can better understand user intent and produce more accurate and useful output.

[0029] This invention proposes a multimodal interactive law enforcement document generation method based on RNN-W2V vectorization. Based on a data set consisting of professional law enforcement data such as historical cases, laws and regulations, this method adopts a joint training method of RNN and Word2Vec algorithms, denoted as RNN-W2V, to build a professional law enforcement vector knowledge base; on this basis, a suitable basic model is selected for training to obtain a large model exclusive to law enforcement, and through user feedback, a reinforcement learning reward model is used to continuously optimize and upgrade the large model capabilities; thereby utilizing text generation capabilities, semantic understanding capabilities, content discrimination capabilities, etc., with the help of multimodal interactive intelligent agents to assist law enforcement personnel in completing intelligent document generation and improve office efficiency and accuracy. The main purposes can be divided into three points: RNN-W2V Data Vectorization: This approach designs a joint training algorithm framework that combines the Skip-Gram framework from the Word2Vec algorithm with a convolutional neural network (RNN). The Skip-Gram framework captures the global semantic information of words, while the RNN extracts local grammatical and contextual features from the text. This method can enrich and comprehensively generate word vectors, building a higher-quality text vector database.

[0030] Build an on-site inquiry assistant: Based on the basic information of the case, provide real-time dynamic auxiliary inquiries, guide the inquiry process, and provide full-link on-site inquiry assistance services before, during, and after the inquiry, so as to achieve the purpose of guiding newcomers on law enforcement process issues and providing more efficient support to experienced law enforcement personnel.

[0031] Build an assistant for generating law enforcement documents for the cultural market: Based on law enforcement datasets, train a large law enforcement-specific model to assist in matching causes of law enforcement cases, citing legal clauses, generating penalty content, etc., enabling the intelligent generation of case-assisted law enforcement documents, and significantly improving the efficiency and accuracy of document production.

[0032] This invention proposes a method for generating multimodal interactive law enforcement documents based on RNN-W2V vectorization. First, the RNN-W2V algorithm is used to build a law enforcement vector database, and a large law enforcement-specific model is trained based on the data set; then, the Agent intelligent body is used to decompose the document production task, and different large model capabilities and tools are called to complete the generation of corresponding field content information; then, the generated content needs to be reviewed by the system and manually reviewed and verified, including sensitive words, typos, etc.; finally, according to the format of the document, the generation preview of various documents is completed to ensure the standardization of the format and the compliance of the content. The specific method flow diagram is as follows. Figure 1 As shown, the specific technical implementation steps are as follows: Step 1: Data Processing. Based on the collection and aggregation of law enforcement data from multiple channels, standardized operations such as cleaning, extraction, and text segmentation are performed on structured law enforcement datasets. For unstructured law enforcement datasets, OCR technology is used to extract text from images, perform data cleaning, and segmentation. The processed data is then standardized and normalized, and finally classified and stored.

[0033] Step 2: Constructing a law enforcement vector database. Using the RNN-W2V joint training algorithm, we integrate their outputs through a weighted average strategy. The resulting word vectors can simultaneously incorporate both semantic and grammatical information about the word, making the text vectors more expressive.

[0034] The RNN-W2V joint loss function is , To maximize the similarity between the text vector and the target word vector, Optimizing similarity for negative sampling:

[0035]

[0036]

[0037] in and are the word vectors of the target word and the context word respectively, K is the number of negative samples, is a negative sample, is a text vector, is the target word vector.

[0038] The word vectors obtained by RNN and Word2Vec algorithm are fused through the weighted average strategy. The weighted average formula used is:

[0039] in It is a hyperparameter that controls the output weight ratio.

[0040] Through the above joint training, by fully leveraging the advantages of the two different models, we can simultaneously utilize word-level contextual information and text-level local features to improve the quality of the final word vector and build a vector database for the professional field of law enforcement.

[0041] Step 3: Training the large law enforcement model. Based on a dataset consisting of historical case files, laws and regulations, and other law enforcement texts, we selected a basic model for training and fine-tuned it using the LoRA method. We adjusted parameters such as the learning rate, number of training epochs, batch size, and LoRA rank to achieve optimal performance for the large model. The gradient update method in Lora fine-tuning is as follows: , , ,

[0042] Where B is the initialization zero matrix and A is the random Gaussian initialization matrix; During training, Frozen, the gradient is not updated, and A and B contain trainable parameters, then the weight update is as follows:

[0043] Step 4: Inquiry model training. Based on the historical "Investigation and Interrogation Records" and on-site inquiry audio as the basic data set, select the corresponding basic model for fine-tuning training, adjust the parameter settings, and compare the loss function gradient convergence to determine the optimal weight of the large model, thereby selecting the best parameter settings to ensure the performance of the large model. Step 5: Build an on-site inquiry assistant. Based on the inquiry big model, we will build an on-site inquiry assistant to realize intelligent assisted law enforcement inquiry in the entire process before, during, and after the inquiry. Before the inquiry, an inquiry outline will be generated according to the case cause to provide inquiry ideas for on-site law enforcement personnel. During the inquiry, through the voice recognition transcription function, the law enforcement inquiry big model will be called to provide dynamic inquiry guidance for law enforcement personnel and judge the completeness of the case inquiry in real time. After the inquiry, the on-site inquiry voice transcription data will be saved, and the text generation capability of the big model will be used to generate an investigation and inquiry record document from the investigation and inquiry record. Step 6: Agent. After obtaining the user's demand for law enforcement document production, the agent will simulate the human cognitive process and use its deep language cognitive reasoning and complex situational understanding capabilities to analyze the user's intentions. Based on the understood needs, it will plan the process of executing the task. Depending on the type of document, it will call different large model capabilities, prompt engineering design, and interactive tools to complete the document production task. For the specific process design, see Figure 2 .

[0044] Step 7: Model Optimization. The model's output for each task is scored and evaluated. Based on the scoring of the task content, targeted optimizations are performed, including adjustments to training data, optimization prompts, and post-processing, to further enhance the model's performance and accuracy. This feedback loop mechanism helps continuously improve the model's overall capabilities and enhance the effectiveness of assisting law enforcement document generation.

[0045] Step 8: Content Review and Verification. To further ensure the quality of law enforcement documents, we will conduct content review and verification on the text generated by the model, including the following three points: 1) Using the DFA (Deterministic Finite Automaton) algorithm to screen and filter sensitive words. When sensitive words are detected, the system can process them according to preset rules, such as replacing them with asterisks (*) or highlighting them to prompt corrections; 2) Automatic detection and correction of typos in law enforcement documents; 3) Accuracy verification and standardization of law enforcement terminology; Step 9: Document Preview. After review and verification, the field content generated by the large model is intelligently filled into the corresponding document format to generate a preview file. It also supports online editing, saving, and printing of files.

[0046] Based on the technical process of the multimodal interactive law enforcement document generation method disclosed in the above embodiment, a multimodal interactive law enforcement document generation device can be developed, which consists of an extraction module, a parsing module, a model building module, a document generation module, and a preview module. The extraction module is used to collect law enforcement structured data sets and extract a text data set from it; the text data set includes law enforcement text data and basic data. The parsing module is used to parse the text data set, extract word vectors and grammatical structure information from it, and integrate them through a weighted average strategy to construct a law enforcement vector database. The model building module uses the law enforcement vector database to fine-tune and train the open source large model to obtain an on-site inquiry model. The document generation module is used to input the case cause into the on-site inquiry model, and the on-site inquiry model automatically generates an inquiry outline based on the case cause, provides inquiry guidance to law enforcement personnel, and generates an inquiry record document. The preview module is used to perform content review and verification based on the generated inquiry record document, and generate a preview file after review and verification.

[0047] The multimodal interactive law enforcement document generation method and device disclosed in the above embodiments have the following technical effects: 1. Vector database based on the RNN-W2V algorithm: RNN and Word2Vec are jointly trained. By combining RNN's ability to capture long-term contextual dependencies with Skip-Gram's strengths in modeling word semantic relationships, richer text vectors are generated. This joint training improves the model's understanding of long texts and its expression of fine-grained semantics, enhancing the accuracy and robustness of text representation, thereby building a text vector database that is more conducive to training large models.

[0048] 2. Build an on-site inquiry assistant: Based on multimodal big model technology, a real-time on-site inquiry assistant is built to assist law enforcement officers in conducting on-site inquiries. This includes three key features: 1) Real-time speech transcription: Leveraging the big model's speech recognition capabilities, real-time inquiry recordings are converted into text by role, reducing manual data entry time and ensuring the accuracy of the records; 2) Recommended question list: Based on the real-time responses of the person being questioned, the dynamic inquiry big model is invoked to provide law enforcement officers with a list of recommended questions, dynamically guiding the inquiry process; 3) Inquiry progress assessment: Based on the real-time inquiry content, the completeness of key investigative dimensions of the case is determined, and the progress of the case inquiry is evaluated.

[0049] 3. Law Enforcement Agent: Leveraging the agent's autonomy, adaptability, and intelligent decision-making capabilities, this system integrates a dedicated law enforcement model, the Law Enforcement Prompt project, and other law enforcement tools. Based on evolving user needs and information, it dynamically adjusts its behavior and strategies to adapt to different scenarios and task requirements. This system efficiently handles complex case information, improves document quality and processing speed, enhances overall law enforcement efficiency, and achieves efficient workflow and system optimization.

[0050] 4. Intelligent Generation of Law Enforcement Documents: Based on over 1,100 different case causes and case file information, we continuously train and optimize a large law enforcement-specific model. Leveraging the model's text generation and semantic understanding capabilities, we automatically generate different types of law enforcement documents, including case information, case elements, clause basis, and judgment results. This further assists law enforcement personnel in improving document processing efficiency, promoting the application of smart law enforcement, and enhancing the quality and effectiveness of law enforcement.

[0051] The technical process of the multimodal interactive law enforcement document generation method disclosed in the above embodiment can be implemented in whole or in part through software, hardware, firmware or any other combination.

[0052] When implemented in hardware, the aforementioned embodiments can all or partly compile the operating logic and computational processes into software and then run them on an electronic device. The electronic device includes a processor, memory, a communication interface, and a communication bus. The processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction, which causes the processor to execute the technical process of the multimodal interactive law enforcement document generation method disclosed in the aforementioned embodiments.

[0053] When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. If the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be embodied in the form of a software product, which is essentially or contributes to the relevant technology. The software product is stored in a storage medium and includes several instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0054] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes may be made to it in form and detail without departing from the spirit and scope of the present invention as defined in the appended claims.

Claims

1. A multimodal interactive law enforcement document generation method, characterized in that: The steps include: Collecting a law enforcement structured data set and extracting a text data set from it; the text data set includes law enforcement text data and basic data; The text data set is parsed using a convolutional neural network (RNN) and a Word2Vec model, word vectors and grammatical structure information are extracted from the text data set, and the word vectors and grammatical structure information are integrated through a weighted average strategy to construct a law enforcement vector database; The law enforcement vector database is used to fine-tune the convolutional neural network (RNN) and Word2Vec model to obtain an on-site inquiry model; Inputting the cause of action into the on-site inquiry model, the on-site inquiry model automatically generates an inquiry outline based on the cause of action, provides inquiry guidance to law enforcement personnel, and generates an inquiry record document; Based on the generated interrogation record document, the content is reviewed and verified, and a preview file is generated after the review and verification.

2. A multimodal interactive law enforcement document generation method according to claim 1, characterized in that: Using a convolutional neural network (RNN) to extract grammatical structure information from the text data set; The Word2Vec model is used to extract word vectors of the text data set.

3. A multimodal interactive law enforcement document generation method according to claim 2, characterized in that: The joint loss function of the convolutional neural network RNN ​​and Word2Vec model is L joint : L joint =λ1L RNN +λ2L W2V Among them, L RNN Maximize the similarity between the text vector and the target word vector; L W2V Optimize similarity for negative sampling; λ1 and λ2 are their respective weights; Where, and are the word vectors of the target word and the context word respectively; K is the number of negative samples; w i is a negative sample; v text is a text vector; is the target word vector.

4. A multimodal interactive law enforcement document generation method according to claim 2, characterized in that: The grammatical structure information v extracted by the convolutional neural network RNN RNN And the word vector v extracted by the Word2Vec model w2v The feature fusion is performed through the weighted average strategy, and the weighted average formula used is: v fin =αv RNN +(1-a)v w2v Where α∈[0, 1] is a hyperparameter that controls the output weight ratio.

5. A multimodal interactive law enforcement document generation method according to claim 1, characterized in that: The law enforcement vector database is used to perform Lora fine-tuning on the convolutional neural network RNN ​​and Word2Vec model. The gradient update method in Lora fine-tuning is as follows: W0+ΔW=W0+BA,B∈R d×r ,A∈R r×k ,r< <min(d,k) Where B is the initialization zero matrix; A is the random Gaussian initialization matrix; W0 represents the weight matrix of the original pre-trained model; ΔW represents the updated weight; B is a matrix of shape d×r; A is a matrix of shape r×k; d, r, and k represent the matrix latitude; min(d, k) represents the rank of the low-rank matrix.

6. A multimodal interactive law enforcement document generation method according to claim 5, characterized in that: During training, W0 is frozen and the gradient is not updated. The random Gaussian initialization matrix A and the initialization zero matrix B contain trainable parameters, and the weights are updated as follows: h=(W0+ΔW)x=(W0+BA)x=W0x+BAx Where h represents the hidden layer output; x represents the input vector.

7. A multimodal interactive law enforcement document generation method according to claim 1, characterized in that: Based on the on-site inquiry model, intelligent assisted law enforcement inquiry is realized throughout the entire process before, during and after the inquiry; Before questioning, generate an inquiry outline based on the case to provide on-site law enforcement officers with questioning ideas; During the interrogation, the voice recognition and transcription function is used to call the on-site interrogation model to provide law enforcement officers with dynamic interrogation guidance and determine the completeness of the case interrogation in real time; After the questioning, the on-site questioning voice transcription data will be saved, and the on-site questioning model will be used to generate the investigation and questioning record into an investigation and questioning transcript.

8. A multimodal interactive law enforcement document generation device, characterized in that: include: An extraction module is used to collect law enforcement structured data sets and extract text data sets therefrom; the text data sets include law enforcement text data and basic data; A parsing module, composed of a convolutional neural network (RNN) and a Word2Vec model, is used to parse the text data set, extract word vectors and grammatical structure information, and integrate them through a weighted average strategy to construct a law enforcement vector database; A model building module, which uses the law enforcement vector database to fine-tune the convolutional neural network RNN ​​and Word2Vec model to obtain an on-site inquiry model; A document generation module is used to input the case cause into the on-site inquiry model, and the on-site inquiry model automatically generates an inquiry outline based on the case cause, provides inquiry guidance to law enforcement personnel, and generates an inquiry record document; The preview module is used to review and verify the content of the generated interrogation record document, and generate a preview file after review and verification.

9. An electronic device, characterized in that: include: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the multimodal interactive law enforcement document generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The storage medium stores at least one executable instruction, and when the executable instruction is executed on the electronic device, the electronic device executes the multimodal interactive law enforcement document generation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method and device for generating records based on large model and computer product

    CN121031541A

  • Method and device for generating a transcript based on a large model and computer product

    CN121031541B