Public security record making method based on AI multi-model cooperative processing and related equipment
By employing a multi-model collaborative processing method, this approach utilizes multimodal visual language models and speech AI models to generate customized transcript templates and verify their legality and compliance. This solves the problems of low efficiency and low accuracy in traditional transcript production, achieving efficient and accurate transcript production.
Patent Information
- Application Number
- CN202511654500.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-03-20
AI Technical Summary
Traditional police record taking relies on manual input by police officers, which is inefficient and prone to errors. Existing AI applications suffer from insufficient compliance checks, low voice recognition accuracy, and inadequate information.
A multi-model collaborative processing approach is adopted, which uses multimodal visual language models and speech AI models to convert evidence images and audio recordings into text. Combined with open-source large language models, customized transcript templates are generated and legality and compliance are verified to produce accurate transcript text.
It improved the efficiency and accuracy of record-keeping, reduced the workload of police officers, ensured the legality, compliance, and completeness of the records, and reduced the risk of a broken chain of evidence.
Smart Images

Figure CN121706751A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and related equipment for producing public security transcripts based on AI multi-model collaborative processing. Background Technology
[0002] In related technologies, traditional record-keeping mainly relies on police officers manually typing on a keyboard, recording simultaneously during questioning. For some officers unfamiliar with computer operation, this method is inefficient and carries the risk of recording errors. With the popularization of artificial intelligence, especially large language model technology, some public security units have begun to try applying it to record-keeping work. Currently, there are three main application directions: First, based on large language models and professional knowledge bases, completed records are checked for legality and compliance, prompting for supplementary questions or modifications. This method is a post-event remedy, not only unable to avoid rework but also potentially increasing the subsequent workload of police officers. Second, speech AI models are used to transcribe questioning content in real time to assist police officers in record-keeping, and large language models are used to prompt for missing questions. However, in areas where dialects are commonly used, the accuracy of speech recognition is low and its applicability is poor. Third, information about the parties involved in a case is collected through pre-filled forms, and then combined with large language models and knowledge bases to generate preliminary records. Police officers then supplement and improve the content based on this. In this method, if the initial information provided by the party involved is insufficient, it will significantly increase the workload of police officers in subsequent supplementary questioning.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main objective of this application is to propose a method and related equipment for producing public security records based on AI multi-model collaborative processing, which can effectively improve the efficiency and accuracy of public security record production.
[0005] To achieve the above objectives, one aspect of this application proposes a method for creating public security records based on AI multi-model collaborative processing, the method comprising the following steps: The system retrieves the case type, case category, case characteristic information, case evidence images, and case audio recordings input into a preset visual interface. The case evidence images are converted into text using a preset multimodal visual language model to obtain the first transcribed text; Based on the case type, case category, case feature information, and the first transcribed text, a customized transcript template is generated using a preset open-source large language model; The case recording material is converted into text using a preset speech AI model to obtain a second transcribed text; Based on the customized transcript template and the second transcribed text, a transcript to be verified is generated using the preset open-source large language model; Based on the public security legal knowledge base, the legality and compliance of the transcript to be verified are checked using the preset open-source large language model to obtain the target transcript text.
[0006] In some embodiments, the case audio recordings include recordings of staff repeating key information.
[0007] In some embodiments, the step of converting the case evidence images into text using a preset multimodal visual language model to obtain the first transcribed text includes: Analyze the data to be analyzed in the evidence images of the case; If the data to be analyzed does not include preset sensitive data, the case evidence images are converted into text using a preset multimodal visual language model to obtain the first transcribed text; If the data to be analyzed includes preset sensitive data, the case evidence image is converted into text using a preset multimodal visual language model to obtain the first text information to be processed; and the first text information to be processed is corrected using a preset open-source large language model to obtain the first transcribed text.
[0008] In some embodiments, the step of converting the case recording material into text using a preset speech AI model to obtain a second transcribed text includes: Analyze the language information in the audio recordings of the case; If the language information is a single language, the case audio material is converted into text using the Paraformer model to obtain the second transcribed text. If the language information is multilingual, the case audio material is converted into text using the Whisper model and the Paraformer model respectively to obtain the second text information to be processed; and the second text information to be processed is fused and logically corrected using a preset open-source large language model to obtain the second transcribed text.
[0009] In some embodiments, generating a customized transcript template based on the case type, the case category, the case feature information, and the first transcribed text using a preset open-source large language model includes: The first transcript template is obtained from the public security legal knowledge base based on the case type and the case category. Based on the case feature information and the first transcribed text, the target question is retrieved from the public security legal knowledge base; The information in the first transcript template is updated based on the target problem to obtain the customized transcript template.
[0010] In some embodiments, generating a transcript to be verified using the preset open-source large language model based on the customized transcript template and the second transcribed text includes: Each question to be matched in the customized transcript template is matched with the content of the second transcribed text, and the target content is extracted from the second transcribed text as the answer content of the question to be matched using the preset open-source large language model; All the questions to be matched and all the answers are recombined in a preset order to obtain the transcript to be verified.
[0011] In some embodiments, the method further includes the following steps: Obtain the inspection results of the target transcript text; If the inspection results determine that the target transcript contains content that does not meet the requirements, the target transcript will be adjusted.
[0012] To achieve the above objectives, another aspect of this application proposes a public security record-keeping device based on AI multi-model collaborative processing, the device comprising: The acquisition module is used to acquire case type, case category, case feature information, case evidence images, and case audio materials input from the preset visual interface; The first conversion module is used to convert the case evidence images into text using a preset multimodal visual language model to obtain the first transcribed text. The first generation module is used to generate a customized transcript template based on the case type, the case category, the case feature information and the first transcribed text, using a preset open-source large language model. The second conversion module is used to convert the case audio material into text using a preset speech AI model to obtain a second transcribed text. The second generation module is used to generate a transcript to be verified based on the customized transcript template and the second transcribed text, using the preset open-source large language model. The verification module is used to perform legality and compliance verification on the transcript to be verified based on the public security legal knowledge base and the preset open-source large language model to obtain the target transcript text.
[0013] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0014] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0015] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0016] The embodiments of this application include at least the following beneficial effects: This application provides a method and related equipment for producing public security records based on AI multi-model collaborative processing. This method involves inputting case type, case category, case feature information, case evidence images, and case audio recordings through a preset visual interface. Then, a preset multimodal visual language model is used to convert the case evidence images into text, obtaining a first transcribed text. Next, based on the case type, case category, case feature information, and the first transcribed text, a preset open-source large language model is used to generate a customized record template, ensuring that the customized record template conforms to the actual situation of the current case type and category. Then, a preset speech AI model is used to convert the case audio recordings into text, obtaining a second transcribed text. Based on the customized record template and the second transcribed text, a preset open-source large language model is used to generate a record to be verified, ensuring that the content of the record to be verified is consistent with the audio recording. Finally, based on the public security legal knowledge base, the preset open-source large language model is used to verify the legality and compliance of the record to be verified, ensuring that the target record text meets the requirements of laws and regulations, thereby improving the efficiency and accuracy of public security record production. Attached Figure Description
[0017] Figure 1 This is a flowchart of the public security record production method based on AI multi-model collaborative processing provided in the embodiments of this application; Figure 2 This is a schematic diagram of a visual interface provided in an embodiment of this application; Figure 3 This is another visual interface diagram provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the public security record-making device based on AI multi-model collaborative processing provided in the embodiments of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0019] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0020] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0022] Before providing a detailed description of the embodiments of this application, some of the nouns and terms used in the embodiments of this application will be explained first. The nouns and terms used in the embodiments of this application shall be interpreted as follows: Qwen3 is a new generation of open-source large language model series. Through hybrid inference, multi-language support, flexible scalability design and innovative architecture, it achieves a balance between high performance and low cost, is suitable for multi-scenario applications, and promotes the development of open-source large model technology.
[0023] Qwen2.5-VL is a multimodal visual language model designed specifically for visual tasks. Qwen2.5-VL can simultaneously process visual content such as images, text, and charts, supporting tasks such as object recognition, text parsing, and layout analysis.
[0024] ParaFormer is a novel shallow Transformer architecture that achieves high performance and low resource consumption through parallelization and progressive approximation. ParaFormer is suitable for resource-constrained devices and supports adaptive continuous learning and model scaling.
[0025] Whisper is an open-source Automatic Speech Recognition (ASR) model developed by OpenAI, featuring multilingual support, high robustness, and a wide range of applications. This model is suitable for scenarios such as voice assistants, meeting recording, caption generation, and multilingual communication.
[0026] In related technologies, traditional record-keeping mainly relies on police officers manually typing on a keyboard, recording simultaneously during questioning. For some officers unfamiliar with computer operations, this method is inefficient and carries the risk of errors. Taking telecommunications fraud cases as an example, due to the long crime chain and complex fund flows, the requirements for the accuracy of the record content are high. Police officers need to simultaneously record key information such as the process of being defrauded, the flow of funds, and the accounts involved during questioning. Under the current record-keeping method, errors in recording key information such as amounts and account numbers are prone to occur. These problems are often only discovered during the investigation stage, not only increasing the cost of evidence correction but also potentially leading to a break in the chain of evidence due to inaccurate key information, causing the case to be returned for supplementary investigation. Such fundamental errors not only prolong the case-handling period but may also cause irreversible losses due to missing the critical window for fund recovery.
[0027] With the popularization of artificial intelligence, especially large language model technology, some public security units have begun to try to apply it to the work of taking statements. Currently, there are three main application directions: First, based on large language models and professional knowledge bases, the legality and compliance of completed statements are checked, and the content that needs to be supplemented or modified is suggested. This method is a post-event remedy, which not only cannot avoid rework, but may also increase the subsequent workload of police officers. Second, the speech AI model is used to transcribe the content of the inquiry in real time to assist the police officers in making statements, and the large language model is used to suggest missing questions. However, in areas where dialects are commonly used, the accuracy of speech recognition is low and the applicability is poor. Third, information of the parties involved in the case is collected through pre-filled forms, and then a preliminary statement is generated by combining large language models and knowledge bases. The police officers supplement the questions and improve the content accordingly. In the process of making statements, if the information initially filled in by the parties is insufficient, it will significantly increase the workload of the police officers in subsequent supplementary inquiries.
[0028] In view of this, this application provides a method and related equipment for producing public security records based on AI multi-model collaborative processing, which can effectively improve the efficiency and accuracy of producing public security records and reduce the workload of relevant staff.
[0029] The method for creating police records based on AI multi-model collaborative processing provided in this application relates to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the method for creating police records based on AI multi-model collaborative processing, but is not limited to the above forms.
[0030] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0031] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0032] The embodiments of this application will be described in detail below with reference to the accompanying drawings: Figure 1 This is an optional flowchart of the public security record-keeping method based on AI multi-model collaborative processing provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S110 to S160: Step S110: Obtain the case type, case category, case feature information, case evidence images, and case audio recordings input into the preset visual interface; Step S120: Convert the case evidence images into text using a preset multimodal visual language model to obtain the first transcribed text; Step S130: Based on the case type, case category, case characteristic information and the first transcribed text, generate a customized transcript template using a preset open-source large language model; Step S140: Convert the case recording material into text using a preset speech AI model to obtain the second transcribed text; Step S150: Generate a transcript to be verified based on the customized transcript template and the second transcribed text using a preset open-source large language model; Step S160: Based on the public security legal knowledge base, the legality and compliance of the transcript to be verified is checked by a preset open-source large language model to obtain the target transcript text.
[0033] It is understood that the preset visual interface in this embodiment can be a pre-set system interface used to assist in the production of police records. For example, such as... Figure 2 As shown, this visual interface is used for recording statements in telecommunications fraud cases. When using the interface, relevant personnel can select the type of the current case (e.g., criminal, administrative, civil, or other) in the case type box; select the category of the current case (e.g., theft, telecommunications fraud, intentional injury) in the case category selection box; fill in case information (e.g., "near a shopping mall, high foot traffic") in the case characteristic information input box; and upload relevant materials (e.g., images of case evidence) in the relevant evidence upload box. Figure 3 As shown, you can also upload case audio files in the case audio recording upload box and select the number of languages included in the audio file according to the actual situation.
[0034] Specifically, in this embodiment, after all the information is filled in and uploaded in the visual interface, relevant personnel can click on the corresponding controls in the visual interface, such as clicking the "Start Processing" control. The data processing terminal can then automatically create and verify the transcript based on the information filled in or uploaded in the visual interface, effectively improving the efficiency and accuracy of transcript creation.
[0035] Understandably, when receiving a dispatch request, the responding police officer goes to the scene and uses a body camera to record video, and uses a police mobile phone to take photos to collect key elements such as the environment, items, people, and address involved in the case, and also conducts audio collection. Specifically, the responding police officer needs to conduct detailed questioning of the involved persons according to a standardized questioning process to fully understand the course of the case; if the on-site recording is noisy, contains dialects, or has a low proportion of key information, the police officer can record the key information in Mandarin to facilitate accurate information extraction by the system later. After completing the collection and questioning of relevant case data at the scene, the police officer uploads the relevant case data. In this embodiment, the police officer first fills in the case type, case category, and case characteristic information in the system's visual interface, and then uploads the case evidence pictures collected at the scene, as well as the recorded audio and video data. Among them, the case audio materials include the recording of the staff repeating key information; the case characteristic information is used to supplement the case details, such as writing "theft occurred between 5 pm and 6 pm, near a shopping mall, with high traffic" in a theft case, to enrich the information as much as possible, thereby generating targeted questioning and improving the quality of the record.
[0036] It is understood that, in this embodiment, after obtaining the case evidence image, the content of the case evidence image is converted into text using a preset multimodal visual language model to form the first transcribed text. The preset multimodal visual language model can be the Qwen2.5-VL model. Taking the Qwen2.5-VL model for converting case evidence images into text information as an example, this embodiment can analyze the data to be analyzed in the case evidence image. If the data to be analyzed does not include preset sensitive data, the case evidence image is converted into text using the preset multimodal visual language model to obtain the first transcribed text; if the data to be analyzed includes preset sensitive data, the case evidence image is converted into text using the preset multimodal visual language model to obtain the first text information to be processed, and then the first text information to be processed is corrected using a preset open-source large language model to obtain the first transcribed text. Specifically, the preset open-source large language model can be the Qwen3 large language model. For example, the preset sensitive data may include, but is not limited to, data such as bank account numbers and transaction records. If the data to be analyzed includes such sensitive data, the preset multimodal visual language model will encrypt this sensitive data during text conversion. For example, it will encrypt the sensitive data using "**" to obtain the first text information to be processed. Since the first text information to be processed contains encrypted data, it is incomplete. Therefore, in this embodiment, the Qwen3 large language model is used to perform context analysis and matching on the first text information to correct the missing data, thereby obtaining complete financial flow data to form the first transcribed text.
[0037] It is understood that, in this embodiment, after obtaining the first transcribed text, a first record template is retrieved from the public security legal knowledge base based on the case type and case category. Then, based on the case feature information and the first transcribed text, a target question is retrieved from the public security legal knowledge base. Finally, the information in the first record template is updated based on the target question to obtain a customized record template. For example, this embodiment first retrieves a coarser-grained first record template from the knowledge base based on the case type and case category. Then, it combines the first transcribed text and case feature information to retrieve and supplement customized target questions from the knowledge base, forming a fine-grained, customized record template as the customized record template for the current case.
[0038] Since this embodiment also acquires case audio recordings, it further performs text conversion on the recordings to obtain a second transcribed text. Specifically, when a video file is input into the visual interface, this embodiment automatically extracts the audio content from the audio file. In this embodiment, the extracted case audio recordings can be converted to text using a preset speech AI model. During the conversion process, this embodiment analyzes the language information in the case audio recordings. If the language information is a single language, the Paraformer model is used to convert the recordings to text to obtain the second transcribed text. If the language information is multilingual, the Whisper and Paraformer models are used to convert the recordings to text to obtain the second text to be processed. Then, a preset open-source large language model is used to fuse and logically correct the second text to obtain the second transcribed text. Specifically, in this embodiment, when the language information is determined to be multilingual, the Whisper and Paraformer models are used to convert the case audio materials into text, respectively, to obtain the corresponding second text information to be processed. Then, the second text information from both models is input into a preset open-source large language model for information fusion. Logical corrections are made to erroneous transcriptions, and Chinese numerals are converted to Arabic numerals, resulting in an accurate and standardized Mandarin transliteration as the second transliterated text. For example, for the parties' mobile phone numbers and ID numbers, which are Chinese numerals, the Qwen3 model will convert them to Arabic numerals. Furthermore, if illogical transcription errors occur, the Qwen3 model will correct these errors and standardize terminology, thereby improving the readability and standardization of the text.
[0039] It is understood that, in this embodiment, after obtaining the customized transcript template and the second transcribed text, a transcript to be verified is generated through a preset open-source large language model. Specifically, this embodiment can match each question to be matched in the customized transcript template with the content of the second transcribed text, and then extract the target content from the second transcribed text as the answer content of the question to be matched through the preset open-source large language model; then, all questions to be matched and all answer content are recombined in a preset order to obtain the transcript to be verified. For example, each "Question:" in the customized transcript template is matched with the content of the second transcribed text, and Qwen3 extracts the speech-to-text content to generate the corresponding "Answer:"; if a question has no corresponding answer in the transcribed content, the "Answer:" is left blank. Finally, each question-answer pair generated is recombined in the original order to ensure that all questions and answers in the transcript conform to the inquiry logic, thereby forming a completed transcript as the transcript to be verified.
[0040] It is understood that, after obtaining the transcript to be verified, this embodiment combines the transcript with the public security legal knowledge base and uses the Qwen3 model to check the transcript for legality and compliance, and makes corrections and reminders. Specifically, the transcript to be verified is used as input to the Qwen3 model, and a legality and compliance review is conducted in conjunction with the public security legal knowledge base; content that does not conform to the standards is automatically corrected or prompted, and the final legal, compliant, and directly usable transcript text is output as the target transcript text.
[0041] Understandably, after obtaining the automatically generated target transcript text, this embodiment will also manually check the target transcript text to verify the accuracy of key information such as the name and address of the party concerned, and upload the check results through a visual interface. If the check results determine that the content of the target transcript text contains content that does not meet the requirements, such as some questions not being answered or the content being incomplete, relevant personnel will be reminded to conduct a second inquiry with the party concerned. The target transcript text can then be adjusted based on the content of the second inquiry, or irrelevant questions in the target transcript text can be directly deleted to ensure the integrity, legality, and usability of the target transcript text.
[0042] In some embodiments, the complete implementation process of the method in this embodiment includes, but is not limited to, the following steps: Step 1: At the scene of the incident, police officers use their mobile phones to take photos of key evidence such as screenshots of the account numbers of the funds involved, images of the suspect's account information, and screenshots of bank transaction records. They also record conversations between the police officer and the victim, or recordings of the police officer recounting the key elements of the incident in Mandarin afterward.
[0043] Step 2: The police officer fills in the case type (such as "Criminal"), case category (such as "Telecom Fraud"), and case feature information (such as "Brush Order Part-time Telecom Fraud") in the visualization interface, and uploads the corresponding audio file of the case evidence pictures and case recording materials that have been collected.
[0044] Step 3: Invoke the Qwen2.5-VL multimodal model to parse the uploaded case evidence pictures, and extract the involved parties' account numbers, suspect account numbers, transaction amounts, time, etc. in the fund flow data. Since the middle part of the account number in some bank transaction records is omitted by "*", Qwen3 will also be used to automatically match it with the involved parties' account numbers and suspect account numbers, accurately identify and output the account number information in the transaction record, and then obtain the first transcript text.
[0045] Step 4: Match the first transcript text extracted by Qwen2.5-VL, the filled case type and case feature information with the public security record knowledge base (such as "Brush Order Part-time Telecom Fraud Record Template"), and input them into the Qwen3 large language model, so that the Qwen3 large language model can automatically generate a customized record template for telecom fraud cases based on the above information. Among them, the customized record template not only includes standard inquiry items, but also automatically supplements customized questions such as "Verification of Multi-account Fund Flow" and "Analysis of the Relevance of Involved Accounts" according to the complexity of the fund flow data, ensuring that the record content covers the key details of the case.
[0046] Step 5: Use speech AI models such as Whisper model or Paraformer to transcribe the uploaded audio into text to generate a preliminary transcript text. Subsequently, input the initial transcript text into the Qwen3 model for correction, including correcting transcription errors, converting Chinese numerals to Arabic numerals (such as correcting "one thousand yuan" to "1000 yuan"), and standardizing term expressions, etc., and output an accurate and standardized speech transcription content as the second transcript text.
[0047] Step 6: Use the customized record template, the corrected second transcript text, and the case information filled in by the police officer as the input of the Qwen3 model. The Qwen3 model will perform intelligent filling, record all the fund flow data accurately and without error in the record, and ensure that the questions and answers before and after in the record conform to the inquiry (interrogation) logic, and finally form a preliminary record with complete content and clear logic as the record to be verified.
[0048] Step 7: Compare the generated preliminary record with the public security legal knowledge base, use the Qwen3 model to conduct legal compliance review and automatically make corrections and prompts, and finally output a legal, compliant, and directly usable record text for handling cases as the target record text.
[0049] As can be seen from the above, in the information entry stage, this application embodiment effectively avoids the problem of inaccurate dialect recognition by introducing a police officer's Mandarin repetition recording mechanism, thereby improving the applicability and reliability of the application system corresponding to the method of this application embodiment in diverse practical scenarios. In the content generation stage, by integrating the Qwen3 large language model, the Qwen2.5-VL multimodal model, and various speech AI models, it can deeply understand and integrate multi-source information such as speech, images, and text. Combined with the public security professional knowledge base, it automatically generates complete and logically clear transcripts, improving the accuracy and structure of the transcripts from the source, and effectively avoiding the problem of recording errors of key information (such as the amount involved in the case, accounts, etc.) that is prone to occur in traditional manual entry. Furthermore, in terms of compliance assurance, this embodiment automatically conducts legality and standardization reviews after the transcript is generated through an embedded legal compliance knowledge base. This enables intelligent prompts and corrections for potential omissions and non-standard expressions. Compared to police officers making transcripts and then using AI to check and correct them, this is more efficient and fundamentally reduces the risk of a broken chain of evidence or a case being returned for supplementary investigation due to transcript quality issues.
[0050] In summary, the embodiments of this application have achieved a revolutionary change in the case-handling model from "police officers inputting text word by word" to "system generation - police officers checking and revising". Through multi-model collaboration and process optimization, it has achieved comprehensive breakthroughs in improving the quality of written records, ensuring procedural compliance, and reducing the burden on police officers, providing efficient and reliable technical support for police operations.
[0051] Please see Figure 4 This application also provides a public security record-making device based on AI multi-model collaborative processing, the device comprising: The acquisition module is used to acquire case type, case category, case feature information, case evidence images, and case audio materials input from the preset visual interface; The first conversion module is used to convert case evidence images into text using a preset multimodal visual language model to obtain the first transcribed text. The first generation module is used to generate customized transcript templates based on case type, case category, case feature information and the first transcribed text, using a preset open-source large language model. The second conversion module is used to convert the case audio materials into text using a preset speech AI model to obtain the second transcribed text. The second generation module is used to generate a transcript to be verified based on a customized transcript template and a second transcribed text, using a preset open-source large language model. The verification module is used to perform legality and compliance verification on the transcript to be verified based on the public security legal knowledge base and a preset open-source large language model, so as to obtain the target transcript text.
[0052] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0053] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0054] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0055] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0056] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0057] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0058] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0059] This application provides a method and related equipment for creating police records based on AI multi-model collaborative processing. After inputting case type, case category, case characteristic information, case evidence images, and case audio materials through a preset visual interface, the method uses a preset multimodal visual language model to convert the case evidence images into text, obtaining a first transcribed text. Then, based on the case type, case category, case characteristic information, and the first transcribed text, a customized record template is generated using a preset open-source large language model, ensuring that the customized record template conforms to the actual situation of the current case type and category. Next, a preset speech AI model converts the case audio materials into text, obtaining a second transcribed text. Based on the customized record template and the second transcribed text, a record to be verified is generated using a preset open-source large language model, ensuring that the content of the record to be verified is consistent with the audio content. Finally, based on a public security legal knowledge base, the record to be verified is validated for legality and compliance using a preset open-source large language model, ensuring that the target record text meets the requirements of laws and regulations, thereby improving the efficiency and accuracy of police record creation.
[0060] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0061] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0062] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0063] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0064] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0065] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0066] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0067] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0068] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0069] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0070] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for producing police transcripts based on AI multi-model collaborative processing, characterized in that, The method includes the following steps: The system retrieves the case type, case category, case characteristic information, case evidence images, and case audio recordings input into a preset visual interface. The case evidence images are converted into text using a preset multimodal visual language model to obtain the first transcribed text; Based on the case type, case category, case feature information, and the first transcribed text, a customized transcript template is generated using a preset open-source large language model; The case recording material is converted into text using a preset speech AI model to obtain a second transcribed text; Based on the customized transcript template and the second transcribed text, a transcript to be verified is generated using the preset open-source large language model; Based on the public security legal knowledge base, the legality and compliance of the transcript to be verified are checked using the preset open-source large language model to obtain the target transcript text.
2. The method according to claim 1, characterized in that, The audio recordings of the case include recordings of staff repeating key information.
3. The method according to claim 1, characterized in that, The step of converting the case evidence images into text using a preset multimodal visual language model to obtain the first transcribed text includes: Analyze the data to be analyzed in the evidence images of the case; If the data to be analyzed does not include preset sensitive data, the case evidence images are converted into text using a preset multimodal visual language model to obtain the first transcribed text; If the data to be analyzed includes preset sensitive data, the case evidence image is converted into text using a preset multimodal visual language model to obtain the first text information to be processed; and the first text information to be processed is corrected using a preset open-source large language model to obtain the first transcribed text.
4. The method according to claim 1, characterized in that, The step of converting the case audio material into text using a preset speech AI model to obtain the second transcribed text includes: Analyze the language information in the audio recordings of the case; If the language information is a single language, the case audio material is converted into text using the Paraformer model to obtain the second transcribed text. If the language information is multilingual, the case audio material is converted into text using the Whisper model and the Paraformer model respectively to obtain the second text information to be processed; and the second text information to be processed is fused and logically corrected using a preset open-source large language model to obtain the second transcribed text.
5. The method according to claim 1, characterized in that, The step of generating a customized transcript template based on the case type, case category, case feature information, and the first transcribed text using a preset open-source large language model includes: The first transcript template is obtained from the public security legal knowledge base based on the case type and the case category. Based on the case feature information and the first transcribed text, the target question is retrieved from the public security legal knowledge base; The information in the first transcript template is updated based on the target problem to obtain the customized transcript template.
6. The method according to claim 1, characterized in that, The step of generating a transcript to be verified based on the customized transcript template and the second transcribed text using the preset open-source large language model includes: Each question to be matched in the customized transcript template is matched with the content of the second transcribed text, and the target content is extracted from the second transcribed text as the answer content of the question to be matched using the preset open-source large language model; All the questions to be matched and all the answers are recombined in a preset order to obtain the transcript to be verified.
7. The method according to claim 1, characterized in that, The method further includes the following steps: Obtain the inspection results of the target transcript text; If the inspection results determine that the target transcript contains content that does not meet the requirements, the target transcript will be adjusted.
8. A police record-taking device based on AI multi-model collaborative processing, characterized in that, The device includes: The acquisition module is used to acquire case type, case category, case feature information, case evidence images, and case audio materials input from the preset visual interface; The first conversion module is used to convert the case evidence images into text using a preset multimodal visual language model to obtain the first transcribed text. The first generation module is used to generate a customized transcript template based on the case type, the case category, the case feature information and the first transcribed text, using a preset open-source large language model. The second conversion module is used to convert the case audio material into text using a preset speech AI model to obtain a second transcribed text. The second generation module is used to generate a transcript to be verified based on the customized transcript template and the second transcribed text, using the preset open-source large language model. The verification module is used to perform legality and compliance verification on the transcript to be verified based on the public security legal knowledge base and the preset open-source large language model to obtain the target transcript text.
9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.