Intelligent record generation method, system and terminal
Through the intelligent transcript generation method, voice recognition and case analysis models are used to record speeches in real time and generate legal suggestions, solving the problem that the existing transcript system relies on manual operations, and improving the case handling efficiency and document quality of the political and legal departments.
Patent Information
- Application Number
- CN202510094071.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
The existing transcript systems rely mostly on manual operations and have low intelligence, resulting in low case handling efficiency and poor experience in political and legal departments, making it difficult to meet the needs of modern political and legal departments to handle cases quickly and efficiently.
Provide an intelligent transcript generation method, by collecting original Q&A voice data in real time, using the pre-trained speech recognition model to identify spokespersons and generate speech text data; based on the pre-trained case analysis model, identify basic case information, determine applicable legal provisions and similar cases, generate legal suggestions, and automatically generate a specified type of transcript document.
It realizes real-time recording of speeches by all parties, supports multi-role identification, automatically labels spokespersons, provides effective legal references, ensures the consistency of legal application, shortens case handling time, improves case handling efficiency, and ensures the professionalism and format of transcripts.
Smart Images

Figure CN119988710A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an intelligent transcript generation method, system and terminal. Background Art
[0002] Traditional transcript systems used in political and legal departments mainly rely on manual operations, requiring manual recording, transcription, and document generation. This approach will result in a large workload, error-proneness, and low efficiency in completing the overall transcript process, making it difficult to meet the needs of modern political and legal departments to handle cases quickly and efficiently. In addition, as the number of cases increases year by year, political and legal departments, including public security departments and judicial departments, are in urgent need of intelligent transcript systems so that police officers and judicial personnel can quickly generate and archive interrogation transcripts, law enforcement transcripts, and trial records, thereby improving case handling efficiency and ensuring the accuracy of the transcript content.
[0003] At present, the intelligent generation technology of transcript documents is mostly limited to speech recognition technology, which mainly includes collecting speech data and transcribing speech data into text data. However, it lacks the ability to deeply understand the case, intelligent analysis ability, and automatic document generation ability. Therefore, in actual applications, police officers or judicial personnel still need to manually organize transcripts, review transcripts, and search for relevant legal provisions and similar precedents, and the efficiency of case handling is still low. Summary of the invention
[0004] In view of the shortcomings of the prior art mentioned above, the purpose of this application is to provide an intelligent transcript generation method, system and terminal to solve the technical problems of low case handling efficiency and poor experience of the political and legal departments caused by the existing transcript system's reliance on manual operation and low level of intelligence.
[0005] To achieve the above-mentioned purpose and other related purposes, the first aspect of the present application provides an intelligent transcript generation method, which includes: real-time collection of multiple original question and answer voice data, and based on a pre-trained speech recognition model, identifying the speakers of each original question and answer voice data, and generating speech text data of each speaker; based on a pre-trained case analysis model, identifying the basic case information of the current case according to the speech text data of each speaker, and determining the applicable legal provisions and similar precedents, and generating legal advice for the current case, so that political and legal personnel can continue to enforce the law based on the legal advice; generating a transcript document of a specified type based on the generated speech text data of each speaker, the basic case information of the current case, and the legal advice.
[0006] In some embodiments of the first aspect of the present application, based on a pre-trained speech recognition model, the speaker of each original question and answer speech data is identified, and the speech text data of each speaker is generated. The method includes: transcribing each original question and answer speech data into text to generate corresponding initial speech text data, and performing grammatical error correction and semantic error correction on each initial speech text data to generate corresponding speech text data; extracting voiceprint features in each original question and answer speech data to generate corresponding voiceprint feature data; matching each generated voiceprint feature data with each registered voiceprint feature data in a voiceprint feature library to determine the speaker of each original question and answer speech data and the role identity of the speaker; matching each speaker with each speaker text data, and marking the speaker corresponding to each speaker text data.
[0007] In some embodiments of the first aspect of the present application, the method of performing grammatical correction and semantic correction on the initial speech text data includes: performing semantic structure analysis on the initial speech text data, and automatically segmenting the initial speech text data according to pauses in the original question and answer voice data to obtain one or more text sentences; performing grammar detection on each text sentence, and performing grammatical correction on the detected grammatical errors; obtaining context information of the initial speech text data, and performing context analysis and semantic understanding on each text sentence after grammatical correction based on the context information, correcting homophones, ambiguous words and spelling errors therein, screening out repeated content therein, supplementing missed information, and generating speech text data.
[0008] In some embodiments of the first aspect of the present application, the generated voiceprint feature data is matched with each registered voiceprint feature data in the voiceprint feature library to determine the speaker of the original question and answer voice data and the role identity of the speaker, including: respectively calculating the similarity between the voiceprint feature data and each registered voiceprint feature data in the voiceprint feature library, and obtaining the registered voiceprint feature data with the largest similarity value as the target registered voiceprint feature data; if the similarity value is greater than or equal to a preset similarity threshold, obtaining the speaker of the target registered voiceprint feature data and the role identity of the speaker, and using the speaker as the speaker of the voiceprint feature data; if the similarity value is less than the preset similarity threshold, determining that the speaker of the voiceprint feature data is a new speaker, and after the staff of the political and legal system confirms the role identity of the new speaker, storing the voiceprint feature data and the confirmed new speaker in the voiceprint feature library.
[0009] In some embodiments of the first aspect of the present application, based on a pre-trained case analysis model, according to the speech text data of each speaker, the basic case information of the current case is identified, and the applicable legal provisions and similar precedents are determined, and the method of generating legal advice for the current case includes: extracting keywords from each speech text data respectively to identify the basic case information of the current case; wherein the basic case information includes: the time of the incident, the location of the incident, the persons involved, one or more case facts, and one or more case evidence; performing logical analysis on the obtained basic case information, confirming the time relationship and causal relationship between the facts of each case, and obtaining the case process of the current case; performing semantic analysis on the obtained basic case information, identifying the disputed facts and undisputed facts in each case fact, and confirming the relationship between the evidence of each case; performing legal reasoning on the undisputed facts and multiple dispute directions of the disputed facts respectively, generating multiple corresponding reasoning paths, and retrieving the pre-constructed legal knowledge base to obtain the legal provisions applicable to each reasoning path, and retrieving the pre-constructed case vector database and the case knowledge base to obtain similar precedents of each reasoning path; according to the preset transcript type, generating legal advice on the current case in a corresponding format.
[0010] In some embodiments of the first aspect of the present application, the method of retrieving a pre-constructed case vector database and a case knowledge base to obtain similar precedents for each reasoning path includes: generating corresponding case summary information based on the undisputed facts of the current case and multiple dispute directions of the disputed facts; obtaining the case summary information of each historical case in the case vector database, and calculating the similarity between the current case and each historical case respectively; screening one or more historical cases whose similarity values are greater than or equal to a preset similarity threshold as similar precedents for the current case; searching the case knowledge base to obtain the basic case information, case judgment results and applicable legal provisions of each similar precedent.
[0011] In some embodiments of the first aspect of the present application, the intelligent transcript generation method also includes: searching for relevant legal provisions and similar cases based on one or more keywords input by the user; wherein the specific method includes: traversing the input information of the same user, obtaining the context information of the current input, and performing context analysis and semantic understanding on each keyword currently input based on the context information, correcting homophones, ambiguous words and spelling errors therein, and obtaining rewritten retrieval information; searching a pre-constructed legal knowledge base based on the retrieval information to obtain relevant legal provisions; searching a pre-constructed case vector database based on the retrieval information to obtain one or more similar cases, and case summary information of each similar case; obtaining related cases of each similar case based on a pre-constructed case knowledge graph; wherein the graph nodes of the case knowledge graph are historical cases, and the edges are similarity relationships between historical cases; searching a pre-constructed case knowledge base to obtain basic case information, case judgment results and applicable legal provisions of each similar case and each related case; and outputting each similar case and each related case in the form of a semantic network.
[0012] In some embodiments of the first aspect of the present application, the intelligent transcript generation method also includes: storing the speech text data of each speaker, the basic case information of the current case, legal advice and the generated transcript document in a pre-constructed text database; after the current case is closed, storing the basic case information of the current case, the case judgment result and the applicable legal provisions in a pre-constructed case knowledge base; generating corresponding case summary information based on the basic case information, the case judgment result and the applicable legal provisions of the current case, and performing vector conversion on the case summary information based on the pre-constructed embedding model to generate corresponding semantic vector data to store in the pre-constructed case vector database; updating the pre-constructed case knowledge graph based on the similarity relationship between the current case and each historical case in the case vector database.
[0013] To achieve the above-mentioned purpose and other related purposes, the second aspect of the present application provides an intelligent transcript generation system, which includes: a client and a server; wherein the server is communicatively connected to the client; the server includes: a speech recognition module, which is used to identify the speakers of multiple original question and answer voice data collected in real time by the client based on a pre-trained speech recognition model, and generate speech text data of each speaker to send to the client for display; a case analysis module, connected to the speech recognition module, which is used to identify the basic case information of the current case based on the pre-trained case analysis model and the speech text data of each speaker, and determine the applicable legal provisions and similar precedents, generate legal advice for the current case, so that the political and legal staff can continue to enforce the law according to the legal advice, and display it through the client; a document generation module, connected to the case analysis module, which is used to generate a specified type of transcript document based on the generated speech text data of each speaker, the basic case information of the current case and the legal advice, and send it to the client for display.
[0014] To achieve the above-mentioned purpose and other related purposes, the third aspect of the present application provides an intelligent note generation terminal, which includes: a processor and a memory; the memory is used to store computer programs; the processor is used to execute the computer programs stored in the memory, so that the terminal executes any one of the intelligent note generation methods provided in the above embodiments.
[0015] As described above, the present application provides an intelligent transcript generation method, system and terminal, which have the following beneficial effects: by identifying the speakers of multiple collected original question-and-answer voice data based on a pre-trained speech recognition model, and generating the speech text data of each speaker, it is possible to record the speeches of all parties in real time, and support multi-role recognition, and automatically annotate the speech text data of different speakers; by identifying the basic case information of the current case based on a pre-trained case analysis model, and determining the applicable legal provisions and similar precedents, and generating corresponding legal suggestions, thereby providing effective legal references for political and legal personnel and ensuring the consistency of legal application; by automatically generating transcript documents of specified types, the uniformity and professionalism of the transcript document format are ensured, and the case handling time is effectively shortened and the case handling efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Shown is a flow chart of an intelligent transcript generation method in one embodiment of the present application.
[0017] Figure 2 Shown is a schematic diagram of a process for generating speech text data of each speaker in an embodiment of the present application.
[0018] Figure 3 Shown is a flowchart of generating legal advice for the current case in one embodiment of the present application.
[0019] Figure 4 Shown is a structural diagram of an intelligent transcript generation system in one embodiment of the present application.
[0020] Figure 5 Shown is a schematic diagram of the structure of a client in one embodiment of the present application.
[0021] Figure 6 Shown is a flowchart of case intelligent analysis and suggestion in one embodiment of the present application.
[0022] Figure 7 Shown is a flowchart of user intelligent retrieval in one embodiment of the present application.
[0023] Figure 8 Shown is a schematic diagram of the structure of an intelligent note generation terminal in one embodiment of the present application. DETAILED DESCRIPTION
[0024] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0025] In legal activities, public security, procuratorate, judicial and other political and legal personnel often need to question, interrogate, investigate, collect evidence or hold trials against criminal suspects, offenders or other persons involved in the case or informed persons during investigation and trial activities. In order to record information such as the case, evidence, behavior, speech and statements, different types of transcripts will be involved, such as: interrogation transcripts generated when questioning witnesses, victims, etc. during the investigation process, interrogation transcripts generated when interrogating criminal suspects; investigation reports generated by recording the investigation process during the investigation and evidence collection process; transcript documents generated by punishing offenders and recording case information during public security management; accident reports generated by recording traffic accident scene information during traffic accident handling; and trial transcripts generated by recording the entire trial process during the court trial.
[0026] Traditional transcript systems mainly rely on manual operation, and recording, transcription and document generation need to be completed manually. This method will result in a large workload, prone to errors and low efficiency in completing the overall transcript process, which is difficult to meet the needs of modern political and legal departments to handle cases quickly and efficiently. The intelligent generation technology of transcript documents is mostly limited to speech recognition technology, lacking the ability to understand the case in depth, intelligent analysis capabilities and automatic document generation capabilities. In order to solve the above technical problems, this application provides an intelligent transcript generation method, system and terminal. Through the AI large language model, deep learning technology, speech recognition technology and natural language processing technology, it can record the speeches of all parties in real time when political and legal staff conduct investigations, inquiries or trials, and analyze the transcribed text content, automatically extract the basic information of the current case, and prompt political and legal staff to the applicable legal provisions and similar precedents for the current case, thereby solving the technical problems of low efficiency and poor experience in case handling by political and legal departments caused by the existing transcript system that relies heavily on manual operation and low intelligence.
[0027] At the same time, in order to make the purpose, technical solutions and advantages of this application more clear, the technical solutions in the embodiments of this application are further described in detail through the following embodiments and in combination with the accompanying drawings. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit the invention.
[0028] like Figure 1 As shown, a flow chart of the intelligent transcript generation method in the embodiment of the present application is shown. The intelligent transcript generation method in the present embodiment mainly includes the following steps.
[0029] Step S1: collect multiple original question-and-answer voice data in real time, and identify the speakers of each original question-and-answer voice data based on a pre-trained speech recognition model to generate speech text data of each speaker.
[0030] Specifically, each original question and answer voice data can be collected in real time through a microphone or other voice collection equipment to obtain the original speeches of the parties during the legal activities as the original material for the final transcript document.
[0031] It should be noted that the speech recognition model can be trained based on a deep neural network, a convolutional neural network, a recurrent neural network or a long short-term memory network, and adopts an unsupervised learning or semi-supervised learning method to realize the automatic conversion of speech signals into corresponding texts. However, it should be noted that the network structure of the speech recognition model is not limited in this application.
[0032] In one embodiment, if Figure 2 As shown, step S1 includes the following steps.
[0033] Step S11: transcribe each original question-and-answer voice data into text to generate corresponding initial speech text data, and perform grammatical and semantic correction on each initial speech text data to generate corresponding speech text data.
[0034] In a specific embodiment, after the original question-answer voice data is transcribed using voice recognition technology, the voice recognition model can continue to optimize the generated initial speech text data based on natural language processing technology. Specifically, the method of performing grammatical and semantic correction on the initial speech text data includes the following steps.
[0035] ① Performing semantic structure analysis on the initial speech text data, and automatically segmenting the initial speech text data according to pauses in the original question-and-answer voice data to obtain one or more text sentences.
[0036] ② Perform grammar check on each text sentence and make grammatical corrections for detected grammatical errors.
[0037] ③ Obtain the context information of the initial speech text data, and based on the context information, perform context analysis and semantic understanding on each text sentence after grammatical correction, correct homophones, ambiguous words and spelling errors, filter out repeated content, supplement missed information, and generate speech text data.
[0038] In this embodiment, the application can automatically correct grammatical errors, spelling errors, homophones, ambiguous words in each initial speech text data, delete semantically repeated content, supplement details, and mark logical relationships without changing the original meaning of the sentence. It can improve the accuracy, readability, fluency, logic and contextual consistency of the generated speech text data, ensure the high quality of the generated speech text data, and thus greatly improve the accuracy of speech recognition by the speech recognition model.
[0039] Step S12: extracting voiceprint features from each original question-and-answer speech data, and generating corresponding voiceprint feature data.
[0040] Step S13: Match the generated voiceprint feature data with the registered voiceprint feature data in the voiceprint feature library to determine the speaker of each original question and answer voice data and the role identity of the speaker.
[0041] In a specific embodiment, a voiceprint feature library is pre-constructed, and voice data of multiple political and legal personnel can be collected in advance, and voiceprint features in each voice data can be extracted to generate corresponding voiceprint feature data, which are registered and stored in the voiceprint feature library as registered voiceprint feature data, so that the corresponding spokesperson can be matched in time during subsequent transcription, and relevant legal suggestions in a specified format can be automatically pushed according to the role identity of the spokesperson.
[0042] Based on the voiceprint feature library, the method of sequentially matching each voiceprint feature data with each registered voiceprint feature data in step S13 to determine the speaker corresponding to the original question and answer voice data and the speaker's role identity mainly includes the following steps.
[0043] ① Calculate the similarity between the voiceprint feature data and each registered voiceprint feature data in the voiceprint feature library respectively, and obtain the registered voiceprint feature data with the largest similarity value as the target registered voiceprint feature data.
[0044] ② If the similarity value is greater than or equal to the preset similarity threshold, the speaker of the target registered voiceprint feature data and the role identity of the speaker are obtained, and the speaker is used as the speaker of the voiceprint feature data.
[0045] ③ If the similarity value is less than the preset similarity threshold, the speaker of the voiceprint feature data is determined to be a new speaker. After the political and legal staff confirms the role identity of the new speaker, the voiceprint feature data and the confirmed new speaker are stored in the voiceprint feature library.
[0046] For example, when political and legal personnel, such as police officers, are interrogating criminal suspects, they usually have registered voiceprint feature data in the voiceprint feature library in advance. Therefore, when the police officers speak, the speech recognition model can directly recognize it and mark its role identity; when the criminal suspect speaks, since it has not been registered in the voiceprint feature library in advance, the speech recognition model will automatically identify it as a new speaker, such as marked as speaker 1. After the police officer enters and confirms his name and role identity through the user interface, all the original question and answer voice data and speech text data of the criminal suspect can be automatically marked, and the role information, identity information and voiceprint feature data of the criminal suspect can be registered in the voiceprint feature library for automatic recognition when the criminal suspect continues to speak.
[0047] Step S14: Match each speaker with the speaker text data, and mark the speaker corresponding to each speaker text data.
[0048] Step S2: Based on the pre-trained case analysis model, according to the speech text data of each speaker, the basic case information of the current case is identified, and the applicable legal provisions and similar precedents are determined, and legal suggestions for the current case are generated so that the political and legal staff can continue to enforce the law according to the legal suggestions.
[0049] In one embodiment, the case analysis model is trained based on a large language model, aiming to understand the input speech text data of each speaker, so as to conduct in-depth analysis and reasoning on the current case, provide legal advice to political and legal personnel, assist political and legal personnel in conducting case interrogations and completing investigation reports, etc.
[0050] Preferably, the network structure of the large language model includes a plurality of stacked Transformer layers, each Transformer layer includes a combination of a multi-head attention layer and a fully connected feedforward neural network layer, so that the trained case analysis model can capture language information at different levels and process complex language tasks. It should be noted that this application does not limit the network structure of the case analysis model, and users can set it according to their needs.
[0051] In one embodiment, if Figure 3 As shown, step S2 includes the following steps.
[0052] Step S21: extract keywords from each speech text data respectively to identify the basic case information of the current case.
[0053] The basic case information includes: the time of the incident, the location of the incident, the persons involved, one or more case facts, and one or more case evidence. Specifically, the persons involved include but are not limited to: parties, witnesses, lawyers, etc.; the case evidence includes but is not limited to: physical evidence, documentary evidence and other relevant evidence materials involved in the current case.
[0054] In this embodiment, the speech text data of each speaker generated based on the speech recognition model can be combined with the preset case prompt words to be input into the case analysis model to understand the case and obtain the basic case information of the current case. The preset case prompt words can include the information that needs to be extracted, such as "time of the incident" and "location of the incident".
[0055] It should be noted that if the input speech text data have not yet marked the specific speakers and the speakers' role identities, the case prompt words can be used to prompt the case analysis model to distinguish different speakers and the identities of the parties according to the contextual information of each speech text data, thereby effectively identifying the key information of the current case, and ensuring the comprehensiveness and accuracy of the basic case information obtained, laying the foundation for subsequent case analysis.
[0056] Step S22: Perform a logical analysis on the obtained basic case information, confirm the time relationship and causal relationship between the facts of each case, and obtain the case process of the current case.
[0057] Step S23: Perform semantic analysis on the obtained basic case information, identify disputed facts and undisputed facts in each case, and confirm the relationship between the evidence in each case.
[0058] Specifically, the disputed facts refer to the different interpretations of the same case facts by the parties, and the undisputed facts refer to the case facts recognized by both or all parties.
[0059] In this embodiment, based on the case analysis model, the thinking chain technology is used to simulate the thinking path of political and legal personnel, so as to achieve accurate and complete analysis of the current case. It should be understood that the thinking chain technology can provide an overall thinking framework from problem definition to solution execution, and can use task decomposition to decompose complex legal issues into several detailed steps, prompting the case analysis model to gradually perform logical analysis, semantic analysis, etc. on the basic information of the case, effectively improving the accuracy and efficiency of the model in case analysis.
[0060] In a preferred embodiment, the identified disputed facts and undisputed facts of the current case can be marked so as to be highlighted in the subsequently generated transcript document, thereby improving the user experience of the intelligent transcript.
[0061] Step S24: Perform legal reasoning on the undisputed facts and multiple dispute directions of disputed facts respectively, generate multiple corresponding reasoning paths, and search the pre-built legal knowledge base to obtain the legal provisions applicable to each reasoning path, and search the pre-built case vector database and case knowledge base to obtain similar cases for each reasoning path.
[0062] For disputed facts, this application can determine the respective dispute directions based on the different interpretations of the parties, and generate corresponding reasoning paths respectively, so as to provide legal advice for each reasoning path and provide methods to resolve disputes, so as to help political and legal personnel identify disputes and take corresponding measures to uncover the truth and verify the facts of the case.
[0063] In a specific embodiment, a method of retrieving a pre-constructed case vector database and a case knowledge base to obtain similar precedents for each reasoning path includes: generating corresponding case summary information based on the undisputed facts of the current case and multiple dispute directions of the disputed facts; obtaining the case summary information of each historical case in the case vector database, and calculating the similarity between the current case and each historical case respectively; screening one or more historical cases whose similarity values are greater than or equal to a preset similarity threshold as similar precedents for the current case; and retrieving the case knowledge base to obtain the basic case information, case judgment results, and applicable legal provisions of each similar precedent.
[0064] Step S25: Generate legal advice on the current case in a corresponding format according to the preset transcript type.
[0065] In this embodiment, the understanding, analysis and reasoning results of the current case are aggregated and further processed to generate a structured and operational legal advice document, which is pushed to political and legal personnel in a specified format, such as generating trial suggestions and pushing them to judges or generating reference precedents and pushing them to case handlers, to provide direct support for case handling; and, through information aggregation, the logical consistency, completeness, accuracy, applicability and timeliness of the generated legal advice can be ensured, providing effective legal references for public security personnel and judicial personnel.
[0066] Step S3: Generate a specified type of transcript document based on the generated speech text data of each speaker, basic case information of the current case, and legal advice.
[0067] Specifically, according to the transcript type specified by the political and legal staff, the corresponding transcript template is obtained; and according to the speech text data of each speaker, the basic case information of the current case, and the key information in the legal advice supplementary transcript template, the final transcript document is generated; thereby realizing the standardized automatic generation of transcript documents, reducing the burden of manual sorting, improving case handling efficiency, and ensuring the professionalism of the transcript documents and the standardization of the format. The types of transcript documents include but are not limited to: interrogation transcripts, interrogation transcripts, investigation reports, trial transcripts, etc.
[0068] In one embodiment, the intelligent transcript generation method further includes:
[0069] According to one or more keywords input by the user, relevant legal provisions and similar cases are retrieved, which specifically includes the following steps.
[0070] ① Traverse the input information of the same user, obtain the context information of the current input, and perform context analysis and semantic understanding on each keyword currently input based on the context information, correct homophones, ambiguous words and spelling errors, and obtain rewritten retrieval information.
[0071] ② According to the search information, search the pre-built legal knowledge base to obtain relevant legal provisions.
[0072] ③ According to the search information, search the pre-constructed case vector database to obtain one or more similar cases and case summary information of each similar case.
[0073] Specifically, the similarities between the search information and the case summary information of each historical case in the case vector database are calculated respectively, and one or more historical cases with a similarity greater than or equal to a preset similarity threshold are output as similar cases.
[0074] ④ Based on the pre-built case knowledge graph, obtain related cases of similar cases.
[0075] Among them, the graph nodes of the case knowledge graph are historical cases, and the edges are similarity relationships between historical cases.
[0076] ⑤ Search the pre-built case knowledge base to obtain basic case information, case judgment results and applicable legal provisions of each similar case and each related case.
[0077] ⑥ Output similar cases and related cases in the form of semantic network.
[0078] In this embodiment, political and legal staff can quickly search for legal provisions and cases through the case analysis model, which can help political and legal staff make judicial decisions. In addition, this application can also provide Internet search, and call other search tools to search for corresponding content, so as to combine the search results and aggregate and output them to political and legal staff. It should be noted that this application can output each search result in the form of a semantic network so that political and legal staff can better understand the case background. Among them, each node of the semantic network is a legal provision or a similar case. After clicking, it can be expanded to display detailed legal content or basic case information, case judgment results, etc.
[0079] In one embodiment, the intelligent transcript generation method also includes the archiving of transcript documents and case archiving. Specifically, it includes: storing the speech text data of each speaker, the basic case information of the current case, legal advice and the generated transcript documents in a pre-constructed text database; after the current case is closed, storing the basic case information, case judgment results and applicable legal provisions of the current case in a pre-constructed case knowledge base; generating corresponding case summary information based on the basic case information, case judgment results and applicable legal provisions of the current case, and performing vector conversion on the case summary information based on the pre-constructed embedding model to generate corresponding semantic vector data to store in the pre-constructed case vector database; updating the pre-constructed case knowledge graph based on the similarity relationship between the current case and each historical case in the case vector database.
[0080] The text database is used to store the transcript documents and the speech text data generated during the transcript process, basic case information of the relevant cases, legal advice and other text contents.
[0081] The case knowledge base is used to store the basic case information, case judgment results and applicable legal provisions of closed historical cases. It should be noted that the application also pre-constructs a legal knowledge base to store all legal provisions, including the Constitution, Criminal Law, Civil Law and other legal content.
[0082] The case vector database is used to store case summary information of historical cases in the case knowledge base.
[0083] The case knowledge graph is used to store the association relationship between historical cases in the case knowledge base to support cross-case association. The graph nodes of the case knowledge graph are historical cases, and the edges are similarity relationships between historical cases.
[0084] In this embodiment, by pre-building the text database, the legal knowledge base, the case knowledge base, the case vector database and the case knowledge graph, and providing an incremental storage method, it not only supports efficient case filing management, but also supports rapid retrieval of relevant legal provisions and similar cases when needed, provides data support for judicial decision-making, further improves the standardization of case management, and improves retrieval efficiency.
[0085] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" represent examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0086] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.
[0087] like Figure 4 , which shows a schematic diagram of the structure of the intelligent transcript generation system 400 in the embodiment of the present application. The intelligent transcript generation system 400 includes: a client 401 and a server 402.
[0088] The client 401 is used to collect multiple original question-and-answer voice data in real time, and display the speech text data of each speaker, the legal advice of the current case, the transcript document and other contents generated by the server 402. Specifically, the client 401 includes but is not limited to: mobile phones, tablets and other terminal devices. Figure 5As shown, the client 401 includes, in addition to voice collection devices such as microphones, a central processing unit CPU, a read-only memory ROM, a random access memory RAM, an AI acceleration module, a high-speed bus, and an IO interface. The IO interface communicates with each database that implements the storage function and the server 402 to ensure the input and output of the client 401.
[0089] The server side 402 is communicatively connected to the client side 401 , and the server side 402 includes: a speech recognition module 4021 , a case analysis module 4022 , and a document generation module 4023 .
[0090] Specifically, the speech recognition module 4021 is used to identify the speakers of multiple original question and answer voice data collected in real time by the client 401 based on a pre-trained speech recognition model, and generate speech text data of each speaker to send to the client 401 for display.
[0091] In a preferred embodiment, when the client 401 displays each speech text data, it can display the corresponding marked speaker and the speaker role identity, and can view the transcription progress of each original question and answer voice data into text in real time through the user interface, ensuring that each speech of each speaker is accurately recorded.
[0092] The case analysis module 4022 is connected to the speech recognition module 4021, and is used to identify the basic case information of the current case based on the pre-trained case analysis model and the speech text data of each speaker, and determine the applicable legal provisions and similar precedents, and generate legal suggestions for the current case so that political and legal personnel can continue to enforce the law based on the legal suggestions, and display them through the client 401.
[0093] In a specific embodiment, the case analysis model is trained based on a large language model. The case analysis module 4022 includes a case understanding unit, a case analysis unit, a case reasoning unit, and a push suggestion unit, which are used to perform intelligent analysis on the current case and provide legal suggestions. Figure 6As shown, through the case understanding unit, the speech text data written by each original question and answer voice data and the preset case prompt words are input into the large language model to understand the current case and identify its basic case information; through the case analysis unit, the large language model uses thinking chain timing and task decomposition technology to analyze the current case and confirm the case process, disputed facts and undisputed facts of the current case; through the case reasoning unit, based on the pre-built legal knowledge base, case knowledge base, case vector database and retriever, the legal provisions and similar precedents applicable to the current case are retrieved; through the push suggestion unit, the obtained case information is aggregated, and according to the output prompt words, legal suggestions in a specified format are generated to be pushed to the political and legal staff, who can view them on the client 401.
[0094] Preferably, when the generated legal advice is displayed on the client 401, the disputed facts of the current case, applicable legal provisions, similar precedents, etc. can be highlighted. This application can reduce the analysis burden of political and legal staff, such as clerks, by semantically marking the text content, significantly improve the case handling efficiency, and ensure that case information is not omitted during the recording process, thereby improving the accuracy of the record.
[0095] It should be noted that the retriever in the case analysis model can also be used to interact intelligently with users and provide users with intelligent retrieval functions, including legal retrieval, case retrieval, Internet retrieval, and calling other types of retrieval tools. Figure 7 As shown, the client 401 can provide a user interface. After the user inputs, the server 402 inputs the large language model according to the current user's question and answer history and the prompt words to understand the user input and rewrite the user input, thereby generating accurate search information; the large language model is used to route the input questions and execute the corresponding processes according to the content of the questions; the large language model combines the results of these processes and aggregates the final output to feedback to the user. Among them, each execution chain includes but is not limited to: Internet search, knowledge base case base search (legal knowledge base search, case knowledge base search), case intelligent analysis and suggestions, tool call, etc.
[0096] Therefore, this application can automatically recommend relevant legal provisions based on the case description input by the user, and associate similar cases to provide effective legal references for political and legal personnel; and through intelligent retrieval and associated recommendations, the performance of the intelligent transcript generation system in terms of the standardization of case data management, retrieval efficiency, and case handling efficiency is also improved.
[0097] The document generation module 4023 is connected to the case analysis module 4022, and is used to generate a specified type of transcript document based on the generated speech text data of each speaker, the basic case information of the current case and legal advice, and send it to the client 401 for display.
[0098] In this embodiment, the document generation module 4023 can automatically generate standardized transcript documents without manual intervention, effectively shortening the case handling time and reducing the workload of manual sorting; and the document generation module 4023 can ensure the uniformity and professionalism of the transcript document format.
[0099] It should be understood that the specific process of each module executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0100] It should also be understood that the division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present application may be integrated into a processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0101] Figure 8 8 is a schematic diagram of the structure of the intelligent transcript generation terminal 800 provided in the embodiment of the present application. Figure 8 As shown, the intelligent transcript generation terminal 800 includes: at least one processor 801, a memory 802, at least one network interface 803 and a user interface 805. The various components in the device are coupled together through a bus system 804. It can be understood that the bus system 804 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 804 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 8 In the specification, various buses are labeled as bus systems.
[0102] The user interface 805 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.
[0103] It is understood that the memory 802 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of exemplary but not limiting explanation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present application is intended to include but is not limited to these and any other suitable categories of memory.
[0104] The memory 802 in the embodiment of the present application is used to store various categories of data to support the operation of the intelligent transcript generation terminal 800. Examples of these data include: any executable program for operating on the intelligent transcript generation terminal 800, such as an operating system 8021 and an application 8022; the operating system 8021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 8022 may include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The intelligent transcript generation method provided in any of the embodiments of the present application may be included in the application 8022.
[0105] The method disclosed in the above-mentioned embodiment of the present application can be applied to the processor 801, or implemented by the processor 801. The processor 801 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the intelligent transcript generation method can be completed by the hardware integrated logic circuit in the processor 801 or the instructions in the form of software. The above-mentioned processor 801 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 801 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor 801 can be a microprocessor or any conventional processor, etc. In combination with the steps of the intelligent transcript generation method provided in the embodiments of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the information in the memory and completes the steps of the aforementioned method in combination with its hardware.
[0106] In an exemplary embodiment, the intelligent transcript generation terminal 800 can be composed of one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD) to execute any one of the intelligent transcript generation methods provided in the above embodiments.
[0107] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0108] In summary, the present application provides an intelligent transcript generation method, system and terminal, which identifies the speakers of multiple collected original question-and-answer voice data based on a pre-trained speech recognition model, generates the speech text data of each speaker, so that the speeches of all parties can be recorded in real time, and supports multi-role recognition, and automatically annotates the speech text data of different speakers; identifies the basic case information of the current case based on a pre-trained case analysis model, determines the applicable legal provisions and similar precedents, and generates corresponding legal suggestions, thereby providing effective legal references for political and legal personnel and ensuring the consistency of legal application; by automatically generating transcript documents of specified types, the uniformity and professionalism of the transcript document format are ensured, and the case handling time is effectively shortened and the case handling efficiency is improved.
[0109] Therefore, the present application effectively overcomes various shortcomings in the prior art and has high industrial utilization value.
[0110] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.
Claims
1. An intelligent transcript generation method, characterized in that: include: Collect multiple original question-and-answer voice data in real time, and identify the speakers of each original question-and-answer voice data based on the pre-trained speech recognition model, and generate speech text data of each speaker; Based on the pre-trained case analysis model, the basic information of the current case is identified according to the speech text data of each speaker, and the applicable legal provisions and similar precedents are determined, and legal suggestions for the current case are generated so that the political and legal staff can continue to enforce the law according to the legal suggestions; Generate a specified type of transcript document based on the generated speech text data of each speaker, basic case information of the current case, and legal advice.
2. The intelligent transcript generation method according to claim 1, characterized in that: Based on the pre-trained speech recognition model, the speaker of each original question-and-answer speech data is identified, and the speech text data of each speaker is generated in the following manner: Transcribing each original question-and-answer voice data into text to generate corresponding initial speech text data, and performing grammatical and semantic correction on each initial speech text data to generate corresponding speech text data; Extracting voiceprint features from each original question-and-answer voice data to generate corresponding voiceprint feature data; Matching each generated voiceprint feature data with each registered voiceprint feature data in the voiceprint feature library to determine the speaker of each original question and answer voice data and the speaker's role identity; Each speaker is matched with each speaker's text data, and the speaker corresponding to each speaker's text data is marked.
3. The intelligent transcript generation method according to claim 2, characterized in that: The method of performing grammatical correction and semantic correction on the initial speech text data includes: Performing semantic structure analysis on the initial speech text data, and automatically segmenting the initial speech text data according to pauses in the original question-and-answer voice data to obtain one or more text sentences; Perform grammar check on each text sentence and make grammar corrections for detected grammatical errors; The context information of the initial speech text data is obtained, and based on the context information, the context analysis and semantic understanding of each text sentence after grammatical error correction are performed, homophones, ambiguous words and spelling errors therein are corrected, repeated content therein is screened out, missing information is supplemented, and speech text data is generated.
4. The intelligent transcript generation method according to claim 2, characterized in that: The generated voiceprint feature data is matched with each registered voiceprint feature data in the voiceprint feature library to determine the speaker of the original question-and-answer voice data and the role identity of the speaker, including: Calculate the similarity between the voiceprint feature data and each registered voiceprint feature data in the voiceprint feature library respectively, and obtain the registered voiceprint feature data with the largest similarity value as the target registered voiceprint feature data; If the similarity value is greater than or equal to a preset similarity threshold, the speaker of the target registered voiceprint feature data and the role identity of the speaker are obtained, and the speaker is used as the speaker of the voiceprint feature data; If the similarity value is less than the preset similarity threshold, the speaker of the voiceprint feature data is determined to be a new speaker. After the political and legal system staff confirms the role identity of the new speaker, the voiceprint feature data and the confirmed new speaker are stored in the voiceprint feature library.
5. The intelligent transcript generation method according to claim 1, characterized in that: Based on the pre-trained case analysis model, the basic information of the current case is identified according to the speech text data of each speaker, and the applicable legal provisions and similar precedents are determined. The legal suggestions for the current case are generated in the following ways: Extract keywords from each speech text data respectively to identify basic case information of the current case; wherein the basic case information includes: the time of the incident, the location of the incident, the persons involved in the case, one or more case facts, and one or more case evidences; Conduct logical analysis on the basic case information obtained, confirm the time relationship and causal relationship between the facts of each case, and obtain the course of the current case; Conduct semantic analysis on the basic case information obtained, identify disputed facts and undisputed facts in each case, and confirm the relationship between the evidence in each case; Conduct legal reasoning for undisputed facts and multiple dispute directions of disputed facts, generate multiple corresponding reasoning paths, search the pre-built legal knowledge base to obtain the legal provisions applicable to each reasoning path, and search the pre-built case vector database and case knowledge base to obtain similar cases for each reasoning path; Generate legal advice on the current case in a corresponding format based on the preset transcript type.
6. The intelligent transcript generation method according to claim 5, characterized in that: Methods for searching the pre-built case vector database and case knowledge base to obtain similar cases for each reasoning path include: Generate corresponding case summary information based on the undisputed facts of the current case and multiple dispute directions of disputed facts; Obtaining case summary information of each historical case in the case vector database, and calculating the similarity between the current case and each historical case respectively; Select one or more historical cases whose similarity values are greater than or equal to a preset similarity threshold as similar cases of the current case; The case knowledge base is searched to obtain basic case information, case judgment results and applicable legal provisions of each similar case.
7. The intelligent transcript generation method according to claim 1, characterized in that: Also includes: Retrieve relevant legal provisions and similar cases based on one or more keywords entered by the user; The specific methods include: Traverse the input information of the same user, obtain the context information of the current input, and perform context analysis and semantic understanding on each keyword currently input based on the context information, correct homophones, ambiguous words and spelling errors, and obtain rewritten search information; According to the search information, a pre-built legal knowledge base is searched to obtain relevant legal provisions; According to the search information, a pre-constructed case vector database is searched to obtain one or more similar cases and case summary information of each similar case; Based on the pre-built case knowledge graph, the related cases of each similar case are obtained; wherein the graph nodes of the case knowledge graph are each historical case, and the edges are the similarity relationships between each historical case; Search the pre-built case knowledge base to obtain basic case information, case judgment results, and applicable legal provisions of similar and related cases; Output similar cases and related cases in the form of semantic networks.
8. The intelligent transcript generation method according to claim 1, characterized in that: Also includes: The speech text data of each speaker, basic case information of the current case, legal advice, and generated transcript documents are stored in a pre-built text database; After the current case is closed, the basic case information, case judgment results and applicable legal provisions of the current case are stored in the pre-built case knowledge base; Generate corresponding case summary information based on the basic case information, case judgment results and applicable legal provisions of the current case, and perform vector conversion on the case summary information based on the pre-built embedding model to generate corresponding semantic vector data to be stored in the pre-built case vector database; Update the pre-built case knowledge graph based on the similarity relationship between the current case and each historical case in the case vector database.
9. An intelligent transcript generation system, characterized in that: include: Client and server side; The server is connected to the client via communication; the server includes: A speech recognition module, for identifying the speakers of a plurality of original question-and-answer speech data collected in real time by the client based on a pre-trained speech recognition model, and generating speech text data of each speaker to be sent to the client for display; A case analysis module, connected to the speech recognition module, is used to identify the basic information of the current case based on the pre-trained case analysis model and the speech text data of each speaker, and determine the applicable legal provisions and similar precedents, generate legal suggestions for the current case, so that the political and legal staff can continue to enforce the law according to the legal suggestions, and display them through the client; The document generation module is connected to the case analysis module and is used to generate a specified type of transcript document based on the generated speech text data of each speaker, the basic case information of the current case and the legal advice, and send it to the client for display.
10. An intelligent transcript generation terminal, characterized in that: include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory so that the terminal executes the intelligent transcript generation method as described in any one of claims 1 to 8.
Citation Information
Cited By
Intelligent auditing method and system for digital file and storage medium
CN120807231A
Conversation contradiction recognition method and system based on thinking chain of acoustic model and large language model
CN121884791A