A method for investigating the interrogation record and the synchronous recording and video content of the link

By using large language models and text matching algorithms to conduct consistency reviews of interrogation transcripts and synchronized audio and video recordings, the problem of low efficiency in traditional manual review is solved, realizing an automated, consistent, and efficient review method to ensure the authenticity of evidence and litigation efficiency.

CN121233742BActive Publication Date: 2026-02-06MICROPATTERN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411592189.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2026-02-06
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

Traditional manual review methods suffer from manpower shortages and low review efficiency when faced with a large number of interrogation transcripts, making it difficult to achieve automated review of the consistency between interrogation transcripts and synchronous audio and video recordings during the investigation process.

Method used

Large language models and text matching algorithms are used to conduct consistency reviews of interrogation transcripts and synchronized audio and video recordings. The dialogue is converted into text, the information is structured, and consistency is judged using large language models and pinyin comparison or text matching algorithms.

Benefits of technology

It has enabled automated review of the consistency of interrogation transcripts and synchronous audio and video recordings during the investigation process, improving review efficiency and accuracy, reducing manpower shortages, and ensuring the authenticity of evidence and litigation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121233742B_ABST
    Figure CN121233742B_ABST
Patent Text Reader

Abstract

The application discloses a kind of methods for examining the contents of interrogation records and synchronous audio and video recordings in investigation links, and relates to the technical field of intelligent examination of consistency of interrogation records and synchronous audio and video recordings in investigation links.By converting the dialogue between the interrogated person and the person being interrogated recorded by synchronous audio and video recordings in investigation links into text, the first part and the main text of the interrogation record are respectively structured.Then, the consistency of the first part information of the interrogation record and the contents of synchronous audio and video recordings is examined by using a large language model and pinyin comparison.The consistency of the main text of the interrogation record and the contents of synchronous audio and video recordings is examined by using a large language model and a text matching algorithm, i.e., whether the text information recorded in the main text of the interrogation record and the audio content recorded in synchronous audio and video recordings are consistent.The standard for consistency is that the meanings are the same, and it is not required that the texts are exactly the same word by word.The method realizes the automatic examination of the consistency of interrogation records and synchronous audio and video recordings in investigation links, and is accurate and efficient.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent examination of consistency between interrogation records and synchronous audio-video recording contents in the investigation link, and particularly relates to a method for examining the consistency between interrogation records and synchronous audio-video recording contents in the investigation link. BACKGROUND

[0002] Interrogation records are important evidence materials in the investigation link, and the authenticity and objectivity of their contents are crucial for the trial of a case. Synchronous audio-video recording in the investigation link refers to real-time and uninterrupted synchronous audio-video recording of the interrogation process between the interrogator and the person being interrogated in the investigation link, in order to ensure the legality and fairness of the investigation process. The use of this technical means aims to record the overall situation of the investigation interrogation truly and completely, so as to be used as evidence in the subsequent litigation process.

[0003] The importance of the examination of the consistency between interrogation records and synchronous audio-video recording contents in the investigation link cannot be ignored, mainly reflected in the following aspects: 1) ensuring the authenticity of evidence: interrogation records and synchronous audio-video recording are both important evidence materials, used to prove the statements of the person being interrogated and the legality of the investigation link. Through the examination of the consistency of their contents, it can be ensured that the recorded statements are true and accurate, thereby ensuring the authenticity and credibility of the evidence. 2) improving litigation efficiency: in the trial process, if the contents of the interrogation records and the synchronous audio-video recording are inconsistent, it may cause the evidence chain to be broken, affecting the trial and judgment of the case. Through the examination of the consistency, problems can be found and solved in time, avoiding unnecessary disputes and delays in the litigation process, and improving the litigation efficiency. 3) promoting judicial justice: through the examination of the consistency between the contents of the interrogation records and the synchronous audio-video recording, it can ensure the fairness and legality of the trial of the case, reduce the occurrence of false cases, and maintain social fairness and justice. 4) improving the standardization of the investigation link: the investigation link is an important link of criminal litigation, and its standardization directly affects the quality and efficiency of the trial of the case. Through the examination of the consistency between the contents of the interrogation records and the synchronous audio-video recording, problems and deficiencies in the investigation link can be found in time, promoting the standardization and professionalization of the investigation link.

[0004] However, the traditional manual examination method has the problems of manpower shortage and low examination efficiency when facing a large number of interrogation records. In some major cases, there are many people involved, many interrogation records, and long interrogation time, and it is difficult to examine them one by one due to manpower shortage.

[0005] Therefore, it is necessary to propose a method and a supporting hardware system for automatically checking the consistency of interrogation records and synchronous recording and video content in the investigation link. The consistency of the first information of the interrogation record and the synchronous recording and video content is checked, and the standard for consistency is that there is no difference in words; the consistency of the text of the interrogation record and the synchronous recording and video content is checked, and there is a certain subjective factor in the judgment of the objectivity of the content, that is, the record content should be consistent with the meaning of the confession of the person being interrogated, but it does not need to reach the degree of complete consistency, one word does not differ. SUMMARY

[0006] In order to solve the technical problem of automatic consistency checking of interrogation records and synchronous recording and video content in the investigation link, the present application provides a method for checking interrogation records and synchronous recording and video content in the investigation link. The following technical solutions are adopted:

[0007] A method for checking interrogation records and synchronous recording and video content in the investigation link, comprising the following steps:

[0008] Step 1, converting the dialogue between the interrogator and the person being interrogated recorded in the synchronous recording and video in the investigation link into text;

[0009] Step 2, information structuring of the interrogation record; text recognition of the interrogation record; information structuring of the first part of the interrogation record, and the extracted key-value pair is represented as , M is the number of key-value pairs, is the i-th key, is the value corresponding to the i-th key; information structuring of the text of the interrogation record, and the extracted question-answer pair is represented as , N is the number of question-answer pairs, is the i-th question, is the answer to the i-th question;

[0010] Step 3, consistency checking of the interrogation record and the synchronous recording and video content, comprising the following sub-steps:

[0011] Step 31, consistency checking of the first part of the interrogation record and the synchronous recording and video content:

[0012] The key-value pair of the first part of the interrogation record is extracted from the dialogue text of the interrogator and the person being interrogated by using a large language model, and the value corresponding to each key is marked as , and and obtained in step 2 are compared, and the comparison result is obtained based on whether the pinyin of and is the same;

[0013] Step 32: consistency checking of the text of the interrogation record and the synchronous recording and video content:

[0014] Extracting answers of each question in the text part of the interrogation record from the dialogue text of the interrogator and the interrogated person by using a large language model , and comparing the answers obtained by the large language model with the answers recorded on the interrogation record by using a text matching algorithm , and determining whether the content in the interrogation record text and the synchronous recording and video is consistent according to a threshold.

[0015] By adopting the above technical solutions, the dialogue between the interrogator and the interrogated person recorded by the synchronous recording and video in the investigation link is converted into text, and the first part and the text part of the interrogation record are respectively structured. Then, the consistency of the first part information of the interrogation record and the content of the synchronous recording and video is checked by using a large language model and pinyin comparison, i.e., whether the voice content recorded in the synchronous recording and video is consistent with the first part information recording the text part in the interrogation record. The consistency standard is that there is no difference in words. The consistency of the text part of the interrogation record and the content of the synchronous recording and video is checked by using a large language model and a text matching algorithm, i.e., whether the voice content recorded in the synchronous recording and video is consistent with the text information recording the text part in the interrogation record. The consistency standard is that the meanings are the same, and it is not required that there is no difference in words. The consistency of the interrogation record and the content of the synchronous recording and video in the investigation link is checked automatically, which is accurate and efficient.

[0016] Optionally, step 1 includes the following sub-steps:

[0017] Step 10: extracting audio from the audio and video file obtained from the synchronous recording and video in the investigation link;

[0018] Step 11: converting the audio into text by using a speech transcription algorithm.

[0019] By adopting the above technical solutions, the audio is converted into text by using a speech transcription algorithm. It is not limited to using which speech transcription algorithm, for example, a two-pass method can be used to convert the streaming and non-streaming end-to-end speech recognition model (Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition).

[0020] Optionally, step 2 includes the following sub-steps:

[0021] Step 20: performing character recognition on the interrogation record by using a character recognition algorithm;

[0022] Step 21: structuring the first part information of the interrogation record, and extracting the key-value pair information of the first part of the interrogation record from the character recognition result by using an information extraction algorithm, and representing the extracted key-value pair as , M is the number of key-value pairs,​​ is the i-th key, is the value corresponding to the i-th key;

[0023] Step 22: Structuring the main text information of the interrogation record. Regular expression method is used to extract the question and answer pairs in the interrogation record, and the extracted question and answer pairs are expressed as N is the number of question and answer pairs, is the i-th question, is the answer to the i-th question.

[0024] Any text recognition algorithm can be used for text recognition of the interrogation record, for example, CRNN (Convolutional Recurrent Neural Network) algorithm can be used. CRNN is a representative framework in the field of text recognition, which combines convolutional neural network (CNN) and recurrent neural network (RNN), and is particularly suitable for recognizing sequential text in images.

[0025] Step 21: Structuring the main text information of the interrogation record. The main text information of the interrogation record generally includes record name, interrogation times, interrogation time and place, interrogator's name, recorder's name, and the name of the person being interrogated. Step 21 needs to extract these information from the text recognition result. Any information extraction algorithm can be used, for example, the universal information extraction framework UIE algorithm (Universal Information Extraction, Yaojie Lu, ACL-2022) can be used. This framework realizes the unified modeling of entity extraction, relationship extraction, event extraction, sentiment analysis and other tasks, and makes different tasks have good migration and generalization ability. The extracted key-value pairs are expressed as M is the number of key-value pairs, is the i-th key, is the value corresponding to the i-th key.

[0026] Step 22: Structuring the main text information of the interrogation record. The main text information of the interrogation record adopts the form of question and answer, and has a fixed and obvious "question: xxx. Answer: xxx" format in writing. Therefore, regular expression method can be used to extract the question and answer pairs in the interrogation record, and the extracted question and answer pairs are expressed as N is the number of question and answer pairs, is the i-th question, is the answer to the i-th question.

[0027] Optionally, in step 20, if the interrogation record is an electronic version, the content of the electronic version of the interrogation record is directly read.

[0028] By adopting the technical scheme, if the interrogation record is not a paper version but an electronic version, OCR recognition is not needed, and only the content of the electronic version of the interrogation record needs to be directly read.

[0029] Optionally, the large language model in step 3 is a large language model provided by a commercial company, an open source large language model, or a large language model fine-tuned by using the interrogation record and the synchronous audio and video data.

[0030] Optionally, in step 32, no text matching algorithm is limited to be used, for example, a semantic matching model ernie_matching based on a single tower Point-wise mode or a semantic matching model ernie_matching based on a single tower Pair-wise mode can be used.

[0031] A computer-readable storage medium stores a content consistency review program designed by using a review method of interrogation record and synchronous audio and video content in an investigation link.

[0032] An interrogation record and synchronous audio and video content consistency review system in an investigation link includes a data input module, a computer-readable storage medium, and a processor. The data input module inputs the interrogation record and the synchronous audio and video content in the investigation link into the computer-readable storage medium. The computer-readable storage medium stores a content consistency review program designed by using a review method of interrogation record and synchronous audio and video content in an investigation link. The processor calls the interrogation record and the synchronous audio and video content in the investigation link, runs the content consistency review program, and obtains a content consistency review result.

[0033] Optionally, the system further includes a display, and the processor controls the display to display the content consistency review result.

[0034] In summary, the present application has the following at least beneficial technical effects:

[0035] The present application can provide a kind of investigation link interrogation record, synchronous audio and video content examination method, the dialogue of the interrogation person and the person interrogated recorded in synchronous audio and video recording of investigation link is converted into text, the first part and the main body of interrogation record are structured respectively;Then, the first part information of interrogation record and synchronous audio and video content are checked for consistency using large language model and pinyin comparison, that is, whether the voice content recorded in the first part information of the main body recorded in interrogation record and synchronous audio and video is consistent, and the standard of consistency is that there is no difference in words;The main body of interrogation record and synchronous audio and video content are checked for consistency using large language model and text matching algorithm, that is, whether the voice content recorded in the main body recorded in interrogation record and synchronous audio and video is consistent, and the standard of consistency is that meaning is the same, not requiring that there is no difference in words, which realizes automatic consistency check of interrogation record and synchronous audio and video content in investigation link, and is accurate and efficient. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 It is a flowchart of the interrogation record, synchronous audio and video content examination method of the present application in investigation link. DETAILED DESCRIPTION

[0037] The present application will be further described in detail below with reference to the accompanying drawings. The present application discloses a kind of interrogation record, synchronous audio and video content examination method in investigation link.

[0038] REFERENCE Figure 1 A kind of interrogation record, synchronous audio and video content examination method in investigation link, comprising the following steps:

[0039] Step 1, the dialogue of the interrogation person and the person interrogated recorded in synchronous audio and video recording of investigation link is converted into text;

[0040] Step 2, information structure is carried out to interrogation record. Text recognition is carried out to interrogation record, the first part information structure of interrogation record, and the key-value pair extracted is represented as M is the number of key-value pairs, is the i th key, is the value corresponding to the i th key;

[0041] Step 3, consistency check of interrogation record and synchronous audio and video content, comprising the following sub-steps:

[0042] Step 31, consistency check of the first part of interrogation record and synchronous audio and video content:

[0043] The value corresponding to each key of the first part of interrogation record is extracted from the dialogue text of the interrogation person and the person interrogated using large language model, marked as and compared with obtained in step 21, and the comparison result is obtained based on whether the pinyin of and is same;

[0044] Step 32: consistency check of the interrogation record text and the synchronous recording and video content, using a large language model to extract the answers to each question in the interrogation record text from the dialogue text of the interrogator and the interrogated person, and using a text matching algorithm to compare the answers obtained by the large language model with the answers recorded in the interrogation record, and determining whether the interrogation record text and the content in the synchronous recording and video are consistent according to a threshold.

[0045] By adopting the above technical solution, the dialogue between the interrogator and the interrogated person recorded by the synchronous recording and video in the investigation link is converted into text, and the first part and the text part of the interrogation record are respectively structured. Then, the consistency of the first part information of the interrogation record and the content of the synchronous recording and video is checked by using a large language model and pinyin comparison, i.e., whether the first part information of the interrogation record and the voice content recorded in the synchronous recording and video are consistent, and the standard for consistency is that there is no difference in words. The consistency of the interrogation record text and the content of the synchronous recording and video is checked by using a large language model and a text matching algorithm, i.e., whether the text information recorded in the interrogation record and the voice content recorded in the synchronous recording and video are consistent, and the standard for consistency is that the meanings are the same, and it is not required that there is no difference in words. The consistency of the interrogation record, the synchronous recording and video content in the investigation link is automatically checked, which is accurate and efficient.

[0046] Step 1 includes the following sub-steps:

[0047] Step 10: extracting audio from the audio and video files obtained from the synchronous recording and video in the investigation link;

[0048] Step 11: converting the audio into text using a speech transcription algorithm.

[0049] By adopting the above technical solution, the audio is converted into text using a speech transcription algorithm. For example, a two-pass method can be used to convert the audio into text, and a unified streaming and non-streaming two-pass end-to-end model for speech recognition can be used.

[0050] Step 2 includes the following sub-steps:

[0051] Step 20: performing text recognition on the interrogation record using a text recognition algorithm;

[0052] Step 21: information structuring of the first part of the interrogation record, extracting the first part information of the interrogation record from the text recognition result using an information extraction algorithm, and representing the extracted key-value pairs as , M is the number of key-value pairs, is the i-th key, is the value corresponding to the i-th key.

[0053] Step 22: Structuring the main text information of the interrogation record, using a regular expression method to extract the question and answer pairs in the interrogation record, and expressing the extracted question and answer pairs as , N is the number of question and answer pairs, is the ith question, is the answer to the ith question.

[0054] By using the above technical solution, the text recognition of the interrogation record can use the CRNN (Convolutional Recurrent Neural Network) algorithm, which is a representative framework in the field of text recognition. It combines convolutional neural network (CNN) and recurrent neural network (RNN), and is particularly suitable for recognizing sequential text in images.

[0055] Step 21: Structuring the main text information of the interrogation record. The main text information of the interrogation record generally includes the name of the record, the number of interrogations, the time and place of the interrogation, the name of the interrogator, the name of the recorder, the name of the person being interrogated, etc. Step 21 needs to extract these information from the text recognition result. It is not limited to using any information extraction algorithm, for example, the universal information extraction framework UIE algorithm (Universal Information Extraction, Yaojie Lu, ACL-2022) can be used. This framework realizes the unified modeling of entity extraction, relationship extraction, event extraction, sentiment analysis and other tasks, and enables good migration and generalization between different tasks. The extracted key-value pairs are expressed as , M is the number of key-value pairs, is the ith key, is the value corresponding to the ith key.

[0056] Step 22: Structuring the main text information of the interrogation record. The main text of the interrogation record adopts the form of question and answer, and has a fixed and obvious format of "question: xxx. Answer: xxx" in writing. Therefore, a regular expression method can be used to extract the question and answer pairs in the interrogation record, and the extracted question and answer pairs are expressed as , N is the number of question and answer pairs, is the ith question, is the answer to the ith question.

[0057] In step 20, if the interrogation record is an electronic version, the content of the electronic version of the interrogation record is directly read.

[0058] If the interrogation record is not a paper version, but an electronic version, then no OCR recognition is needed, and only the content of the electronic version of the interrogation record needs to be directly read.

[0059] The large language model in step 3 is a large language model provided by a commercial company, an open source large language model, or a large language model fine-tuned on the interrogation record, synchronous audio and video data.

[0060] In step 32, the text matching algorithm is a semantic matching model ernie_matching based on a single tower Point-wise mode, or a semantic matching model ernie_matching based on a single tower Pair-wise mode.

[0061] A computer-readable storage medium stores a content consistency review program designed by an interrogation record, synchronous audio and video content review method.

[0062] An interrogation record, synchronous audio and video content consistency review system includes a data entry module, a computer-readable storage medium, and a processor. The data entry module enters the interrogation record and the synchronous audio and video content into the computer-readable storage medium. The computer-readable storage medium stores a content consistency review program designed by an interrogation record, synchronous audio and video content review method. The processor calls the interrogation record and the synchronous audio and video content, runs the content consistency review program, and obtains a content consistency review result.

[0063] It also includes a display controlled by the processor to display the content consistency review result.

[0064] The following specific implementation cases are used to illustrate the implementation principle of an interrogation record, synchronous audio and video content review method and system:

[0065] In order to make the purpose, technical solution and advantages of the present application clearer and more obvious, the following interrogation record is fictitious, all information is fictitious and does not correspond to the real case.

[0066] The specific interrogation record is as follows:

[0067] First

[0068] Interrogation record

[0069] Time May 21, 2024 15:10-16:05

[0070] Place xx City Public Security Bureau xx Police Station

[0071] Interrogator Zhang San Work unit xx Police Station

[0072] Recorder Li Si Work unit xx Police Station

[0073] Suspect Wang Wu Gender Male Age 24 Date of Birth xx month xx day, 2000

[0074] Suspect ID Type and Number ID Card 1xxxxxxxxxxxxxxxx4

[0075] Suspect Current Address Building 2, Unit 3, Floor 4, Room 4xx, First Street, xx City

[0076] Question: We are police officers from the xx City Public Security Bureau, and we are now legally interrogating you. Please answer our questions truthfully. Did you hear that?

[0077] Answer: I heard it.

[0078] Question: What is your name?

[0079] Answer: My name is Wang Wu.

[0080] Question: Gender?

[0081] Answer: Male.

[0082] Question: Date of birth?

[0083] Answer: xx month xx day, 2000.

[0084] Question: What is your ID number?

[0085] Answer: 1xxxxxxxxxxxxxxxx4.

[0086] Question: What is your home address?

[0087] Answer: Building 2, Unit 3, Floor 4, Room 4xx, First Street, xx City.

[0088] Question: Do you know why we arrested you?

[0089] Answer: I don't know.

[0090] Question: Look at this photo. Whose phone is this in the photo?

[0091] Answer: It's Zhao Liu's phone.

[0092] Question: Why is there your fingerprint on this phone?

[0093] Answer: I admit to stealing Zhao Liu's phone.

[0094] Page 1 of 1

[0095] Step 1: Convert the dialogue between the interrogator and the suspect recorded in the synchronous audio and video of the investigation link into text.

[0096] Step 10: Extract audio from the audio and video files obtained from the synchronous recording of the investigation link.

[0097] Step 11: Convert the audio to text using a speech transcription algorithm. Specifically, a two-pass method is used in the embodiment to convert streaming and non-streaming end-to-end speech recognition models (Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition).

[0098] An example of obtaining a speech transcription result is as follows:

[0099] We are the police of xx city public security bureau, now we will interrogate you according to law, you must answer our questions truthfully, do you understand? I understand. What's your name? I'm Wang Wu. Gender? Male. Date of birth? 2000 xx month xx day. What is your ID number? 1xxxxxxxxxxxxxxxx4. What is your home address? No. 2, Building 3, Unit 4, Floor 4, Room xx, First Avenue, xx City. Do you know why we arrested you? I don't know. Show me a photo, do you know whose phone this is? I think it's my colleague Zhao Liu's phone. Zhao Liu's phone has your fingerprints, why? Officer, I was wrong. I admit that after I picked up Zhao Liu's phone, I didn't return it to him.

[0100] Step 2: Information structure of the interrogation record.

[0101] Step 20: Text recognition of the interrogation record. Here, CRNN (Convolutional Recurrent Neural Network) algorithm is used, which is a representative framework in the field of text recognition, combining convolutional neural network (CNN) and recurrent neural network (RNN), especially suitable for recognizing sequential text in images.

[0102] Step 21: Structuring the first information of the interrogation record. The first information of the interrogation record generally includes record name, interrogation times, interrogation time and place, interrogator's name, recorder's name, and the name of the person being interrogated. Step 21 needs to extract these information from the text recognition result. It is not limited to using which information extraction algorithm, for example, the universal information extraction framework UIE (Universal Information Extraction, Yaojie Lu, ACL-2022) can be used, which realizes the unified modeling of entity extraction, relationship extraction, event extraction, sentiment analysis and other tasks, and makes different tasks have good migration and generalization ability. The extracted key-value pairs are represented as , M is the number of key-value pairs, is the i-th key, is the value corresponding to the i-th key.

[0103] In this embodiment, the extracted key-value pairs are as follows:

[0104] : Time

[0105] : 15:10, May 21, 2024 to 16:05, May 21, 2024

[0106] : Place

[0107] : xx City Public Security Bureau xx Police Station

[0108] : Interrogator's name

[0109] : Zhang San

[0110] : Interrogator's work unit

[0111] : xx Police Station

[0112] : Recorder's name

[0113] : Li Si

[0114] : Recorder's work unit

[0115] : xx Police Station

[0116] : Interrogated person's name

[0117] : Wang Wu

[0118] : Interrogated person's gender

[0119] : Male

[0120] : Interrogated person's age

[0121] : 24

[0122] : Interrogated person's date of birth

[0123] : 2000 xx xx

[0124] : The interrogated person's identity card type and number

[0125] : ID card 1xxxxxxxxxxxxxxxx4

[0126] : The interrogated person's current address

[0127] : Building 2, Unit 3, Floor 4, Room 4xx, First Street, xx City

[0128] Step 22: Structuring the interrogation record text information. The interrogation record text adopts the form of question and answer, and has a fixed and obvious "question: xxx. answer: xxx" format in writing. Therefore, the method of regular expression can be used to extract the question and answer pairs in the interrogation record, and the extracted question and answer pairs are represented as , N is the number of question and answer pairs, is the i-th question, is the answer to the i-th question.

[0129] This embodiment uses python language programming to realize the method of regular expression to extract the question and answer pairs in the interrogation record, and the code segment is as shown below:

[0130] import re

[0131] text = "question: we are the police of xx city public security bureau, now we will interrogate you according to law, you must answer our questions truthfully, have you heard clearly? Answer: I have heard clearly. Question: what is your name? Answer: my name is wangwu. Question: gender? Answer: male. Question: date of birth? Answer: 2000 xx xx. Question: what is your identity card number? Answer: 1xxxxxxxxxxxxxxxx4. Question: what is your home address? Answer: building 2, unit 3, floor 4, room 4xx, first street, xx city. Question: do you know why we arrested you? Answer: I don't know. Question: show you a photo, whose phone is this? Answer: it's Zhao Liu's phone. Question: why is there your fingerprint on this phone? Answer: I admit stealing Zhao Liu's phone. "

[0132] pattern = re.compile(r'question: (.*?) answer: ((?(?!question: ).) *)', re.MULTILINE | re.DOTALL)

[0133] matches = pattern.findall(text)

[0134] In this embodiment, the extracted question-answer pairs are as follows:

[0135] : We are the police of xx city public security bureau, now according to law to you for interrogation, you must answer our questions, listen clearly?

[0136] : I listen clearly.

[0137] : What's your name?

[0138] : My name is Wangwu.

[0139] : Gender?

[0140] : Male.

[0141] : Date of birth?

[0142] : 2000 xx month xx day.

[0143] : What is your ID number?

[0144] : 1xxxxxxxxxxxxxxxx4.

[0145] : What is your home address?

[0146] : No. 2 building, unit 3, floor 4, room 4xx, first street, xx city.

[0147] : Do you know why we arrested you?

[0148] : I don't know.

[0149] : Show you a photo, whose phone is this on the photo?

[0150] : It's Zhao Liu's phone.

[0151] : Why is there your fingerprint on this phone?

[0152] : I admit stealing Zhao Liu's phone.

[0153] Step 3: Inconsistency examination between interrogation record and synchronous recording content.

[0154] Step 31: Inconsistency examination between interrogation record header and synchronous recording content. As mentioned earlier, the information contained in the header is objective, and the standard for examination is that the text must be word-for-word. Using a large language model, extract each key value (labeled ) of the interrogation record header from the dialogue text of the interrogator and the person being interrogated, and compare it with and obtained in step 21. Considering that the voice transcription result and the text recognition result may be homophonic but different in characters, we do not directly compare and whether they are the same, but compare whether the pinyin of and is the same. Here, we do not limit the use of which large language model, it can be a large language model provided by a commercial company (such as Baidu's Wenxin Yiyang, Ali's Tongyi Qianwen), or an open source large language model (such as ChatGLM3, Baichuan2, etc.), or a large language model fine-tuned on the interrogation record, synchronous recording data.

[0155] In this embodiment, we constructed 10,000 interrogation record text recognition results and interrogation and dialogue text of the person being interrogated, and fine-tuned the open source large language model Bachuan2-7B.

[0156] For i in 1,2,3,...,M

[0157] prompt = "In the following dialogue between the interrogator and the person being interrogated, what is it?" text = Step 1 obtained interrogation and dialogue text of the person being interrogated

[0158] input_text = prompt + text

[0159] output = large_language_model.generate(input_text)

[0160] output_pingyin = output's pinyin

[0161]

[0162] _pinyin = 's pinyin

[0163] ​if output_pingyin == _pinyin

[0164] interrogation record first part of and consistent with the content of the synchronous recording video.

[0165] else

[0166] interrogation record first part of and inconsistent with the content of the synchronous recording video.

[0167] when i=7,

[0168] Prompt="In the following dialogue between the interrogator and the interrogatee, what is the name of the interrogatee?"

[0169] Text="We are the police of xx city public security bureau, now according to law to you for interrogation, you want to answer our questions truthfully, listen clearly? I listen clearly. What's your name? I'm called Wang Wu. Gender? Male. Date of birth? 2000 xx month xx day. What is your ID number? 1xxxxxxxxxxxxxxxx4. What is your family address? No. 2, building 3, unit 4, room 4, first street, xx city. Do you know why we arrested you? Not clear. Show you a photo, do you know whose phone this is? I think it's my colleague Zhao Liu's phone. Zhao Liu's phone has your fingerprint, why? Officer, I was wrong. I admit that I didn't return Zhao Liu's phone after I picked it up."

[0170] input_text="In the following dialogue between the interrogator and the interrogatee, what is the name of the interrogatee? We are the police of xx city public security bureau, now according to law to you for interrogation, you want to answer our questions truthfully, listen clearly? I listen clearly. What's your name? I'm called Wang Wu. Gender? Male. Date of birth? 2000 xx month xx day. What is your ID number? 1xxxxxxxxxxxxxxxx4. What is your family address? No. 2, building 3, unit 4, room 4, first street, xx city. Do you know why we arrested you? Not clear. Show you a photo, do you know whose phone this is? I think it's my colleague Zhao Liu's phone. Zhao Liu's phone has your fingerprint, why? Officer, I was wrong. I admit that I didn't return Zhao Liu's phone after I picked it up."

[0171] Output=Wang Wu

[0172] Output_pinyin=wangwu

[0173] _pinyin=wangwu

[0174] because output_pingyin == _pinyin, so the conclusion is that the and in the first part of the interrogation record are consistent with the content in the synchronous recording and video.

[0175] Step 32: Consistency review of interrogation record text and synchronous recording and video content. The interrogation record text is recorded by the recorder according to the content of the interrogation and the interrogation of the person being interrogated. It is impossible to be word-for-word, as long as the content is consistent. Use a large language model to extract the answers to each question in the interrogation record text from the dialogue text of the interrogator and the person being interrogated, and use a text matching algorithm to compare the answers obtained by the large language model with the answers recorded on the interrogation record. According to the threshold, determine whether the interrogation record text and the content in the synchronous recording and video are consistent. Similarly, this does not limit the use of any large language model, which can be a large language model provided by a commercial company (such as Baidu's Wenxin Yiyang, Ali's Tongyi Qianwen), or an open source large language model (such as ChatGLM3, Baichuan2, etc.), or fine-tune the open source large language model using interrogation record, synchronous recording and video data. It does not limit the use of any text matching algorithm, such as the single-tower Point-wise based semantic matching model ernie_matching, or the single-tower Pair-wise based semantic matching model ernie_matching.

[0176] In this embodiment, we construct a batch of interrogation record text recognition results and interrogation and dialogue text of the person being interrogated, and fine-tune the open source large language model Bachuan2-7B.

[0177] For i in 1,2,3,...,N

[0178] prompt = "In the following dialogue between the interrogator and the person being interrogated, how did the person being interrogated answer this question?"

[0179] text=Step 1 obtained interrogation and dialogue text of the person being interrogated

[0180] input_text=prompt + text

[0181] output=large_language_model.generate(input_text)​

[0182] score = TextMatching(output, )

[0183] if score > thr

[0184] the content in the interrogation record text and is consistent with the content in the synchronous recording video.

[0185] else

[0186] the content in the interrogation record text and is not consistent with the content in the synchronous recording video.

[0187] when i = 9, prompt = "In the following dialogue between the interrogator and the interrogatee, how did the interrogatee answer the question "Why is there your fingerprint on this phone?""

[0188] text = "We are police officers from the xx City Public Security Bureau, and we are now legally interrogating you. You must answer our questions truthfully. Did you hear that? I heard it. What is your name? My name is Wang Wu. Gender? Male. Date of birth? 2000 xx month xx day. What is your ID number? 1xxxxxxxxxxxxxxxx4. What is your home address? 2 No. 3 Unit 4 Layer 4 Room xx City First Street Building. Do you know why we arrested you? Not clear. Show you a photo, do you know whose phone this is? I think it should be my colleague Zhao Liu's phone. Zhao Liu's phone has your fingerprint, why? Officer, I was wrong. I admit I stole Zhao Liu's phone."

[0189] input_text = "In the following dialogue between the interrogator and the interrogatee, how did the interrogatee answer the question "Why is there your fingerprint on this phone?" We are police officers from the xx City Public Security Bureau, and we are now legally interrogating you. You must answer our questions truthfully. Did you hear that? I heard it. What is your name? My name is Wang Wu. Gender? Male. Date of birth? 2000 xx month xx day. What is your ID number? 1xxxxxxxxxxxxxxxx4. What is your home address? 2 No. 3 Unit 4 Layer 4 Room xx City First Street Building. Do you know why we arrested you? Not clear. Show you a photo, do you know whose phone this is? I think it should be my colleague Zhao Liu's phone. Zhao Liu's phone has your fingerprint, why? Officer, I was wrong. I admit I picked up Zhao Liu's phone and didn't give it back to him."

[0190] output = police officer, I was wrong. I admit that after I picked up Zhao's cell phone, I did not return it to him.

[0191] Score = TextMatching("police officer, I was wrong. I admit that after I picked up Zhao's cell phone, I did not return it to him.", "I admit stealing Zhao's cell phone."), score = 0.528;

[0192] Set thr = 0.9;

[0193] Because score < thr, the content in the interrogation record text and is inconsistent with the content in the synchronous recording video.

[0194] The above are preferred embodiments of the present application, not to limit the protection scope of the present application, therefore: any equivalent changes made according to the structure, shape, principle of the present application should be covered within the protection scope of the present application.

Claims

1. A method for reviewing interrogation transcripts and synchronized audio and video recordings during the investigation process, characterized in that, Includes the following steps: Step 1: Convert the dialogue between the interrogator and the interrogated person recorded in the synchronous audio and video recording during the investigation into text. Step 2: Structure the interrogation transcript. The interrogation transcript was subjected to text recognition; the header information of the interrogation transcript was structured, and the extracted key-value pairs were represented as follows: M is the number of key-value pairs. For the i-th key, The value corresponding to the i-th key; the interrogation transcript is structured, and the extracted question-and-answer pairs are represented as... N is the number of question-answer pairs. For the i-th question, This is the answer to the i-th question; Step 3, consistency review of the interrogation transcript and synchronized audio and video recordings, includes the following sub-steps: Step 31, consistency review of the interrogation transcript and the synchronized audio and video recordings: This method utilizes a large language model to extract each key from the beginning of the interrogation transcript. The corresponding value is marked as and to and the information obtained in step 2 Comparison, based on and To obtain the comparison result, check if the pinyin is the same; Step 32: Consistency review of the interrogation transcript and the synchronized audio and video recordings: Using a large language model, extract each question from the main body of the interrogation transcript. The answer And use text matching algorithms to analyze the answers obtained from the large language model. And the answers recorded in the interrogation transcript A comparison is made, and the content of the interrogation transcript and the synchronized audio and video recordings are judged to be consistent based on the threshold.

2. The method for reviewing interrogation transcripts and synchronized audio and video recordings during the investigation process according to claim 1, characterized in that, Step 1 includes the following sub-steps: Step 10: Extract audio from the audio and video files obtained from the synchronous audio and video recording during the investigation phase; Step 11: Use a speech-to-text algorithm to convert the audio into text.

3. The method for reviewing interrogation transcripts and synchronized audio and video recordings during the investigation process according to claim 1, characterized in that, Step 2 includes the following sub-steps: Step 20: Use a text recognition algorithm to perform text recognition on the interrogation transcript; Step 21: Structure the header information of the interrogation transcript. An information extraction algorithm is used to extract key-value pairs from the text recognition results of the header of the interrogation transcript. The extracted key-value pairs are represented as follows: M is the number of key-value pairs. For the i-th key, This is the value corresponding to the i-th key; Step 22: Structure the main body of the interrogation transcript. Use regular expressions to extract question-and-answer pairs from the interrogation transcript and represent the extracted question-and-answer pairs as follows: N is the number of question-answer pairs. For the i-th question, This is the answer to the i-th question.

4. The method for reviewing interrogation transcripts and synchronized audio and video recordings during the investigation process according to claim 3, characterized in that: In step 20, if the interrogation record is in electronic format, the content of the electronic version of the interrogation record is read directly.

5. The method for reviewing interrogation transcripts and synchronized audio and video recordings during the investigation process according to claim 1, characterized in that: The large language model in step 3 can be a large language model provided by a commercial company, an open-source large language model, or a large language model that has been fine-tuned using interrogation transcripts and synchronized audio and video recording data.

6. A computer-readable storage medium, characterized in that: The storage adopts a content consistency review procedure designed using the review method for interrogation transcripts and synchronous audio and video recordings during the investigation process as described in any one of claims 1-5.

7. A system for verifying the consistency of interrogation transcripts and synchronized audio and video recordings during the investigation process, characterized in that: The system includes a data entry module, a computer-readable storage medium, and a processor. The data entry module enters interrogation transcripts and synchronized audio and video recordings from the investigation process into the computer-readable storage medium. The computer-readable storage medium stores a content consistency review program designed using the review method for interrogation transcripts and synchronized audio and video recordings from the investigation process as described in any one of claims 1-5. The processor calls the interrogation transcripts and synchronized audio and video recordings from the investigation process, runs the content consistency review program, and obtains the content consistency review result.

8. The system for reviewing the consistency of interrogation transcripts and synchronized audio and video recordings during the investigation process according to claim 7, characterized in that: It also includes a display, the processor of which controls the display to show the content consistency review results.

Citation Information

Patent Citations

  • Interrogation information auditing method and device, computer device and storage medium

    CN109766474A

  • Apparatus and method for analysis of audio recordings

    US20220148584A1