Automatic record generation method based on AI voice interaction

Through time domain segmentation and deep neural network technology, speech features are automatically processed to generate formatted records, solving the problems of irregular transcript generation and insufficient information in the existing technology, and achieving efficient and informative automatic generation of transcripts.

CN120375809AActive Publication Date: 2025-07-25GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510602675.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-25
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The existing transcript generation methods are difficult to generate transcripts with standardized formats and rich information, especially in complex voice scenarios, which lack the ability to adapt to emotional labels, resulting in inefficiency and information omissions.

Method used

The voice clip collection is generated through time domain segmentation technology, tone, speech speed and volume characteristics are extracted, emotional labels are determined using deep neural networks, and corresponding templates are selected according to the speech type to generate text, realizing automatic transcript generation.

Benefits of technology

It realizes automated processing from original voice to formatted records, generates standardized and informative records, improving interrogation efficiency and credibility of records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375809A_ABST
    Figure CN120375809A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic record generation method based on AI voice interaction, and belongs to the technical field of voice analysis, and the method comprises the steps: obtaining original voice data, recording a continuous audio stream from a target environment through a preset audio collection module, and generating a preliminarily divided voice segment set through a time domain segmentation technology; extracting a tone feature, a speech speed change and a volume feature for the speech segment set, and generating a speech feature set; inputting the voice feature set into a deep neural network model, determining an emotion label through multi-dimensional feature mapping, and generating a voice unit set with the emotion label; if the voice type of the voice unit set is a statement type, generating a text by adopting a narrative template; if the voice type of the voice unit set is an inquiry type, adopting an inquiry type template to generate a text; and obtaining a formatted record. According to the AI voice interaction-based record automatic generation method, the problem that a record with a standard format and rich information is difficult to generate in an existing record generation mode is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech analysis, and particularly to a method for automatically generating transcripts based on AI voice interaction. Background Art

[0002] The generation of interrogation transcripts is a crucial link in the judicial field, directly related to the efficiency of case investigation and the credibility of evidence. However, current reliance on manual recording or other semi-automated solutions generally has significant limitations. Manual recording is time-consuming and inefficient, and it is easy to miss key information or misjudge the type of speech due to the subjective judgment of the recorder.

[0003] In the field of interrogation, the core challenges mainly focus on the accuracy of speech type recognition and the intelligent capture of emotional information. In the interrogation scenario, the speech is often accompanied by background noise, and the emotional information contained in the characteristics of the speech such as pitch, speech rate, and volume changes frequently. Although some existing automated tools can achieve the basic function of speech-to-text conversion, they often lack the adaptability to complex speech scenarios, especially in terms of emotional tags, resulting in difficulty in generating transcripts with standardized formats and rich information, which limits their application value in actual interrogations. Summary of the Invention

[0004] In order to overcome the defects existing in the prior art, the present invention provides a method for automatically generating transcripts based on AI voice interaction to solve the above problems.

[0005] The technical solution adopted by the present invention to solve its technical problems is: a method for automatically generating transcripts based on AI voice interaction, including the following steps: S1: Obtain the original speech data, record a continuous audio stream from the target environment through a preset audio acquisition module, and generate a set of preliminarily divided speech segments by using time-domain segmentation technology; S2: Extract the pitch feature, speech rate change, and volume feature for the set of speech segments to generate a speech feature set; S3: Input the speech feature set into a deep neural network model, determine the emotional label through multi-dimensional feature mapping, and generate a set of speech units with emotional labels; S4: If the speech type of the set of speech units is a statement type, generate text using a narrative template; if the speech type of the set of speech units is an interrogation type, generate text using an interrogation template; to obtain a formatted transcript.

[0006] Preferably, in the step S1, the original speech data is time-domain separated according to a preset silence threshold to generate a set of speech segments with time stamps.

[0007] Optionally, in step S1, obtain the duration of each segment in the speech segment set; if the duration exceeds a preset threshold, use an endpoint detection algorithm based on short-time energy to re-locate the segmentation point for this segment to optimize the speech segment set.

[0008] Specifically, in step S1, extract the Mel-frequency cepstral coefficients of each segment from the optimized speech segment set; Use a pre-trained classifier to classify the Mel-frequency cepstral coefficients to obtain the class label of each segment; Extract the segments with the class label of speech, and merge the continuous segments with the speech class label to further optimize the speech segment set.

[0009] Specifically, in step S2, perform short-time Fourier transform processing on the speech segments in the speech segment set to obtain the time-frequency matrix corresponding to each frame in the speech segment, extract the fundamental frequency trajectory from the time-frequency matrix as the pitch feature, calculate the average energy of each frame as the volume feature, and obtain the speech rate change feature through syllable boundary detection.

[0010] It should be noted that in step S2, combine the pitch feature, volume feature, and speech rate change feature into a three-dimensional vector, input it into a Gaussian mixture model to obtain a clustering label, and form a speech feature set after binding the pitch feature, volume feature, and speech rate change feature to the corresponding clustering label.

[0011] Optionally, in step S3, analyze the speech feature set through a deep neural network to obtain a speech feature vector; According to the speech feature vector, use principal component analysis technology to extract the change type to obtain preliminary emotion change data; Input the preliminary emotion change data into a pre-trained convolutional neural network to determine the emotion label of each frame in the speech segment, then re-combine it into a speech segment with an emotion label, and then combine all speech segments into a speech unit set.

[0012] Specifically, in step S4, obtain the speech segments in the speech unit set, and combine the pitch feature, volume feature, and speech rate change feature of the speech segment to judge whether the speech segment belongs to a statement type or an inquiry type to obtain a type judgment result; If the type judgment result is a statement type, convert the speech segment into text and fill it into a narrative template, and attach the corresponding emotion label; if the type judgment result is an inquiry type, convert the speech segment into text and fill it into an inquiry template, and attach the corresponding emotion label; obtain a formatted record.

[0013] The beneficial effects of the present invention are as follows: In the method for automatically generating a transcript based on AI voice interaction, a set of speech segments is initially divided from the recorded continuous audio stream through time-domain segmentation technology. Then, a set of speech features is extracted from the set of speech segments. Subsequently, the set of speech features is input into a deep neural network model to determine emotion labels, and a set of speech units with emotion labels is generated. Based on this set of speech units, it is determined to select corresponding templates according to different speech types to generate a formatted transcript. This solution realizes the automated processing from the original speech to the formatted transcript, and generates a transcript with standardized format and rich information, ensuring its application value in actual interrogations. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a flowchart of the method for automatically generating a transcript based on AI voice interaction in an embodiment of the present invention; Figure 2 It is a flowchart of the generation of the set of speech features in an embodiment of the present invention; Figure 3 It is a flowchart of the generation of the set of speech units in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The following further describes the specific embodiments of the present invention with reference to the drawings. It should be noted here that the description of these embodiments is for helping to understand the present invention, but does not constitute a limitation to the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0016] As Figures 1-3 shown, a method for automatically generating a transcript based on AI voice interaction includes the following steps: S1: Obtain the original speech data, record a continuous audio stream from the target environment through a preset audio acquisition module, and generate a set of initially divided speech segments by using time-domain segmentation technology; S2: For the set of speech segments, extract pitch features, speech rate changes, and volume features to generate a set of speech features; S3: Input the set of speech features into a deep neural network model, determine emotion labels through multi-dimensional feature mapping, and generate a set of speech units with emotion labels; S4: If the speech type of the set of speech units is a statement type, use a narrative template to generate text; if the speech type of the set of speech units is an interrogation type, use an interrogation template to generate text; obtain a formatted transcript.

[0017] In the above-described method for automatically generating a transcript based on AI voice interaction, a set of speech segments is initially divided from the recorded continuous audio stream through time-domain segmentation technology. Then, a set of speech features is extracted from the set of speech segments. Subsequently, the set of speech features is input into a deep neural network model to determine emotion labels and generate a set of speech units with emotion labels. Based on this set of speech units, it is determined to select corresponding templates according to different speech types to generate a formatted transcript. This solution realizes the automated processing from the original speech to the formatted transcript, and generates a transcript with a standardized format and rich information, ensuring its application value in actual interrogations.

[0018] It should be noted that in step S1, the original speech data is subjected to time-domain separation according to a preset silence threshold to generate a set of speech segments with timestamps.

[0019] In this embodiment, the target environment audio stream is collected through a microphone array to obtain the original audio data with a sampling rate of a preset value. For example, in an interrogation room, the microphone array can be deployed in the center of the room to collect the surrounding speech signals. Assuming the sampling rate is preset to 16 kHz to ensure that the audio quality is suitable for subsequent processing. This sampling rate is relatively common in speech processing and can balance clarity and data volume. A silence detection module is used to perform time-domain segmentation on the original audio data, and a set of speech segments with timestamps is generated according to a preset silence threshold. For example, the silence threshold of -40 dB can be set for the silence detection module, and the audio below this value is considered as a silent area; when someone speaks during the interrogation, the speech part is retained, while the part below the silence threshold is considered as the gap when no one is speaking and will be marked as silent. Based on this, multiple segments with timestamps are segmented and generated. The advantage of this is to initially separate the effective speech and reduce the subsequent computational burden.

[0020] Preferably, in step S1, the duration of each segment in the set of speech segments is obtained; if the duration exceeds a preset threshold, an endpoint detection algorithm based on short-time energy is used to re-locate the segmentation point for this segment to optimize the set of speech segments.

[0021] For a set of speech segments, calculate the duration of each segment. In one possible implementation, assume that the duration of a certain segment is 4 seconds and the preset threshold is 3 seconds. Segments exceeding the preset threshold need to be further processed. For example, during an interrogation, if someone makes a long speech, the segment may contain a complete sentence or even multiple sentences, with a relatively long duration, and requires more refined segmentation to improve the subsequent classification accuracy. If the duration exceeds the preset threshold, an endpoint detection algorithm based on short-time energy is used to re-locate the segmentation points to obtain an optimized set of speech segments. The endpoint detection algorithm based on short-time energy can detect the start and end points of a speech segment. For example, in a 4-second segment, it is determined through energy peaks that the actual speech in the segment starts at 0.5 seconds and ends at 3.5 seconds. By removing the silent parts at the beginning and end, the optimized segment obtained is more in line with the actual speech range. This method can effectively improve the purity of the speech segments and lay a foundation for feature extraction.

[0022] Optionally, in step S1, extract the Mel-frequency cepstral coefficients of each segment from the optimized set of speech segments; it can be understood that the Mel-frequency cepstral coefficients simulate the human ear's perception of frequency and are commonly used in speech recognition; Use a pre-trained classifier to classify the Mel-frequency cepstral coefficients to obtain the class label of each segment; use a pre-trained classifier to classify the Mel-frequency cepstral coefficients to obtain the class label of each segment. In one embodiment, the classifier can be a support vector machine, which can distinguish labels such as "speech", "noise", and "background sound" after pre-training. For example, in an interrogation, a certain audio segment is marked as "speech", and another segment containing the sound of a fan is marked as "noise". This classification helps to understand the audio content and enhance the semantic value of the processing results; Extract the segments with the class label of speech and merge the consecutive segments with the speech class label to further optimize the set of speech segments. For example, during an interrogation, someone says three consecutive sentences. These three sentences correspond to a single segment or three consecutive segments that are all marked as "speech", then these three segments can be merged into a complete speech segment and marked as "speech". This merging can reduce fragmentation, improve the coherence and practicality of the results, and facilitate archiving or subsequent speech-to-text conversion. In one possible implementation, the entire process from acquisition to classification and then to merging forms an efficient audio processing chain; for example, for audio in a noisy environment in an interrogation room, through the above processing, the speech content and noise can be clearly separated.

[0023] Specifically, in step S2, perform short-time Fourier transform processing on the speech segments in the set of speech segments to obtain the time-frequency matrix corresponding to each frame in the speech segment, extract the fundamental frequency trajectory from the time-frequency matrix as the pitch feature, calculate the average energy of each frame as the volume feature, and obtain the speech rate change feature through syllable boundary detection.

[0024] In this embodiment, a Hanning window is used to perform short-time Fourier transform processing on the speech segment. It can be understood that the Hanning window is a smoothing window function that can effectively reduce spectral leakage and improve the accuracy of time-frequency analysis. For example, in an interrogation recording, a 2-second speech segment is divided into multiple 20-millisecond frames. After applying the Hanning window to each frame, short-time Fourier transform is performed to generate a time-frequency matrix.

[0025] Specifically, the time-frequency matrix reflects the energy distribution of the speech segment in terms of time and frequency. For example, if a certain frame has strong energy at 300 Hz, it is considered the fundamental frequency of the speaker. The fundamental frequency trajectory is extracted from the time-frequency matrix as the pitch feature to track the change in the main frequency of each frame. For example, during an interrogation, if someone's pitch rises from 250 Hz to 350 Hz, the fundamental frequency trajectory can reflect this fluctuation. Calculating the average energy of each frame as the volume feature is relatively intuitive. In one possible implementation, the average energy of a certain frame of speech may be -20 dB, indicating a moderate volume. When obtaining the speech rate change feature through syllable boundary detection, the recognition of syllable boundaries depends on the intensity change in the speech. Specifically, a preset intensity threshold is used to determine the syllable boundary positions in the speech segment, the time intervals are extracted from the syllable boundary positions, and the time difference of the time intervals between adjacent boundaries is calculated to obtain the rhythm difference. For example, during an interrogation, someone says 4 syllables per second, and another segment drops to 2 syllables due to more pauses. The speech rate change feature can capture this rhythm difference.

[0026] It should be noted that in step S2, the pitch feature, volume feature, and speech rate change feature are combined into a three-dimensional vector and input into a Gaussian mixture model to obtain a clustering label. After binding the pitch feature, volume feature, and speech rate change feature to the corresponding clustering labels, a speech feature set is formed. After combining the pitch feature, volume feature, and speech rate change feature into a three-dimensional vector and inputting it into the Gaussian mixture model for clustering, for example, the three-dimensional vector of a certain segment of speech may be [300 Hz, -15 dB, 3 syllables per second]. After inputting it into the Gaussian mixture model for clustering, it is obtained that 300 Hz represents "steady-emotion speech", -15 dB represents "moderate-volume speech", and 3 syllables per second represents "more pauses".

[0027] Preferably, in step S3, the speech feature set is analyzed through a deep neural network to obtain a speech feature vector. In one possible implementation, when analyzing the speech features corresponding to each speech segment in the speech feature set through a deep neural network, a recurrent neural network can be used to process the temporal information to generate a speech feature vector. For example, after the speech features corresponding to a 2-second speech segment are input into the deep neural network, a 128-dimensional vector is output, which characterizes features such as pitch, pitch height, and speech rate. Specifically, the deep neural network may focus on the change trend of the pitch from low to high in a certain segment of speech, and the value of a certain dimension in the generated vector rises from 0.3 to 0.7, reflecting the dynamic characteristics. According to the speech feature vector, the principal component analysis technique is used to extract the change type to obtain preliminary emotional change data. When using the principal component analysis technique to extract the change type, the principal component analysis can reduce the high-dimensional features to the main change directions. For example, the first three principal components are extracted from a 128-dimensional vector, which represent the pitch fluctuation, the speech rate, and the volume strength respectively, to obtain the preliminary emotional change data. The preliminary emotional change data is input into a pre-trained convolutional neural network to determine the emotional label of each frame in the speech segment, and then recombined into a speech segment with emotional labels. Then, all speech segments are combined into a speech unit set. When the preliminary emotional change data is input into the convolutional neural network, the network captures local features through the convolutional layer to determine the emotional label. For example, a certain frame of a certain speech segment is labeled as "excited emotion" because the convolutional kernel recognizes the pattern of rising pitch and accelerating speech rate. In one embodiment, a 2-second speech segment is divided into 5 frames, which are respectively labeled as "calm", "excited", "excited", "calm", "excited", and these 5 frames with emotional labels are recombined into a speech segment, thereby generating a speech segment with emotional labels.

[0028] Optionally, in the step S4, the speech segments in the speech unit set are obtained, and the speech segments are judged to be of a statement type or an inquiry type by combining the pitch feature, the volume feature, and the speech rate change feature of the speech segments, and a type judgment result is obtained. If the type judgment result is a statement type, the speech segment is transformed into text and filled into a narrative template, and the corresponding emotional label is attached; if the type judgment result is an inquiry type, the speech segment is transformed into text and filled into an inquiry template, and the corresponding emotional label is attached; a formatted record is obtained.

[0029] It can be understood that the statement type usually focuses on transmitting information, with a stable intonation and logical coherence, while the inquiry type has an inquiring intention, the speech rate may be slightly faster and is often accompanied by interrogative words.

[0030] For example, in an interrogation recording, when someone is making a self-statement, the speech rate is stable at 3 syllables per second and there is no obvious pitch fluctuation, which can be preliminarily judged as a statement type. On the contrary, when someone is asking a question, the speech rate reaches 4 syllables per second and the pitch at the end rises, which tends to be an inquiry type. Specifically, keywords can also be combined during the judgment. For example, "what" and "how" often point to inquiries, while "already" and "therefore" often imply statements. It should be noted that transforming the speech segment into text information is an existing conventional technology, and it will not be elaborated here one by one. In one possible implementation manner, the narrative template is usually designed in the structure of "Answer:", and the inquiry template is usually designed in the structure of "Question:".

[0031] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments still fall within the protection scope of the present invention.

Claims

1. A method for automatically generating a transcript based on AI voice interaction, characterized in that, It includes the following steps: S1: Obtain the original voice data, record a continuous audio stream from the target environment through a preset audio acquisition module, and generate a preliminarily divided set of voice segments using time-domain segmentation technology; S2: For the set of voice segments, extract pitch features, speech rate changes, and volume features to generate a voice feature set; S3: Input the voice feature set into a deep neural network model, determine the emotion label through multi-dimensional feature mapping, and generate a set of voice units with emotion labels; S4: If the speech type of the set of voice units is a statement type, generate text using a narrative template; if the speech type of the set of voice units is an inquiry type, generate text using an inquiry template; Obtain a formatted record.

2. The method for automatically generating a transcript based on AI voice interaction according to claim 1, wherein: In the step S1, perform time-domain separation on the original voice data according to a preset silence threshold to generate a set of voice segments with timestamps.

3. A method for automatically generating a transcript based on AI voice interaction according to claim 2, characterized in that: In the step S1, obtain the duration of each segment in the set of voice segments; If the duration exceeds the preset threshold, use an endpoint detection algorithm based on short-time energy to re-locate the segmentation point for this segment to optimize the set of voice segments.

4. A method for automatically generating a transcript based on AI voice interaction according to claim 3, characterized in that: In the step S1, extract the Mel-frequency cepstral coefficients of each segment from the optimized set of voice segments; Use a pre-trained classifier to classify the Mel-frequency cepstral coefficients to obtain the class label of each segment; Extract the segments with the class label of speech, and merge the continuous segments with the speech class label to further optimize the set of voice segments.

5. The method for automatically generating a transcript based on AI voice interaction according to claim 4, characterized in that: In the step S2, perform short-time Fourier transform processing on the voice segments in the set of voice segments to obtain the time-frequency matrix corresponding to each frame in the voice segment, extract the fundamental frequency trajectory from the time-frequency matrix as the pitch feature, calculate the average energy of each frame as the volume feature, and obtain the speech rate change feature through syllable boundary detection.

6. The method for automatically generating a transcript based on AI voice interaction according to claim 5, wherein: In the step S2, combine the pitch feature, volume feature, and speech rate change feature into a three-dimensional vector, input it into a Gaussian mixture model to obtain a clustering label, and form a voice feature set after binding the pitch feature, volume feature, and speech rate change feature to the corresponding clustering label.

7. The method for automatically generating a transcript based on AI voice interaction according to claim 6, wherein: In the step S3, analyze the voice feature set through a deep neural network to obtain a voice feature vector; According to the voice feature vector, use principal component analysis technology to extract the change type to obtain preliminary emotion change data; Input the preliminary emotion change data into a pre-trained convolutional neural network to determine the emotion label of each frame in the voice segment, then re-combine it into a voice segment with emotion labels, and then combine all voice segments into a set of voice units.

8. A method for automatically generating a transcript based on AI voice interaction according to claim 7, characterized in that: In the step S4, obtain the voice segments in the set of voice units, and combine the pitch feature, volume feature, and speech rate change feature of the voice segment to judge whether the voice segment belongs to the statement type or the inquiry type to obtain a type judgment result; If the type judgment result is the statement type, transform the voice segment into text and fill it into the narrative template, and attach the corresponding emotion label; if the type judgment result is the inquiry type, transform the voice segment into text and fill it into the inquiry template, and attach the corresponding emotion label; obtain a formatted record.

Citation Information

Patent Citations

  • Meeting minute generating method and device, computer equipment and storage medium

    CN109817245A

  • Conference recording method and device, computer equipment and medium

    CN113691382A

  • Case interrogation intelligent auxiliary device and method

    CN119357599A

  • Medical information interaction terminal and interaction method based on AI analysis

    CN119560139A

  • System for Indicating Emotional Attitudes Through Intonation Analysis and Methods Thereof

    US20080270123A1