Voiceprint identification report generation method, device and computer-readable medium

By acquiring speech and frequency spectrum information and using natural language processing models to automatically generate voiceprint identification reports, the problems of long time and unstable quality caused by manual writing in existing technologies are solved, and efficient and automated report generation is achieved.

CN116417002BActive Publication Date: 2025-09-16VOICEAI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310267662.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2025-09-16
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

Existing voiceprint identification reports rely on manual writing, which takes a long time and has uneven quality.

Method used

By obtaining the spectral and frequency spectrum information of multiple specimens and sample characteristic sound segments, descriptive information is generated, and statistical analysis is performed using the natural language processing model. Finally, a voiceprint identification report is automatically generated based on the format template.

Benefits of technology

It realizes the automatic generation of voiceprint identification reports, saves manual writing time and manpower, and improves the quality and efficiency of reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116417002B_ABST
    Figure CN116417002B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, and computer-readable medium for generating a voiceprint identification report, relating to the field of voiceprint identification technology. The method comprises: obtaining multiple atlas information based on multiple pre-acquired sample characteristic sound segments and multiple sample characteristic sound segments; generating descriptive information for each of the atlas information; performing statistical analysis on the multiple descriptive information based on a natural language processing model to generate summary information; and generating a voiceprint identification report based on the atlas information, the descriptive information, and the summary information based on a preset format template. Therefore, the method can automatically generate a voiceprint identification report based on the sample voice and sample voice, greatly saving the manpower and time required to manually write the identification report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of voiceprint identification, and more specifically, to a method, device, and computer-readable medium for generating a voiceprint identification report. Background Art

[0002] Currently, in the voiceprint identification business, voiceprint identification reports are generally written manually and rely heavily on the identification personnel's professional knowledge and practical experience, which makes it take a long time to issue identification reports in actual business and the quality is uneven. Summary of the Invention

[0003] This application proposes a method, device and computer-readable medium for generating a voiceprint identification report to improve the above-mentioned defects.

[0004] In the first aspect, an embodiment of the present application provides a method for generating a voiceprint identification report, the method comprising: obtaining multiple atlas information based on multiple pre-acquired specimen characteristic sound segments and multiple sample characteristic sound segments, the atlas information comprising spectral information and frequency spectrum information, for characterizing the voiceprint characteristics of the specimen characteristic sound segments and the sample characteristic sound segments; generating descriptive information for each of the atlas information, the descriptive information being used to describe the comparison results of the specimen characteristic sound segments and the sample characteristic sound segments; performing statistical analysis on the multiple descriptive information based on a natural language processing model to generate summary information, the summary information being used to evaluate the identity of the specimen voice and the sample voice; and generating a voiceprint identification report based on the atlas information, the descriptive information and the summary information based on a preset format template.

[0005] In the second aspect, the embodiment of the present application also provides a voiceprint recognition device, which includes: a graph information generation unit, a description information generation unit, a summary information generation unit, and an identification report generation unit. The graph information generation unit is used to obtain multiple graph information based on multiple pre-acquired sample characteristic sound segments and multiple sample characteristic sound segments, and the graph information includes spectral information and spectrum information, which is used to characterize the voiceprint characteristics of the sample characteristic sound segments and the sample characteristic sound segments; the description information generation unit is used to generate description information for each of the graph information, and the description information is used to describe the comparison results of the sample characteristic sound segments and the sample characteristic sound segments; the summary information generation unit is used to perform statistical analysis on the multiple description information based on the natural language processing model to generate summary information, and the summary information is used to evaluate the identity of the sample voice and the sample voice; the identification report generation unit is used to generate a voiceprint identification report based on the graph information, the description information and the summary information based on a preset format template.

[0006] In a third aspect, an embodiment of the present application further provides a computer-readable medium, wherein the computer-readable medium stores a program code executable by a processor, and when the program code is executed by the processor, the processor executes the above method.

[0007] The present application provides a method, device and computer-readable medium for generating a voiceprint identification report. The method includes: based on a plurality of pre-acquired specimen characteristic sound segments and a plurality of sample characteristic sound segments, obtaining a plurality of atlas information, the atlas information including spectral information and frequency spectrum information, for characterizing the voiceprint characteristics of the specimen characteristic sound segments and the sample characteristic sound segments; then, generating description information for each of the atlas information, the description information being used to describe the comparison results of the specimen characteristic sound segments and the sample characteristic sound segments; then, based on a natural language processing model, performing statistical analysis on the plurality of the description information to generate summary information, the summary information being used to evaluate the identity of the specimen voice and the sample voice; finally, based on a preset format template, generating a voiceprint identification report according to the atlas information, the description information and the summary information. Therefore, this method can automatically obtain the graph information representing the voiceprint features of the specimen and sample characteristic sound segments based on the pre-acquired specimen characteristic sound segments and sample characteristic sound segments, the descriptive information describing the comparison results of the specimen and sample characteristic sound segments, and the summary information evaluating the identity of the specimen and sample voices. Finally, based on a certain format template, the graph information, the descriptive information and the summary information are sorted to obtain a voiceprint identification report. In other words, this method can automatically generate a voiceprint identification report based on the specimen voice and sample voice, which greatly saves the manpower and time of manually writing the identification report.

[0008] Other features and advantages of the embodiments of the present application will be described in the following description and, in part, will become apparent from the description or be understood by practicing the embodiments of the present application. The objectives and other advantages of the embodiments of the present application can be achieved and obtained through the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 A flow chart of a method for generating a voiceprint identification report provided in an embodiment of the present application is shown.

[0011] Figure 2 A flow chart of a method for generating a voiceprint identification report provided in another embodiment of the present application is shown.

[0012] Figure 3 A flow chart of the method for training and obtaining voiceprint feature portraits provided in this application is shown.

[0013] Figure 4 A flow chart of the method provided in this application for estimating a second voiceprint feature from a voiceprint feature portrait is shown.

[0014] Figure 5 A flow chart of a method for generating a voiceprint identification report provided in another embodiment of the present application is shown.

[0015] Figure 6 A flow chart of the method for updating voiceprint feature portraits provided by this application is shown.

[0016] Figure 7 A unit block diagram of a voiceprint recognition device provided in an embodiment of the present application is shown.

[0017] Figure 8 A schematic diagram of a wearable device provided in one embodiment of the present application is shown.

[0018] Figure 9 A schematic diagram of a storage unit according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.

[0020] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0021] Voiceprint identification, also known as voice identity identification, refers to the scientific judgment of the identity of the voice recorded in audio and video materials through comparison and analysis. It is widely used in judicial practice to determine and confirm the identity information of the recording by examining the recording materials in the case.

[0022] In court cases involving voice identity verification, the written materials submitted to the court primarily consist of two documents: the voice identity verification record and the verification document. These documents contain a large amount of graphical material, as well as text explaining and demonstrating these graphical materials, and the conclusion regarding the voice identity verification. The graphical materials can be obtained from a voiceprint verification workstation, which processes the voice of the test subject and the sample voice, extracting them and converting them into images and tables. However, the text in the verification document must be manually compiled and analyzed based on the verification record, requiring considerable expertise and practical experience. The majority of voiceprint verification reports used in actual court cases are produced by industry experts.

[0023] Therefore, in order to overcome the above-mentioned defects, the embodiments of the present application provide a method, device and computer-readable medium for generating a voiceprint identification report, based on a plurality of pre-acquired specimen characteristic sound segments and a plurality of sample characteristic sound segments, a plurality of atlas information is obtained, the atlas information includes spectral information and spectrum information, which is used to characterize the voiceprint characteristics of the specimen characteristic sound segments and the sample characteristic sound segments, and then, description information of each of the atlas information is generated, the description information is used to describe the comparison results of the specimen characteristic sound segments and the sample characteristic sound segments, and then, based on the natural language processing model, a plurality of the description information is statistically analyzed to generate summary information, the summary information is used to evaluate the identity of the specimen voice and the sample voice, and finally, based on a preset format template, a voiceprint identification report is generated according to the atlas information, the description information and the summary information. The voiceprint identification report generation method provided in the embodiments of the present application greatly saves the manpower and time of manually writing identification reports by automatically generating voiceprint identification reports based on specimen voice and sample voice, and provides a convenient and efficient means to combat the growing number of electronic fraud cases.

[0024] See also Figure 1 , Figure 1 A method for generating a voiceprint identification report provided in an embodiment of the present application is shown. Specifically, the method includes: S101 to S104.

[0025] S101: Based on a plurality of sample characteristic sound segments and a plurality of sample characteristic sound segments acquired in advance, a plurality of atlas information is acquired, wherein the atlas information includes spectral information and frequency spectrum information, and is used to characterize the voiceprint features of the sample characteristic sound segments and the sample characteristic sound segments.

[0026] Inspection material and sample: For two audios that need to be identified for voice identity, the audio whose identity information is already clear is called the sample, and the audio whose identity is pending and needs to be clarified is called the inspection material.

[0027] As an embodiment, the atlas information includes spectral information and frequency spectrum information. Specifically, the spectral information may include a sample speech spectrogram and a sample speech spectrogram. The frequency spectrum information may include a spectrogram and a data table obtained after measuring the spectrogram, wherein the horizontal axis of the spectrogram is time, the vertical axis is frequency, and the coordinate points reflect the energy of the speech signal; the horizontal axis of the spectrogram is frequency, the vertical axis is amplitude, and the coordinate points reflect the relationship between the frequency and energy of the speech signal; the data table includes feature parameter data obtained after measuring the spectrogram. Furthermore, the spectrogram includes the formant frequency data of all frames of complete syllables in the sample speech and the sample speech. The spectrogram is a linear predictive coding (LPC) spectrum curve of the cursor selection frame of the sample feature segment and the sample feature segment. The feature parameters may be the formant center frequency, formant intensity, and formant bandwidth.

[0028] As an implementation method, the atlas information can be obtained based on the characteristic sound segments of the specimen and the characteristic sound segments of the sample through an atlas algorithm and converted into a picture or table for storage. Specifically, the atlas algorithm can be a Fourier transform algorithm.

[0029] S102: Generate description information for each of the atlas information, where the description information is used to describe the comparison result between the sample characteristic sound segment and the sample characteristic sound segment.

[0030] In one embodiment, the description information of each of the atlas information is generated, that is, based on the sample speech spectrogram, the sample speech spectrogram, the spectrogram, and the data table, after performing an acoustic feature comparison between the sample characteristic sound segment and the sample characteristic sound segment, a corresponding comparison result analysis text is generated. Specifically, the description information may include description information of the formant trajectory trend comparison result and description information of the formant parameter comparison result.

[0031] As an embodiment, the method for generating the descriptive information can be based on the comparison results and according to pre-set rules. Specifically, by pre-setting a conditional threshold, when the comparison result reaches the conditional threshold, descriptive information related to the condition is added. The descriptive information added depends on the threshold condition that the comparison result meets. For example, when the comparison result of the resonance peak trajectory trend is obtained, the comparison result is compared with the set resonance peak trajectory trend conditional threshold to determine whether the resonance peak trajectory trends are basically the same or basically different. If the resonance peak trajectory trends are basically the same, the comparison result of the resonance peak trajectory trend is further compared with the set resonance peak trajectory trend similarity conditional threshold to determine whether the resonance peak trajectory trends meet the similarity level of completely consistent, basically consistent, or relatively close. Finally, based on the above conclusion, descriptive information of the resonance peak trajectory trend is generated.

[0032] As an embodiment, the method for generating the descriptive information may be to automatically generate the descriptive information by analyzing and processing the graph information through a pre-acquired descriptive information generation model, wherein the descriptive information generation model is based on a large number of appraisal reports in a historical case library that have been manually reviewed and approved by appraisal experts as training data, and is obtained through deep learning algorithm training.

[0033] S103: Based on a natural language processing model, statistical analysis is performed on the plurality of descriptive information to generate summary information, where the summary information is used to evaluate the identity of the sample voice and the sample voice.

[0034] Natural Language Processing (NLP) is a method for achieving effective communication between humans and computers using natural language. Specifically, NLP conducts quantitative research on language information with the support of computers, and provides language descriptions that can be used by humans and computers, so that computers can not only understand the meaning of natural language texts, but also express given intentions and ideas in natural language texts.

[0035] As an implementation method, the summary information is generated based on the description information, that is, based on the description of the comparison results of the characteristic sound segment of the specimen and the characteristic sound segment of the sample, to analyze and judge the voice identity, that is, whether the identity information of the specimen voice and the identity information of the sample voice are the same person, and based on the identity conclusion, the above comparison results are further analyzed and evaluated, and the analysis and evaluation of the identity conclusion and the comparison results are integrated to generate summary information.

[0036] As an implementation method, the method for generating the summary information may be to use a natural speech processing model to understand the text information in the description information, analyze and process the text information, and then use a natural speech processing algorithm to output the analysis results as synthesized text as summary information.

[0037] S104: Based on a preset format template, generate a voiceprint identification report according to the graph information, the description information and the summary information.

[0038] Therefore, the voiceprint identification report generation method provided in the embodiment of the present application can automatically obtain the graph information representing the voiceprint characteristics of the specimen and sample characteristic sound segments based on the pre-acquired specimen characteristic sound segments and sample characteristic sound segments, descriptive information describing the comparison results of the specimen and sample characteristic sound segments, and summary information evaluating the identity of the specimen and sample voice. Finally, based on a certain format template, the graph information, the descriptive information and the summary information are sorted to obtain a voiceprint identification report. In other words, this method can automatically generate a voiceprint identification report based on the specimen voice and sample voice, greatly saving the manpower and time of manually writing the identification report.

[0039] As an implementation, see Figure 2 , Figure 2 The method for generating description information of each of the graph information in step S102 provided in an embodiment of the present application is shown. Specifically, the method includes: S201 to S203.

[0040] S201: Based on the spectral information, perform formant trajectory comparison on the characteristic sound segment of the sample and the characteristic sound segment of the sample to obtain trajectory comparison data.

[0041] As an embodiment, the trajectory comparison data can be obtained based on the relative magnitude of the frequencies of the start and end points of the formant trajectory, as well as the change in the slope of the formant trajectory curve. Specifically, based on the spectral information, i.e., the sample speech spectrogram and the sample speech spectrogram, the algorithm can automatically generate the formant trajectory, and the slope change of the trajectory curve is determined by calculating the first-order and second-order derivatives (differences) of the formant trajectory curve of the syllable stable segment, thereby determining the trajectory trend.

[0042] S202: Based on the frequency spectrum information, perform formant parameter comparison on the characteristic sound segment of the sample and the characteristic sound segment of the sample to obtain parameter comparison data.

[0043] In one embodiment, the parameter comparison data includes formant parameters calculated using a formant estimation algorithm or manual measurement based on the spectral information, i.e., the spectral curves of the sample characteristic sound segment and the sample characteristic sound segment, as well as deviation data between the sample formant parameters and the sample formant parameters. Specifically, the parameter comparison data may include formant number data, formant center frequency data, formant intensity data, and formant bandwidth data.

[0044] S203: Generate description information based on the trajectory comparison result and the parameter comparison result.

[0045] As an implementation, see Figure 3 , Figure 3 The method for generating description information from the trajectory comparison result and the parameter comparison result in step S203 provided by an embodiment of the present application is shown. Specifically, the method may include: S301 to S304.

[0046] S301: Based on the trajectory comparison data, determine whether the deviation between the specimen resonance peak trajectory trend and the sample resonance peak trajectory trend is greater than a preset threshold;

[0047] S302: If the deviation between the specimen resonance peak trajectory trend and the sample resonance peak trajectory trend is greater than a preset threshold, resonance peak trajectory trend description information is generated.

[0048] As an embodiment, based on the first and second-order derivative features, it is possible to determine whether the resonance peak trajectory trend is rising or falling, and whether it is convex or concave. Since a positive first-order derivative indicates an rising trend, and a negative first-order derivative indicates a falling trend; a positive second-order derivative indicates a convex shape, and a negative second-order derivative indicates a concave shape, then, for example, when the first-order derivative is positive and greater than a preset rising threshold, it indicates that the resonance peak trajectory trend is rising. The generated resonance peak trajectory trend description information is: the resonance peak is rising. Furthermore, a description threshold can be added. When the first-order derivative is positive and greater than the preset rising threshold, the first-order derivative is further compared with the description threshold to obtain different trend levels of the rising trend, such as a slight rise, a large rise, etc. As an embodiment, the generation of the descending, convex, and concave trajectory trend description information can refer to the above embodiment and will not be repeated here.

[0049] As an embodiment, the preset threshold is a threshold value set for the first and second order derivative characteristics of the smooth resonance peak trajectory curve of the syllable stable segment. Specifically, for example, if the difference between the first order derivative of the resonance peak of the test material and the first order derivative of the resonance peak of the test sample is greater than the preset threshold, it means that the rising trend of the resonance peak of the test material is significantly different from the rising trend of the resonance peak of the sample, and thus based on the difference conclusion, the resonance peak trajectory trend description information is generated. As an embodiment, the generation of the descending, convex and concave trajectory trend description information can refer to the aforementioned embodiment and will not be repeated here.

[0050] S303: If the deviation between the specimen resonance peak trajectory trend and the sample resonance peak trajectory trend is less than or equal to a preset threshold, generate resonance peak trajectory similarity description information based on a preset trajectory similarity threshold.

[0051] In one embodiment, if the deviation between the specimen's resonance peak trajectory and the sample's resonance peak trajectory is less than or equal to a preset threshold, it indicates that the specimen's resonance peak trajectory is relatively similar to the sample's resonance peak trajectory. Furthermore, the trajectory similarity threshold is used to determine the degree of similarity between the specimen's resonance peak trajectory and the sample's resonance peak trajectory. For example, based on the trajectory similarity threshold, different levels such as completely consistent, substantially consistent, and relatively close can be classified, and corresponding resonance peak trajectory similarity description information can be generated.

[0052] S304: Based on a preset parameter deviation threshold, generate resonance peak parameter similarity description information according to the parameter comparison data.

[0053] As an implementation method, for each resonance peak parameter, a different deviation interval is set within a reasonable parameter deviation threshold range. Ultimately, the conformance of the resonance peak parameters is classified into different levels, such as complete consistency, basic consistency, and relatively close, based on the actual parameter deviation. Furthermore, exceeding the reasonable deviation threshold range can be described as a large deviation. Specifically, for example, a situation where the deviations of all four resonance peaks are less than 15% of the maximum reasonable deviation is set as complete consistency, a situation where the deviation of one resonance peak is greater than 90% of the maximum reasonable deviation is set as relatively close, and other situations are set as basic consistency.

[0054] As an implementation method, for the case where the deviation directions of the high-order resonance peaks (F3, F4) are inconsistent, a description of the center frequency interval deviation is added, and a corresponding parameter deviation threshold is set.

[0055] Although the method of generating descriptive information based on the comparison results and pre-set rules is easier to implement at the program level, this method of generating descriptive information also has its weaknesses. The inventors found in use that the evaluation and text generation of the comparison results of the characteristic sound segments of the specimen and the characteristic sound segments of the sample based on the preset threshold value makes the entire evaluation result too dependent on the manually set threshold value, and it is impossible to make timely changes to the evaluation criteria according to the different circumstances of different cases, and the descriptive information is too rigid.

[0056] As an implementation, see Figure 4 , Figure 4 Another method for generating description information of each of the graph information in step S102 provided by an embodiment of the present application is shown. Specifically, the method may include: S401 to S402.

[0057] S401: Based on the voiceprint feature data and voiceprint feature description texts in multiple identification reports, a description information generation model is trained;

[0058] S402: Input the graph information into the description information generation model to generate description information.

[0059] As an implementation method, the identification report can be a complete voiceprint identification report manually issued from a large number of cases in a network server. Based on the voiceprint feature data and voiceprint feature description text in a large number of identification reports, a description information generation model is trained through a deep learning algorithm, and then the description information generation model is used to automatically generate description information for the graph information. The description information generation model continuously updates the evaluation criteria for a large amount of graph information during continuous learning iterations, making the generated description information more flexible and changeable, thereby adapting to different cases and improving the accuracy of subsequent identity judgments.

[0060] As an implementation, see Figure 5 , Figure 5 A method provided in an embodiment of the present application is shown, in step S103, for performing statistical analysis on a plurality of description information based on a natural language processing model to generate summary information. Specifically, the method includes: S501 to S503.

[0061] S501: Searching for description information that meets a preset standard among the plurality of description information as key description information, and counting the proportion of the conforming point information and the different point information in the key description information.

[0062] As an embodiment, the preset standard refers to whether the descriptive information plays a leading role in ultimately reaching a conclusion as to whether the specimen and the sample are from the same person. As an embodiment, in the description of resonance peak features, information such as the number of resonance peaks, resonance peak trends, resonance peak center frequencies, and frequency intervals of higher-order resonance peaks can be used as key descriptive information. Furthermore, the determination of key descriptive information can be quantitatively described based on relevant industry standards.

[0063] S502: Generate voiceprint identity summary information based on the natural language processing model and the proportion;

[0064] S503: Based on the natural language processing model, generate analysis summary information for the key description information that does not meet the voiceprint identity conclusion.

[0065] As an embodiment, the method of generating voiceprint identity summary information based on the proportion ratio can be to refer to a preset identity threshold. If the proportion of the matching points exceeds the identity threshold, it is judged that the specimen and the sample are from the same person. Otherwise, it is judged that they are not, and the voiceprint identity summary information is generated based on the judgment conclusion.

[0066] As an implementation method, the minor factors in the points of conformity and difference included in the key description information that are inconsistent with the voiceprint identity summary information are described in combination with the context of the speech and the recording environment to generate analysis summary information. For example, the low-order resonance peaks of phonemes that do not conform to the characteristics of the same person may be consistent, which is a common feature of the same pronunciation and is a non-essential conformity, or the resonance peak intensity and bandwidth of phonemes that conform to the characteristics of the same person may be different, which is caused by different recording equipment and environment and is a non-essential difference.

[0067] See also Figure 6 , Figure 6 A method for generating a voiceprint identification report provided in an embodiment of the present application is shown. Specifically, the method includes: S601 to S606.

[0068] S601: Based on the pre-acquired sample voice and sample voice, obtain sample text information and sample text information, wherein the sample text information includes sample phoneme information and sample word information, and the sample text information includes sample phoneme information and sample word information.

[0069] As an implementation method, the method for obtaining the specimen voice and sample voice is to first obtain the specimen recording and sample recording through recording collection and other means, and then calculate various quality information of the recording through a voice quality detection algorithm, such as effective duration, signal-to-noise ratio, reverberation time, clipping ratio, etc., and screen out the voice that meets specific quality conditions as the specimen voice and sample voice.

[0070] As an embodiment, the method for obtaining the text information is to use automatic speech recognition technology to convert the specimen voice and the sample voice into text information. Specifically, the text information includes phonemes (English phonetic symbols, Chinese pinyin) and word information (English words, Chinese characters). Furthermore, the word information also includes the timestamp of the corresponding audio data (the start and end time of the text segment in the audio data).

[0071] S602: Searching for sound segments that meet preset conditions in the sample voice and the sample voice as sample feature sound segments and sample feature sound segments, wherein the preset conditions are that the sample phoneme information is the same as the sample phoneme information, and the sample word information is the same as the sample word information.

[0072] As an embodiment, the method for obtaining the characteristic sound segments of the sample and the characteristic sound segments of the sample is to search for sound segments that meet the conditions in the sample voice and the sample voice based on preset conditions. The preset conditions are that the sample phoneme information is the same as the sample phoneme information, and the sample word information is the same as the sample word information, thereby realizing automatic selection of characteristic sound segments.

[0073] S603: Based on the pre-acquired plurality of sample characteristic sound segments and the pre-acquired plurality of sample characteristic sound segments, a plurality of atlas information is acquired, wherein the atlas information includes spectral information and frequency spectrum information, and is used to characterize the voiceprint features of the sample characteristic sound segments and the sample characteristic sound segments;

[0074] S604: generating description information for each of the atlas information, wherein the description information is used to describe the comparison result between the sample characteristic sound segment and the sample characteristic sound segment;

[0075] S605: Based on a natural language processing model, statistical analysis is performed on the plurality of description information to generate summary information, wherein the summary information is used to evaluate the identity of the sample voice and the sample voice;

[0076] S606: Based on a preset format template, generate a voiceprint identification report according to the graph information, the description information and the summary information.

[0077] The implementation of steps S603 to S606 may refer to the aforementioned embodiment and will not be described in detail here.

[0078] See Figure 7 , Figure 7 A method for generating a voiceprint identification report provided in an embodiment of the present application is shown. Specifically, the method includes: S701 to S705.

[0079] S701: Based on a plurality of pre-acquired characteristic sound segments of a specimen and a plurality of pre-acquired characteristic sound segments of a sample, a plurality of atlas information is acquired, wherein the atlas information includes spectral information and frequency spectrum information, and is used to characterize the voiceprint features of the characteristic sound segments of the specimen and the sample;

[0080] S702: generating description information for each of the atlas information, wherein the description information is used to describe the comparison result between the sample characteristic sound segment and the sample characteristic sound segment;

[0081] S703: Based on a natural language processing model, statistical analysis is performed on the plurality of description information to generate summary information, wherein the summary information is used to evaluate the identity of the sample voice and the sample voice;

[0082] S704: Based on a preset format template, generate a voiceprint identification report according to the graph information, the description information and the summary information.

[0083] The implementation of step S701 and step S704 may refer to the aforementioned embodiment and will not be described in detail here.

[0084] S705: Optimizing the natural language processing model based on the voiceprint feature description texts and identification conclusion texts in the multiple identification reports.

[0085] As an implementation method, the natural language processing model is optimized based on the voiceprint feature description texts and identification conclusion texts in multiple identification reports, wherein the identification reports can be complete voiceprint identification reports manually issued from a large number of cases in a network server. Through the learning algorithm, the natural language processing model is continuously updated and iterated to reduce the error rate of the final generated summary information, making the sentences more fluent and in line with the reading habits of natural people.

[0086] See also Figure 8 , which shows the structural frame of a voiceprint identification report generating device 800 provided in an embodiment of the present application. The device may include a graph information generating unit 801, a description information generating unit 802, a summary information generating unit 803 and an identification report generating unit 804.

[0087] A graph information generating unit 801 is configured to obtain a plurality of graph information based on a plurality of sample characteristic sound segments and a plurality of sample characteristic sound segments acquired in advance, wherein the graph information includes spectral information and frequency spectrum information, and is configured to represent the voiceprint features of the sample characteristic sound segments and the sample characteristic sound segments;

[0088] A description information generating unit 802 is configured to generate description information for each of the atlas information, wherein the description information is used to describe the comparison result between the sample characteristic sound segment and the sample characteristic sound segment;

[0089] A summary information generating unit 803 is configured to perform statistical analysis on the plurality of description information based on a natural language processing model to generate summary information, wherein the summary information is used to evaluate the identity of the sample voice and the sample voice;

[0090] The identification report generating unit 804 is configured to generate a voiceprint identification report based on a preset format template, the atlas information, the description information and the summary information.

[0091] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0092] In several embodiments provided in this application, the coupling between units may be electrical, mechanical or other forms of coupling.

[0093] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0094] Please refer to Figure 9 , which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable medium 900 stores program code, which can be called by a processor to execute the method described in the above method embodiment.

[0095] The computer-readable storage medium 900 may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 900 may include a non-transitory computer-readable storage medium. The computer-readable storage medium 900 may have storage space for program code 910 for executing the steps of any of the above methods. These program codes may be read from or written to one or more computer program products. The program code 910 may be compressed, for example, in a suitable form.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating a voiceprint identification report, characterized in that: include: Based on the pre-acquired plurality of sample characteristic sound segments and the pre-acquired plurality of sample characteristic sound segments, a plurality of atlas information is acquired, wherein the atlas information includes spectral information and frequency spectrum information, and is used to characterize the voiceprint features of the sample characteristic sound segments and the sample characteristic sound segments; Generating description information for each of the atlas information, wherein the description information is used to describe the comparison result between the characteristic sound segment of the sample and the characteristic sound segment of the sample; Based on a natural language processing model, statistical analysis is performed on the plurality of descriptive information to generate summary information, wherein the summary information is used to evaluate the identity of the sample voice and the sample voice; Generate a voiceprint identification report based on the atlas information, the description information, and the summary information based on a preset format template; The generating of description information of each of the graph information includes: Based on the spectral information, performing formant trajectory comparison on the characteristic sound segment of the sample and the characteristic sound segment of the sample to obtain trajectory comparison data; Based on the spectrum information, performing formant parameter comparison on the sample characteristic sound segment and the sample characteristic sound segment to obtain parameter comparison data; generating description information from the trajectory comparison data and the parameter comparison data; The generating of description information from the trajectory comparison data and the parameter comparison data includes: Based on the trajectory comparison data, determine whether the deviation between the resonance peak trajectory trend of the specimen and the resonance peak trajectory trend of the sample is greater than a preset threshold; If the deviation between the resonance peak trajectory trend of the specimen and the resonance peak trajectory trend of the sample is greater than a preset threshold, generating resonance peak trajectory trend description information; If the degree of deviation between the specimen resonance peak trajectory trend and the sample resonance peak trajectory trend is less than or equal to a preset threshold, resonance peak trajectory similarity description information is generated based on a preset trajectory similarity threshold.

2. The method according to claim 1, characterized in that The generating of description information from the trajectory comparison data and the parameter comparison data further includes: Based on a preset parameter deviation threshold, resonance peak parameter similarity description information is generated according to the parameter comparison data.

3. The method according to claim 1, characterized in that The natural language processing model is used to perform statistical analysis on the plurality of description information to generate summary information, including: Searching for description information that meets a preset standard among the plurality of description information as key description information, and counting the proportion of the conforming point information and the different point information in the key description information; Based on the natural language processing model, generating voiceprint identity summary information according to the proportion; Based on the natural language processing model, analysis summary information is generated for the key description information that does not conform to the voiceprint identity conclusion.

4. The method according to claim 1, wherein Before acquiring a plurality of atlas information based on a plurality of sample characteristic sound segments and a plurality of sample characteristic sound segments acquired in advance, the method further includes: Based on the pre-acquired sample voice and sample voice, obtain sample text information and sample text information, wherein the sample text information includes sample phoneme information and sample word information, and the sample text information includes sample phoneme information and sample word information; Search for sound segments that meet preset conditions in the sample voice and the sample voice as sample feature sound segments and sample feature sound segments, wherein the preset conditions are that the sample phoneme information is the same as the sample phoneme information, and the sample word information is the same as the sample word information.

5. The method according to claim 1, wherein After generating a voiceprint identification report based on the preset format template and the atlas information, the description information, and the summary information, the method further includes: The natural language processing model is optimized based on the voiceprint feature description texts and identification conclusion texts in multiple identification reports.

6. A device for generating a voiceprint identification report, characterized in that: It includes a map information generating unit, a description information generating unit, a summary information generating unit, and an identification report generating unit; The atlas information generating unit is used to obtain a plurality of atlas information based on a plurality of sample characteristic sound segments and a plurality of sample characteristic sound segments acquired in advance, wherein the atlas information includes spectral information and frequency spectrum information, and is used to characterize the voiceprint features of the sample characteristic sound segments and the sample characteristic sound segments; The description information generating unit is used to generate description information of each of the atlas information, and the description information is used to describe the comparison result of the sample characteristic sound segment and the sample characteristic sound segment; the description information of each of the atlas information is generated, including: based on the spectral information, performing a resonance peak trajectory comparison on the sample characteristic sound segment and the sample characteristic sound segment to obtain trajectory comparison data; based on the spectrum information, performing a resonance peak parameter comparison on the sample characteristic sound segment and the sample characteristic sound segment to obtain parameter comparison data; generating a description from the trajectory comparison data and the parameter comparison data. Information; the generating of description information from the trajectory comparison data and the parameter comparison data, including: judging, based on the trajectory comparison data, whether the degree of deviation between the specimen resonance peak trajectory trend and the sample resonance peak trajectory trend is greater than a preset threshold; if the degree of deviation between the specimen resonance peak trajectory trend and the sample resonance peak trajectory trend is greater than the preset threshold, generating resonance peak trajectory trend description information; if the degree of deviation between the specimen resonance peak trajectory trend and the sample resonance peak trajectory trend is less than or equal to the preset threshold, generating resonance peak trajectory similarity description information based on a preset trajectory similarity threshold; The summary information generating unit is used to perform statistical analysis on the plurality of descriptive information based on a natural language processing model to generate summary information, wherein the summary information is used to evaluate the identity of the sample voice and the sample voice; The identification report generating unit is used to generate a voiceprint identification report based on a preset format template, according to the atlas information, the description information and the summary information.

7. A computer-readable medium, characterized in that The computer-readable medium stores a program code executable by a processor, and when the program code is executed by the processor, the processor executes the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voiceprint identification method, device and equipment

    CN110634490A

  • Voiceprint identification method and device, model training method and device, equipment and storage medium

    CN112382300A