Device and method for recognizing voiceprint and computer readable medium

By establishing a voiceprint database and using spatiotemporal tags to select the current context, the problems of insufficient voiceprint recognition speed and high false positive rate were solved, achieving fast and accurate voiceprint recognition.

CN122024741APending Publication Date: 2026-05-12SQ TECH (SHANGHAI) CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SQ TECH (SHANGHAI) CORP
Filing Date
2026-02-12
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing voiceprint recognition technologies suffer from computational latency and high false positive rates when identifying large-scale databases, especially in complex scenarios.

Method used

By establishing a voiceprint database, recording voiceprint models, spatiotemporal contexts, and correlations, the current context is selected using the spatiotemporal labels of speech signals, and voiceprints are identified from candidate voiceprint models, reducing the need for full database comparisons and improving recognition speed and accuracy.

Benefits of technology

It achieves fast and accurate voiceprint recognition in complex situations, reduces computational load and false positive rate, and improves recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024741A_ABST
    Figure CN122024741A_ABST
Patent Text Reader

Abstract

The invention discloses a device and a method for quickly recognizing voiceprints based on a space-time situation and a computer readable medium, and the method comprises the steps: judging a current situation according to a space-time label after the space-time label synchronized with a voice signal is obtained, and selecting a candidate voiceprint model associated with the current situation from a voiceprint database; according to the technical means of identifying the voiceprint in the voice signal according to the candidate voiceprint model, the number of voiceprint comparison models can be reduced, and the technical effect of improving the voiceprint identification speed and accuracy is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A voiceprint recognition device, method, and computer-readable medium, particularly a device, method, and computer-readable medium for rapid voiceprint recognition based on spatiotemporal context. Background Technology

[0002] With the popularization of artificial intelligence and biometric technology, voiceprint recognition has been widely used in fields such as identity verification, smart homes, and meeting recording.

[0003] Traditional voiceprint recognition technology mainly relies on the acoustic features of the speech signal itself as the primary basis for judgment. For example, it compares the speaker's voiceprint features using Mel-Frequency Cepstral Coefficients (MFCC), i-vector, x-vector, or deep neural network models. In particular, voiceprint recognition technology usually assumes that all voiceprint models have the same prior conditions when recognizing voiceprints. Therefore, it is often necessary to perform a full comparison of all voiceprint models in the database.

[0004] However, when the database size reaches tens of thousands or even hundreds of thousands of users, full comparison leads to significant computational latency and is prone to false acceptance due to similar voiceprint features. While existing improvements record timestamps or geographic location information from the voice data, this recorded information is typically used only for supplementary purposes, such as treating time or location information as metadata for the audio file, solely for subsequent file retrieval or post-event analysis. In other words, even after improvements, voiceprint recognition still requires a complete or large-scale comparison of the voiceprint model, resulting in massive computational demands, long response times, and a high risk of misjudgment in complex scenarios.

[0005] In summary, it is evident that existing technologies have long suffered from insufficient voiceprint recognition speed and high false positive rates in complex scenarios. Therefore, it is necessary to propose improved technical methods to address this issue. Summary of the Invention

[0006] In view of the problems of insufficient speed and high false positive rate in complex situations in existing technologies, the present invention discloses a device and method for rapid voiceprint recognition based on spatiotemporal context, and a computer-readable medium, wherein: The device for rapid voiceprint recognition based on spatiotemporal context disclosed in this invention comprises at least: a voiceprint database for recording multiple voiceprint models, multiple spatiotemporal contexts, and the degree of correlation between each voiceprint model and each spatiotemporal context; a data acquisition module for acquiring speech signals and spatiotemporal tags synchronized with the speech signals; a context judgment module for selecting the current context from multiple spatiotemporal contexts based on the spatiotemporal tags; a model selection module for filtering multiple candidate voiceprint models from multiple voiceprint models, wherein the degree of correlation between the multiple candidate voiceprint models and the current context is higher than a correlation threshold; and a voiceprint recognition module for recognizing the voiceprint in the speech signal based on the multiple candidate voiceprint models.

[0007] The method for rapid voiceprint recognition based on spatiotemporal context disclosed in this invention includes at least the following steps: establishing a voiceprint database, which contains multiple voiceprint models, multiple spatiotemporal contexts, and the degree of correlation between each voiceprint model and each spatiotemporal context; obtaining a speech signal and a spatiotemporal tag synchronized with the speech signal; selecting the current context from multiple spatiotemporal contexts based on the spatiotemporal tag; selecting multiple candidate voiceprint models from multiple voiceprint models, wherein the degree of correlation between the candidate voiceprint models and the current context is higher than a correlation threshold; and recognizing the voiceprint in the speech signal based on the multiple candidate voiceprint models.

[0008] The computer-readable medium disclosed in this invention stores a computer program thereon, which, when executed by a device, enables the device to implement the above-described method for rapid voiceprint recognition based on spatiotemporal context.

[0009] The apparatus, method, and computer-readable medium disclosed in this invention are as described above. The difference between this invention and the prior art is that, after obtaining a spatiotemporal tag synchronized with the speech signal, this invention determines the current context based on the spatiotemporal tag and selects candidate voiceprint models associated with the current context from the voiceprint database, and identifies the voiceprint in the speech signal based on the candidate voiceprint models. This solves the problems existing in the prior art and can reduce the model samples for recognition by determining the current context, thereby achieving the technical effect of improving the speed and accuracy of voiceprint recognition. Attached Figure Description

[0010] Figure 1A This is a flowchart of the method for rapid voiceprint recognition based on spatiotemporal context proposed in this invention.

[0011] Figure 1B This is a flowchart of the method for determining the correlation between a voiceprint model and spatiotemporal context proposed in this invention.

[0012] Figure 1C This is a flowchart of the method for identifying voiceprints by distinguishing sub-speech signals from speech signals, as proposed in this invention.

[0013] Figure 1D This is a flowchart of the method for determining the social relationships of the speaker of a sub-speech signal proposed in this invention.

[0014] Figure 1E This is a flowchart of the method for selecting the current context proposed in this invention.

[0015] Figure 2 This is a schematic diagram of data association proposed in an embodiment of the present invention.

[0016] Figure 3 This is a schematic diagram of the module of the device for rapid voiceprint recognition based on spatiotemporal context proposed in this invention.

[0017] Figure 4 This is a schematic diagram of the computer system of the electronic device proposed in this invention.

[0018] Explanation of reference numerals in the attached figures: Step 110: Establish a voiceprint database, which includes a voiceprint model, spatiotemporal context, and the degree of correlation between the voiceprint model and the spatiotemporal context. Step 111: Obtain the voiceprint model and associated spatiotemporal data Step 113: Determine the degree of correlation between the voiceprint model and the spatiotemporal context based on the spatiotemporal data associated with the voiceprint model. Step 115: Record the voiceprint model, spatiotemporal context, and the degree of correlation between the voiceprint model and the spatiotemporal data in the voiceprint database. Step 120: Obtain the speech signal and the spatiotemporal tag synchronized with the speech signal. Step 125: Distinguish multiple sub-speech signals from the speech signal based on voiceprint features. Step 130: Select the current context from the spatiotemporal contexts based on the spatiotemporal labels. Step 131: Extract time and spatial data from the spatiotemporal tags. Step 133: Provide temporal and spatial data to the inference model to calculate the inference probability for each spatiotemporal scenario. Step 135: Select the spatiotemporal scenario with an inference probability higher than a predetermined threshold as the current scenario. Step 150: Select candidate voiceprint models from the voiceprint models. The candidate voiceprint models should have a correlation with the current context that is higher than the correlation threshold. Step 160: Identify the voiceprint in the speech signal based on the candidate voiceprint model. Step 165: Identify the voiceprint of the sub-speech signal based on the candidate voiceprint model. Step 171: Analyze the speech content of the sub-speech signal Step 175: Determine the social relationships between the speakers who produced the sub-speech signals based on the speech content. 210: Voiceprint Database 211: Voiceprint Model 212: Spatiotemporal Context 213: Context Mapping Table 220: Voice signal 230: Time and Space Tag 300: Device 302: Input Unit 303: Communication Interface 304: Storage Media 305: Output Unit 307: Processor 310: Data Maintenance Module 320: Data Acquisition Module 330: Key Computing Module 340: Contextual Judgment Module 350: Model Selection Module 360: Voiceprint Recognition Module 380: Threshold Adjustment Module 390: Relationship Determination Module 400: Computer Systems 401: CPU 402: ROM 403: RAM 404: Bus 405: I / O Interface 406: Input section 407: Output Section 408: Storage Section 409: Communications Section 410: Driver 411: Removable media Detailed Implementation The features and implementation methods of the present invention will be described in detail below with reference to the accompanying drawings and embodiments. The content is sufficient to enable any person skilled in the art to easily and fully understand the technical means used by the present invention to solve the technical problem and to implement it accordingly, thereby achieving the effects that the present invention can achieve.

[0019] It should be noted that the accompanying drawings are incorporated in and constitute a part of this specification, illustrating embodiments consistent with the present invention, and are used together with the specification to explain the principles of the invention. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0020] The accompanying drawings provided in the following embodiments are only schematic illustrations of the basic concept of the present invention. The drawings only show the elements related to the present invention and are not drawn according to the actual number, shape and size of the elements in the actual implementation. In the actual implementation, the form, quantity and proportion of each element can be arbitrarily changed, and the layout of the elements may also be more complex.

[0021] In this invention, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0022] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, known structures and devices are shown schematically rather than in detail to avoid obscuring embodiments of the invention.

[0023] The present invention can select the current context based on the spatiotemporal label of the speech signal, and identify the voiceprint in the speech signal based on the voiceprint model associated with the current context.

[0024] This invention proposes a system for rapid voiceprint recognition based on spatiotemporal context, a method for rapid voiceprint recognition based on spatiotemporal context, and a computer-readable medium, which will be described in detail below.

[0025] The following is a preliminary step. Figure 1A The flowchart of the method for rapid voiceprint recognition based on spatiotemporal context proposed in this invention is used to illustrate the system operation process of this invention. Please refer to the following: Figure 2 A schematic diagram of data association proposed in this invention. The steps of the method for rapid voiceprint recognition based on spatiotemporal context proposed in this invention are as follows: Step 110: Establish a voiceprint database 210. The voiceprint database 210 contains multiple voiceprint models 211, multiple spatiotemporal scenarios 212, and a scenario mapping table 213. The scenario mapping table 213 records the degree of association between each voiceprint model 211 and each spatiotemporal scenario 212. The voiceprint model proposed in this invention can be an i-vector or a d-vector feature vector, but this invention is not limited to these. The spatiotemporal scenario proposed in this invention includes time information and spatial information. The time information can be an absolute time representing a specific time or a specific time period, or a time segment representing a periodic time (such as Monday morning, Tuesday afternoon from 2 pm to 5 pm, etc.). The spatial information can be the actual latitude and longitude and altitude, indoor positioning tags, logical area names (such as meeting room B in office A, etc.), but this invention is not limited to these. The degree of association proposed in this invention can be the frequency of occurrence or association strength of the voiceprint model in a specific spatiotemporal scenario, and / or the probability distribution of the voiceprint model appearing in a specific spatiotemporal scenario. In some embodiments, the same voiceprint model can be associated with multiple different spatiotemporal contexts, and the spatiotemporal contexts associated with the same voiceprint model usually correspond to different degrees of association, which makes the voiceprint model may or may not become a candidate voiceprint model in different spatiotemporal contexts.

[0026] In some embodiments, when establishing a voiceprint database, historical recognition records can be used as data samples. In this case, it is necessary to determine the degree of correlation between each voiceprint model in the historical recognition records and various spatiotemporal contexts in the voiceprint database. That is, for example... Figure 1B As shown in the flowchart, step 110 may also include the following steps: Step 111: Obtain each voiceprint model and the spatiotemporal data associated with each voiceprint model. Generally, voiceprint models and the spatiotemporal data associated with them can be obtained from historical identification records as data samples.

[0027] Step 113: Determine the degree of association between each voiceprint model and each spatiotemporal context based on the obtained spatiotemporal data associated with each voiceprint model. For example, a probability model can be used to calculate the degree of association (occurrence frequency or probability distribution) of each voiceprint model in each spatiotemporal context, or other statistical models, machine learning models (such as random forests), deep learning models (such as Transformer architecture), or knowledge graphs can be used to mine the relationship between people, places, and times to determine the degree of association between each voiceprint model and each spatiotemporal context. For example, if a certain voiceprint model frequently appears in the meeting room on Friday evenings, then the degree of association of that voiceprint model with the spatiotemporal context of the meeting room on Friday evenings will increase significantly.

[0028] Step 115: Record each voiceprint model, each spatiotemporal context, and the degree of correlation between each voiceprint model and each spatiotemporal context in the voiceprint database.

[0029] Step 120: Obtain the voice signal 220 and the spatiotemporal tag 230 synchronized with the voice signal. The spatiotemporal tag proposed in this invention may contain time data and spatial data. The spatial data may be GPS positioning data, location identification data representing a specific space within a building, etc., but this invention is not limited thereto. In this invention, while obtaining the voice signal, the synchronization of spatiotemporal information must also be ensured; the spatiotemporal tag proposed in this invention plays the role of data navigation. In some embodiments, the spatiotemporal tag may also contain hash values ​​of time data and spatial data, tag identification data, or specific encoding, etc., and this invention has no particular limitations.

[0030] In step 120, the acquired speech signal may be considered as a whole; for example, the speech signal may contain only one person's voice. In this case, the acquired spatiotemporal label corresponds to the entire speech signal in time. In step 120, if the acquired speech signal contains multiple sub-speech signals with different voiceprint features, that is, if the speech signal contains a multi-person dialogue, then... Figure 1C As shown in the flowchart, step 120 may also include step 125, which is to distinguish (diarize) multiple sub-speech signals from the speech signal based on the voiceprint features.

[0031] In step 120, the speech signal can also be segmented into multiple audio frames, each with a different spatiotemporal label. For example, in step 120, the sensitivity and signal-to-noise ratio (SNR) threshold of the segmented speech signal can be dynamically adjusted according to the current context. That is, the sensitivity and SNR threshold can be adjusted based on the spatial information represented by the spatiotemporal label. For instance, if the spatiotemporal label indicates that the current spatial information is in a noisy environment (such as a factory), the SNR threshold can be increased to filter background voices; if the spatiotemporal label indicates that the spatial information is in a quiet environment (such as an office), the SNR threshold can be decreased to capture subtle speech.

[0032] It should be noted that step 120 may be receiving the voice signal acquired by the voice acquisition device (not shown in the figure) and the spatiotemporal tag generated by the voice acquisition device when acquiring the voice signal. The voice acquisition device includes, but is not limited to, mobile phones, Internet of Things (IoT) microphones, in-vehicle devices, or video conferencing equipment. For example, when a smartphone starts recording, it can simultaneously record GPS coordinates and system time, and encapsulate the obtained GPS coordinates and system time into a data packet of voice information. Step 120 may also be receiving the voice signal and spatiotemporal tag acquired by different data acquisition devices (not shown in the figure) and simultaneously receiving the received voice signal and spatiotemporal tag. When the voice acquisition device cannot generate a spatiotemporal tag or cannot determine the spatial location, step 120 may also receive the voice signal acquired by a voice acquisition device such as a traditional microphone, extract environmental sound features from the received voice signal to generate an acoustic fingerprint, and determine spatial data based on the extracted acoustic features to generate a spatiotemporal tag, and generate voiceprint feature data corresponding to the voice signal. For example, environmental sound features such as ambient noise spectrum and reverberation characteristics in the speech signal can be analyzed, and the acoustic model of the space that the sound features conform to can be determined based on the analysis results, such as a large atrium, and then the spatial data can be inferred to be the company's first-floor lobby, thus generating a spatiotemporal tag. However, the present invention is not limited to this. For example, step 120 can also detect the device identification data of nearby Bluetooth (BT) devices and compare the detected device identification data with the device identification data existing in the known space to determine the spatial data. Alternatively, background sound can be separated from the speech signal and the spectral characteristics of the background sound can be obtained, thereby determining the acoustic model of the space that conforms to the obtained spectral characteristics to determine the spatial data.

[0033] Step 130: Based on the spatiotemporal tag 230 obtained in step 120, the current context is selected from the spatiotemporal contexts 212 contained in the voiceprint database 210. More specifically, step 130 may include, for example... Figure 1E Steps: Step 131: Extract temporal and spatial data from the spatiotemporal tags. For example, geofencing or indoor positioning technology can be used to convert spatial data into logical area identification data.

[0034] Step 133: Provide the temporal and spatial data to the inference model to calculate the inference probability for each spatiotemporal scenario. For example, when the spatial data is within the company's latitude and longitude range and the temporal data is during working hours, the inference model can calculate the inference probability for spatiotemporal scenarios such as office scenarios, meeting scenarios, and dining scenarios. The inference model mentioned above can be a Bayesian inference model, or it can be the probability model, statistical model, machine learning model, deep learning model, or knowledge graph mentioned in step 113.

[0035] Step 135: Select the spatiotemporal scenario with an inference probability higher than a predetermined threshold as the current scenario. For example, if only one spatiotemporal scenario (such as an office scenario) has an inference probability higher than the predetermined threshold among the inference probabilities generated by the inference model, the office scenario can be selected as the current scenario. If multiple spatiotemporal scenarios have an inference probability higher than the predetermined threshold, such as the office scenario and the meeting scenario generated by the inference model both having an inference probability higher than the predetermined threshold, the spatiotemporal scenario with the highest inference probability can be selected as the current scenario.

[0036] The following will continue to explain. Figure 1A The steps in the process.

[0037] Step 150: Multiple candidate voiceprint models are selected from the voiceprint models 211 contained in the voiceprint database 210. The selected candidate voiceprint models have a higher correlation with the current context selected in step 130 than the correlation threshold. For example, in step 150, probability threshold filtering, Top-N ranking filtering, and weighted summation (integrating time, space, and device weights) can be used to filter voiceprint models to select candidate voiceprint models. If the correlation threshold is set to 0.7, only voiceprint models with a correlation probability greater than 0.7 with the current context (such as an office) (such as the voiceprint models of 20 members of a specific department) can be selected as candidate voiceprint models.

[0038] Step 160: Identify the voiceprint in the speech signal 220 obtained in step 120 based on the candidate voiceprint model selected in step 150, thereby determining the speaker's identity and providing it for subsequent use.

[0039] If the speech signal obtained in step 120 contains multiple sub-speech signals, then as follows Figure 1C As shown in the flowchart, step 160 may also include step 165, which identifies the voiceprint in the sub-speech signal obtained in step 120 based on the candidate voiceprint model.

[0040] Subsequently, the present invention can also be as follows: Figure 1D As shown in the flowchart, after step 165, the following steps are performed: Step 171: Analyze the speech content of each sub-speech signal obtained in step 165.

[0041] Step 175: Determine the social relationship of the speakers who issued each sub-voice signal based on the voice content obtained in Step 171, so as to provide information for subsequent use. For example, if two sub-voice signals are identified, and the voice content of one of the sub-voice signals contains keywords such as "manager" or "please approve leave slip," then the social relationship between the two unspeaking individuals can be determined to be a superior-subordinate relationship. However, the social relationship proposed in this invention is not limited to a superior-subordinate relationship; other relationships such as friendships or family member relationships can also be considered social relationships proposed in this invention.

[0042] Thus, through this invention, a voiceprint recognition architecture based on the current context can be established. The current spatiotemporal label is obtained through the voice signal, and highly relevant candidate voiceprint models are selected from the voiceprint database for comparison, rather than performing a full database comparison, thereby accelerating the efficiency of voiceprint recognition.

[0043] The present invention may also perform the following steps before step 150: adjusting the association threshold used in step 150 based on the reliability (such as the accuracy of spatial information) or completeness of the spatiotemporal tags obtained in step 120. For example, if the error range of the GPS positioning data is too large, it indicates that the reliability of the spatiotemporal tags is low, and the association threshold can be lowered to expand the range of candidate models and avoid omissions.

[0044] Furthermore, this invention can also ensure the timeliness of the voiceprint database. More practically, after step 160, the following steps can be performed: after successfully identifying a voiceprint model, increasing the correlation between the voiceprint model and the spatiotemporal context in the voiceprint database based on the successfully identified voiceprint model and the corresponding spatiotemporal data, and / or periodically reducing the correlation of voiceprint models that were not successfully identified in a specific spatiotemporal context for a certain period of time in the voiceprint database based on the historical records of voiceprint model identification. For example, if a certain voiceprint model has not appeared in a specific time context for a certain period of time, the correlation between the voiceprint model and the spatiotemporal context can be reduced. For example, if an employee leaves the company for a year, the correlation between the employee's voiceprint model and the spatiotemporal context of the office will decrease periodically (e.g., exponentially), avoiding outdated data from occupying identification resources. In this way, the correlation can automatically decrease over time, making the impact of recent behavior on the correlation greater than that of historical behavior, avoiding long-term accumulation leading to stagnation.

[0045] The following continues with Figure 3 The schematic diagram of the device for rapid voiceprint recognition based on spatiotemporal context proposed in this invention illustrates the device for implementing this invention. For example... Figure 3As shown, the device 300 of the present invention includes an input unit 302, a communication interface 303, a storage medium 304, an output unit 305, and a processor 307. The processor 307 is interconnected with the input unit 302, the communication interface 303, the storage medium 304, and the output unit 305 via a bus (not shown).

[0046] Input unit 302 can provide input data through peripheral input devices of device 300. For example, input unit 302 can input data or commands through keyboard, mouse, touchpad, touch screen, or input sound signals through microphone.

[0047] The communication interface 303 can be connected to network storage devices or servers (not shown in the figure) outside the device 300, and can request and download data from the connected network devices.

[0048] Storage medium 304 is typically a large-capacity storage area of ​​device 300. It can store data or signals downloaded by communication interface 303, data or signals provided to processor 307, and data or signals generated by processor 307.

[0049] The output unit 305 can also output the data generated by the processor 307 through the peripheral output device of the device 300. For example, the output unit 307 can display the data through a display or a touch screen.

[0050] The processor 307 may include modules such as a data maintenance module 310, a data acquisition module 320, a context judgment module 340, a model selection module 350, and a voiceprint recognition module 360, and may also include attachable modules such as an association calculation module 330, a threshold adjustment module 380, and a relationship determination module 390. Specifically, the data maintenance module 310, the association calculation module 330, and the model selection module 350 are connected; the data acquisition module 320 is connected to the association calculation module 330, the context judgment module 340, and the relationship determination module 390; and the model selection module 350 is connected to the context judgment module 340, the voiceprint recognition module 360, and the threshold adjustment module 380.

[0051] In some embodiments, processor 307 can execute one or more sets of computer instructions stored in memory (not shown), and can generate the included modules after executing the computer instructions; in other embodiments, the modules included in processor 307 can be generated by one or more circuits and / or complete or partial chips and other hardware elements, that is, processor 307 includes hardware elements that make up the included modules. In other words, the modules included in processor 307 can be software modules or hardware modules, and the present invention has no particular limitations.

[0052] The data maintenance module 310 is responsible for reading, writing, and organizing the voiceprint database to maintain it. The voiceprint database can record multiple voiceprint models, multiple spatiotemporal scenarios, and scenario mapping tables. The scenario mapping table contains the degree of association between each voiceprint model and each spatiotemporal scenario. The data maintenance module 310 can also provide the spatiotemporal data associated with the voiceprint models to the association calculation module 330, thereby allowing the association calculation module 330 to obtain the degree of association between each voiceprint model and each spatiotemporal scenario.

[0053] The data acquisition module 320 is responsible for acquiring the voice signal and the spatiotemporal tag synchronized with the voice signal. In some embodiments, the data acquisition module 320 can be connected to a microphone array, a GPS chip, and a system clock simultaneously, and can have multi-source synchronization capabilities. For example, the data acquisition module 320 can have a preprocessing function, which can characterize the acquired voice signal and encapsulate it with the spatiotemporal data into a single object for use by subsequent modules. More specifically, the data acquisition module 320 can receive the voice signal and spatiotemporal tag acquired by the voice acquisition device; the data acquisition module 320 can also receive the voice signal and spatiotemporal tag acquired by different data acquisition devices, and can synchronize the received voice signal and spatiotemporal tag; the data acquisition module 320 can also receive the voice signal acquired by the voice acquisition device, and can extract the acoustic fingerprint of the voice signal, and can determine the spatial data based on the extracted acoustic features to generate a spatiotemporal tag, and can generate voiceprint feature data corresponding to the voice signal.

[0054] In some embodiments, the data acquisition module 320 can diaarize multiple sub-speech signals from the speech signal based on voiceprint features.

[0055] The association calculation module 330 can calculate and continuously adjust the association degree between the voiceprint model and the spatiotemporal context in the voiceprint database (context mapping table). More specifically, the association calculation module 330 can use a model or algorithm to calculate the association degree between the voiceprint model and the spatiotemporal context based on the voiceprint model and corresponding spatiotemporal data provided by the data maintenance module 310. Alternatively, after the voiceprint recognition module 360 ​​successfully recognizes the voiceprint model, it can recalculate the association degree between the voiceprint model and the spatiotemporal context based on the successfully recognized voiceprint model and corresponding spatiotemporal data (the recalculated association degree will be higher than the original association degree). Or, the association calculation module 330 can periodically reduce the association degree in the voiceprint database of voiceprint models that have not been successfully recognized in a specific spatiotemporal context for a certain period of time based on the historical records of voiceprint model recognition, and periodically clean up outdated association data.

[0056] The context judgment module 340 is responsible for selecting a current context from the spatiotemporal contexts in the voiceprint database based on the spatiotemporal tags obtained by the data acquisition module 320. For example, the context judgment module 340 can extract time data and spatial data from the spatiotemporal tags, and provide the obtained time data and spatial data to the inference model to calculate the inference probability of each spatiotemporal context, and select the spatiotemporal context whose calculated inference probability is higher than a predetermined threshold as the current context.

[0057] The model selection module 350 is responsible for filtering multiple candidate voiceprint models from the voiceprint models in the voiceprint database. The degree of correlation between the candidate voiceprint models selected by the model selection module 350 and the current context selected by the context judgment module 340 is higher than the correlation threshold.

[0058] The voiceprint recognition module 360 ​​is responsible for performing vector comparison operations (such as Cosine Similarity) to identify the voiceprint in the speech signal obtained by the data acquisition module 320 based on the candidate voiceprint models selected by the model selection module 350, thereby providing recognition results for subsequent use. If the data acquisition module 320 obtains multiple sub-speech signals from the speech signal, the voiceprint recognition module 360 ​​can identify the voiceprint of each sub-speech signal based on the candidate voiceprint models. Since the number of candidate voiceprint models has been significantly reduced by the model selection module 350, the voiceprint recognition module 360 ​​can achieve high-precision recognition with extremely low resource consumption.

[0059] The threshold adjustment module 380 can monitor the quality of environmental perception and then adjust the association threshold used by the model selection module 350 to screen voiceprint models based on the reliability or integrity of the spatiotemporal tags obtained by the data acquisition module 320. For example, when the noisy environment causes unstable acoustic fingerprint extraction, the threshold adjustment module 380 can send a signal to the model selection module 350 to lower the association threshold to increase the number of candidate voiceprint models, thereby maintaining the error tolerance of recognition.

[0060] The relationship determination module 390 has a built-in speech-to-text (STT) and natural language processing (NLP) engine. It can identify the speech content of each sub-speech signal acquired by the data acquisition module 320, and determine the social relationship of the speaker who issued each sub-speech signal based on the analysis results of each speech content (such as title and tone). It can also provide the determined social relationship for subsequent use.

[0061] In summary, the difference between this invention and the prior art lies in the technical means of obtaining a spatiotemporal label synchronized with the speech signal, determining the current context based on the spatiotemporal label, selecting candidate voiceprint models associated with the current context from the voiceprint database, and identifying voiceprints in the speech signal based on the candidate voiceprint models. This technical means can solve the problems of insufficient speed for full-database voiceprint recognition and high misjudgment rate in complex situations in the prior art, thereby reducing the model samples for recognition based on the current context to achieve the technical effect of improving the speed and accuracy of voiceprint recognition.

[0062] Figure 4 This is a schematic diagram of an electronic device provided in one embodiment of the present invention. It should be noted that... Figure 4 The computer system 400 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.

[0063] like Figure 4 As shown, the computer system 400 includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 402 or programs loaded from storage portion 408 into Random Access Memory (RAM) 403, such as performing the methods described in the above embodiments. The RAM 403 also stores various programs and data required for system operation. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An Input / Output (I / O) interface 405 is also connected to the bus 404.

[0064] The following components are connected to the input / output interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface controller such as a local area network (LAN) card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output interface 405 as needed. Removable media 411, such as floppy disks, optical disks, magneto-optical disks, semiconductor memory, etc., are installed on the drive 410 as needed so that computer programs read from them can be installed into the storage section 408 as needed.

[0065] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit 401, it performs various functions defined in the system of the present invention.

[0066] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or element, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a floppy disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage, magnetic random access memory (MRAM), or any suitable combination thereof. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband frequency or as part of a carrier wave, wherein a computer-readable computer program is carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or element. A computer program contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0067] The flowcharts and schematic diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in the flowchart or schematic diagram may represent a module, a program segment, or a portion of program code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed simultaneously, or sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the schematic diagram or flowchart, and combinations of blocks in the schematic diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0068] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0069] Another aspect of the present invention provides a computer-readable medium having a computer program stored thereon. When executed by a processor of a device (such as a computer) in a distributed network environment, the computer program enables the device to implement the method for rapid voiceprint recognition based on spatiotemporal context as described above. This computer-readable medium may be included in the electronic device described in the above embodiments, or it may exist independently and not incorporated into the electronic device.

[0070] Another aspect of the present invention provides a computer program product or computer program including computer instructions stored in a computer-readable medium. A processor of a computer device reads the computer instructions from the computer-readable medium and executes the computer instructions, causing the computer device to perform the method for rapid voiceprint recognition based on spatiotemporal context provided in the various embodiments above.

[0071] While the embodiments disclosed in this invention are as described above, the content is not intended to directly limit the scope of patent protection for this invention. Any modifications or refinements made by those skilled in the art to which this invention pertains, in terms of form and detail, without departing from the spirit and scope disclosed herein, shall fall within the scope of patent protection for this invention. The scope of patent protection for this invention shall still be determined by the scope defined in the appended claims.

Claims

1. A method for rapid voiceprint recognition based on spatiotemporal context, applied to a device, the method comprising at least the following steps: Establish a voiceprint database, which contains multiple voiceprint models, multiple spatiotemporal scenarios, and the degree of correlation between each voiceprint model and each spatiotemporal scenario; Obtain the speech signal and the spatiotemporal label synchronized with the speech signal; The current context is selected from multiple spatiotemporal contexts based on the spatiotemporal label; Multiple candidate voiceprint models are selected from these multiple voiceprint models. The correlation between these candidate voiceprint models and the current context is higher than a correlation threshold. The voiceprint in the speech signal is identified based on the multiple candidate voiceprint models.

2. The method for rapid voiceprint recognition based on spatiotemporal context as described in claim 1, characterized in that, The steps of establishing the voiceprint database also include determining the degree of association between each voiceprint model and each spatiotemporal context based on the spatiotemporal data associated with each voiceprint model.

3. The method for rapid voiceprint recognition based on spatiotemporal context as described in claim 1, characterized in that, The steps for obtaining the voice signal and the spatiotemporal tag synchronized with the voice signal are as follows: receiving the voice signal and the spatiotemporal tag obtained by the voice acquisition device; receiving the voice signal and the spatiotemporal tag obtained by different data acquisition devices and synchronizing the voice signal and the spatiotemporal tag; receiving the voice signal obtained by another voice acquisition device and extracting the acoustic fingerprint of the voice signal and judging the spatial data based on the acoustic features to generate the spatiotemporal tag.

4. The method for rapid voiceprint recognition based on spatiotemporal context as described in claim 1, characterized in that, The step of obtaining the speech signal further includes the step of distinguishing multiple sub-speech signals from the speech signal based on the voiceprint features, and the step of identifying the voiceprint in the speech signal based on the multiple candidate voiceprint models is to identify the voiceprint of each sub-speech signal based on the multiple candidate voiceprint models.

5. The method for rapid voiceprint recognition based on spatiotemporal context as described in claim 4, characterized in that, After identifying the voiceprint in the speech signal based on the multiple candidate voiceprint models, the method further includes identifying the speech content of each sub-speech signal and determining the social relationship of the speaker who issued each sub-speech signal based on the multiple speech content.

6. The method for rapid voiceprint recognition based on spatiotemporal context as described in claim 1, characterized in that, The step of selecting the current situation from the multiple spatiotemporal situations based on the spatiotemporal label further includes retrieving time data and spatial data from the spatiotemporal label, providing the time data and spatial data to the inference model to calculate the inference probability of each spatiotemporal situation, and selecting the spatiotemporal situation with the inference probability higher than a predetermined threshold as the current situation.

7. The method for rapid voiceprint recognition based on spatiotemporal context as described in claim 1, characterized in that, Before the step of selecting the multiple candidate voiceprint models from the multiple voiceprint models, the method further includes a step of adjusting the association threshold based on the reliability or integrity of the spatiotemporal tag.

8. A device for rapid voiceprint recognition based on spatiotemporal context, the device comprising at least: A voiceprint database is used to record multiple voiceprint models, multiple spatiotemporal contexts, and the degree of correlation between each voiceprint model and each spatiotemporal context. The data acquisition module is used to acquire the voice signal and the spatiotemporal tag synchronized with the voice signal; The context judgment module is used to select the current context from multiple spatiotemporal contexts based on the spatiotemporal label; The model selection module is used to filter out multiple candidate voiceprint models from the multiple voiceprint models, and the degree of correlation between the multiple candidate voiceprint models and the current context is higher than the correlation threshold. and The voiceprint recognition module is used to identify the voiceprint in the speech signal based on the multiple candidate voiceprint models.

9. A computer-readable medium storing a computer program that, when executed by a device, causes the device to perform the method for rapid voiceprint recognition based on spatiotemporal context as described in any one of claims 1-7.