Voice-based information recognition method, device, equipment, medium and program product

CN117219056BActive Publication Date: 2026-10-09SHENZHEN TENCENT COMP SYST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211513242.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2026-10-09
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

[0003]在相关技术中,对于语音的信息识别,通常是在注册阶段,将主体对象在进行语音注册时的语音特征进行存储,在识别时,通过将待识别语音的语音特征和语音注册时的语音特征进行比较,从而确定待识别语音的目标主体对象,但是由于语音容易受到外界干扰,相关技术中单一的识别方式会导致语音的识别准确性较低

Benefits of technology

通过待识别语音的第一语音特征与各主体对象在进行语音注册时的注册语音的第二语音特征之间的第一特征相似度,当第一特征相似度大于相似度阈值的第二语音特征的第一数量为多个或零个时,且多个主体对象中存在具有衍生语音特征的主体对象时,确定待识别语音的第一语音特征与衍生语音特征的第二特征相似度,基于第一特征相似度和第二特征相似度,从多个主体对象中,确定待识别语音的目标主体对象。如此,通过将待识别语音的第一语音特征分别与多个主体对象在语音注册时的第二语音特征,以及衍生语音特征进行比较,得到第一特征相似度和第二特征相似度,通过第一特征相似度和第二特征相似度,识别待识别语音的目标主体对象。由于主体对象的衍生语音特征是通过对主体对象的历史语音特征进行更新得到的,从而使得主体对象的衍生语音特征在更新过程中,不断的吸纳历史语音特征的语音特性,从而能够更加准确的反映主体对象的语音在不同时段的变化,从而使得在识别待识别语音的目标主体对象时,综合考虑待识别语音的第一语音特征与衍生语音特征的第二特征相似度,从而显著提高了语音的识别准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117219056B_ABST
    Figure CN117219056B_ABST
Patent Text Reader

Abstract

The application provides a voice-based information recognition method, device, equipment, medium and program product; the method comprises: obtaining a first voice feature of a to-be-recognized voice, and a second voice feature of a registration voice of a plurality of subject objects when performing voice registration; respectively determining a first feature similarity of the first voice feature and each second voice feature; comparing each first feature similarity with a similarity threshold, and based on the comparison result, determining a first number of second voice features whose first feature similarity is greater than the similarity threshold; when the first number is a plurality or zero, and there is a subject object with a derived voice feature in the plurality of subject objects, determining a second feature similarity of the first voice feature and the derived voice feature; based on the second feature similarity and the first feature similarity, determining a target subject object of the to-be-recognized voice from the plurality of subject objects. Through the application, the recognition accuracy of the voice can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a voice-based information recognition method, apparatus, electronic device, and storage medium. Background Technology

[0002] Speech, as the acoustic representation of language, is one of the most natural, effective, and convenient means for humans to exchange information. In recent years, speech recognition technology has made great progress. However, when people input speech, they are inevitably interfered with by the voices of different speakers in the same environment.

[0003] In related technologies, speech recognition typically involves storing the speech features of the subject during the registration phase. During recognition, the speech features of the subject to be recognized are compared with those registered to determine the target subject. However, because speech is easily affected by external interference, the single recognition method in related technologies leads to low accuracy in speech recognition. Summary of the Invention

[0004] This application provides a speech-based information recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can significantly improve the accuracy of speech recognition.

[0005] The technical solution of this application embodiment is implemented as follows: This application provides a speech-based information recognition method, including: The first speech feature of the speech to be recognized and the second speech feature of the registered speech of multiple subjects during speech registration are obtained. Determine the first feature similarity between the first speech feature and each of the second speech features; The similarity of each of the first features is compared with a similarity threshold, and based on the comparison results, a first number of second speech features whose first feature similarity is greater than the similarity threshold is determined; When the first quantity is multiple or zero, and there is a subject object with derived speech features among the multiple subject objects, the second feature similarity between the first speech feature and the derived speech feature is determined; The derived speech features are obtained by updating the historical speech features of the subject object; Based on the second feature similarity and the first feature similarity, the target subject of the speech to be recognized is determined from the plurality of subject objects.

[0006] This application provides a voice-based information recognition device, including: The acquisition module is used to acquire the first speech features of the speech to be recognized, and the second speech features of the registered speech of multiple subject objects when they are registering their speech. The first feature similarity determination module is used to determine the first feature similarity between the first speech feature and each of the second speech features respectively; The comparison module is used to compare the similarity of each of the first features with a similarity threshold, and based on the comparison result, determine a first number of second speech features whose first feature similarity is greater than the similarity threshold; The second feature similarity determination module is used to determine the second feature similarity between the first speech feature and the derived speech feature when the first quantity is multiple or zero, and there is a subject object with derived speech features among the multiple subject objects; wherein, the derived speech feature is obtained by updating the historical speech features of the subject object; The target subject object determination module is used to determine the target subject object of the speech to be recognized from the plurality of subject objects based on the second feature similarity and the first feature similarity.

[0007] In some embodiments, the above-described speech-based information recognition device further includes: a supplementary determination module, configured to, when the first quantity is one, determine the subject object corresponding to the second speech feature whose first feature similarity is greater than the similarity threshold as the target subject object of the speech to be recognized; when the first quantity is multiple or zero, and there is no subject object with the derived speech feature among the multiple subject objects, determine the subject object corresponding to the second target speech feature as the target subject object of the speech to be recognized; wherein, the second target speech feature is the second speech feature corresponding to the largest first feature similarity.

[0008] In some embodiments, the target subject object determination module is further configured to determine the maximum value between the second feature similarity and the first feature similarity; and determine the subject object corresponding to the maximum value as the target subject object of the speech to be recognized.

[0009] In some embodiments, the above-described speech-based information recognition device further includes: an update module, configured to acquire derived information of the target subject object, the derived information indicating whether the target subject object has the derived speech feature; when the derived information indicates that the target subject object does not have the derived speech feature, update the second speech feature of the target subject object and generate the derived speech feature of the target subject object; when the derived information indicates that the target subject object has the derived speech feature, update the second speech feature of the target subject object and the derived speech feature of the target subject object.

[0010] In some embodiments, the above-described updating module is further configured to update the second speech feature of the target subject object based on the first speech feature when the derived information indicates that the target subject object does not have the derived speech feature, thereby obtaining the updated second speech feature of the target subject object; and to determine the first speech feature as the derived speech feature of the target subject object.

[0011] In some embodiments, the update module is further configured to obtain a first weight of the first speech feature and a second weight of the second speech feature of the target subject object, wherein the sum of the first weight and the second weight is equal to 1; and to perform a weighted summation of the first speech feature and the second speech feature of the target subject object according to the first weight and the second weight, respectively, to obtain the updated second speech feature of the target subject object.

[0012] In some embodiments, the above-mentioned updating module is further configured to, when the derived information indicates that the target subject object has the derived speech features, update the derived speech features of the target subject object based on the first speech features to obtain updated derived speech features; and update the second speech features of the target subject object based on the updated derived speech features to obtain updated second speech features of the target subject object.

[0013] In some embodiments, the update module is further configured to obtain a second number of derived speech features of the target subject object; when the second number is one, the derived speech features of the target subject object are determined as speech features to be updated, and the speech features to be updated are updated based on the first speech features to obtain the updated derived speech features; when the second number is multiple, a target derived speech feature is selected from the derived speech features of the target subject object based on the second number, and the target derived speech feature is updated based on the first speech feature to obtain the updated derived speech features.

[0014] In some embodiments, the update module is further configured to obtain the third feature similarity between the first speech feature and each derived speech feature of the target subject object; when the second number is less than a number threshold, the derived speech feature corresponding to the largest third feature similarity is determined as the target derived speech feature; when the second number is greater than or equal to the number threshold, the following processing is performed on each derived speech feature of the target subject object: the derived speech feature is determined as a reference speech feature, and the fourth feature similarity between the reference speech feature and each other speech feature is determined, and the target derived speech feature is determined based on the third feature similarity and the fourth feature similarity; wherein, the other speech features are derived speech features of the target subject object other than the reference speech feature.

[0015] In some embodiments, the above-mentioned updating module is further configured to obtain a third weight of the first speech feature and a fourth weight of the speech feature to be updated, wherein the sum of the third weight and the fourth weight is 1; and to perform a weighted summation of the first speech feature and the speech feature to be updated according to the third weight and the fourth weight, respectively, to obtain the updated derived speech feature of the target subject object.

[0016] In some embodiments, the update module is further configured to obtain the fifth weight of the updated derived speech feature, the sixth weight of the second speech feature of the target subject object, and the seventh weight of the first speech feature; wherein the sum of the fifth weight, the sixth weight, and the seventh weight is equal to 1; the updated derived speech feature, the second speech feature of the target subject object, and the first speech feature are weighted and summed according to the fifth weight, the sixth weight, and the seventh weight, respectively, to obtain the updated second speech feature of the target subject object.

[0017] This application provides an electronic device, including: Memory is used to store executable instructions or computer programs. When a processor executes computer-executable instructions or computer programs stored in the memory, it implements the voice-based information recognition method provided in the embodiments of this application.

[0018] This application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the speech-based information recognition method provided in this application.

[0019] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the speech-based information recognition method described above in this application.

[0020] The embodiments of this application have the following beneficial effects: The first feature similarity between the first speech feature of the speech to be identified and the second speech feature of each subject object during speech registration is determined. If the number of second speech features with a first feature similarity greater than a similarity threshold is multiple or zero, and among the multiple subject objects, there exists a subject object with derived speech features, the second feature similarity between the first speech feature of the speech to be identified and the derived speech feature is determined. Based on the first feature similarity and the second feature similarity, the target subject object of the speech to be identified is determined from among the multiple subject objects. Thus, by comparing the first speech feature of the speech to be identified with the second speech feature of multiple subject objects during speech registration, as well as the derived speech feature, the first feature similarity and the second feature similarity are obtained. The target subject object of the speech to be identified is then identified using the first feature similarity and the second feature similarity. Since the derived speech features of the subject are obtained by updating the historical speech features of the subject, the derived speech features of the subject continuously absorb the speech characteristics of the historical speech features during the update process. This allows them to more accurately reflect the changes in the speech of the subject at different times. As a result, when identifying the target subject of the speech to be identified, the similarity between the first speech feature of the speech to be identified and the second feature of the derived speech feature is comprehensively considered, which significantly improves the accuracy of speech recognition. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the architecture of a voice-based information recognition system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of an electronic device for speech-based information recognition provided in an embodiment of this application; Figure 3 This is a schematic flowchart of the speech-based information recognition method provided in the embodiments of this application; Figures 4 to 5 This is a schematic diagram illustrating the principle of the speech-based information recognition method provided in the embodiments of this application; Figures 6 to 9 This is a schematic flowchart of the speech-based information recognition method provided in the embodiments of this application; Figure 10This is a schematic diagram illustrating the effect of the speech-based information recognition method provided in the embodiments of this application. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0024] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0026] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0027] 1) Biometric technology: refers to the use of computers in close combination with high-tech means such as optics, acoustics, biosensors and biostatistics principles to identify information by utilizing the inherent physiological characteristics of the human body (such as fingerprints, face, iris, etc.) and behavioral characteristics (such as handwriting, voice, gait, etc.).

[0028] 2) Speech: As an analog signal carrying specific information, speech has become an important carrier for obtaining and disseminating information in people's social lives. The purpose of speech signal processing is to extract effective speech information in complex speech environments. The impact of environmental interference on the signal during speech propagation should not be underestimated; therefore, the noise resistance of speech signal processing has become an important research direction.

[0029] 3) Subject: refers to the subject that produces the speech to be recognized or registered, such as human subjects, animal subjects, etc.

[0030] 4) Voice Recognition: Voice recognition is an interdisciplinary field. Over the past two decades, voice recognition technology has made significant progress, moving from the laboratory to the market. It will enter various fields such as industry, home appliances, communications, automotive electronics, healthcare, home services, and consumer electronics. The areas involved in voice recognition include signal processing, pattern recognition, probability theory and information theory, vocalization and auditory mechanisms, artificial intelligence, and more.

[0031] 5) Derived Speech Features: Derived speech features are speech features evolved and updated from the historical speech features of the subject. They reflect the speech characteristics of the subject's speech during previous speech recognition stages. The subject's historical speech features include the first speech features of the speech to be recognized during those stages. Before speech recognition, the subject has no corresponding historical speech features and therefore no derived speech features. After a speech recognition operation, the first speech features of the speech to be recognized during that first operation transform into historical speech features. At this point, the subject possesses historical speech features, which then evolve into derived speech features. The generation of derived speech features occurs after each speech recognition operation, when the subject's historical speech features are updated. Derived speech features are obtained by updating the subject's historical speech features, which are extracted from the subject's already recognized speech. For example, when a subject undergoes speech recognition for the first time, it has no corresponding historical speech features. After the first speech recognition is completed, the first speech features of the subject evolve into the derived speech features of the subject. When the subject undergoes speech recognition for the Nth time (… At this point, the subject object has corresponding historical speech features (i.e., the first speech features of the speech to be recognized during the previous N-1 speech recognitions). After the subject object completes the Nth speech recognition, the first speech features of the speech to be recognized during the Nth speech recognition are transformed into historical speech features. At this point, the historical speech features of the subject object (i.e., the first speech features of the speech to be recognized during the previous N speech recognitions) evolve into the derived speech features of the subject object through the first speech features of the speech to be recognized during the previous N speech recognitions.

[0032] During the implementation of the embodiments of this application, the applicant discovered the following problems with the related technology: In related technologies, for speech information recognition, the speech features of the subject object at the time of speech registration are usually stored during the registration stage. During recognition, the speech features of the speech to be recognized are compared with the speech features at the time of speech registration to determine the target subject object of the speech to be recognized.

[0033] In related technologies, speaker identification systems are trained using audio data from a large number of different speakers to obtain discriminative audio features. However, the performance of speaker identification systems using only a single audio file is limited by the registered audio file; the text content, pronunciation style, and timing of the registered audio file all affect the performance of the speech features. Even if the same person speaks two different sentences at the same time, the similarity score between the two audio features will not be 100%. Therefore, the limitation of single-template-based speaker recognition systems is that they are limited by the scope of the registered template's speech coverage and cannot be dynamically updated, resulting in a gradual deterioration in performance and user experience over time.

[0034] This application provides a speech-based information recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can significantly improve the accuracy of speech recognition. The following describes exemplary applications of the speech-based information recognition system provided in this application.

[0035] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of a voice-based information recognition system 100 provided in an embodiment of this application. The terminal (terminal 400 is shown as an example) connects to the server 200 through a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0036] Terminal 400 is used by a user to access client 410 and display the target object on a graphical interface 410-1 (graphical interface 410-1 is shown as an example). Terminal 400 and server 200 are interconnected via wired or wireless network.

[0037] In some embodiments, server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal 400 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, smartwatch, in-vehicle terminal, etc., but is not limited to these. The electronic device provided in this application embodiment can be implemented as a terminal or a server. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application embodiment.

[0038] In some embodiments, server 200 obtains the speech to be recognized from terminal 400, determines the first speech feature of the speech to be recognized and the second speech feature of the registered speech of multiple subject objects when they register their speech, determines the target subject object of the speech to be recognized from the multiple subject objects based on the first speech feature and the second speech feature, and sends the target subject object to terminal 400.

[0039] In other embodiments, terminal 400 obtains the speech to be recognized from server 200, determines the first speech feature of the speech to be recognized, and the second speech feature of the registered speech of multiple subject objects when they register their speech, determines the target subject object of the speech to be recognized from the multiple subject objects based on the first speech feature and the second speech feature, and sends the target subject object to server 200.

[0040] In other embodiments, the embodiments of this application can be implemented with the aid of cloud technology, which refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to realize the computation, storage, processing, and sharing of data.

[0041] Cloud technology is a general term encompassing network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form resource pools, allowing for on-demand use with flexibility and convenience. Cloud computing technology will become a crucial support. The backend services of cloud computing systems require substantial computing and storage resources.

[0042] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device 500 based on voice information recognition provided in an embodiment of this application, wherein, Figure 2 The electronic device 500 shown can be Figure 1Server 200 or terminal 400 in the middle, Figure 2 The illustrated electronic device 500 includes at least one processor 410, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.

[0043] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0044] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0045] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0046] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0047] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, such as Bluetooth, WiFi, and Universal Serial Bus.

[0048] In some embodiments, the voice-based information recognition device provided in this application can be implemented in software. Figure 2 A speech-based information recognition device 455 stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: an acquisition module 4551, a first feature similarity determination module 4552, a comparison module 4553, a second feature similarity determination module 4554, and a target subject object determination module 4555. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0049] In other embodiments, the voice-based information recognition device provided in this application can be implemented in hardware. As an example, the voice-based information recognition device provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the voice-based information recognition method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0050] In some embodiments, the terminal or server can implement the voice-based information recognition method provided in this application by running a computer program or computer-executable instructions. For example, the computer program can be a native program in the operating system (e.g., a dedicated voice recognition program) or a software module, such as a voice recognition module that can be embedded in any program (e.g., an instant messaging client, a photo album program, an electronic map client, a navigation client); or it can be a native application (APP), i.e., a program that needs to be installed in the operating system to run. In summary, the above-mentioned computer program can be any form of application, module, or plugin.

[0051] The speech-based information recognition method provided in this application will be described in conjunction with exemplary applications and implementations of the server or terminal provided in the embodiments of this application.

[0052] See Figure 3 , Figure 3This is a flowchart illustrating the speech-based information recognition method provided in this application embodiment, which will be combined with... Figure 3 Steps 101 to 105 are described below. The voice-based information recognition method provided in this application embodiment can be implemented by the server or the terminal alone, or by the server and the terminal working together. The following description will take the implementation by the server alone as an example.

[0053] In step 101, the first speech feature of the speech to be identified and the second speech feature of the registered speech of multiple subjects during speech registration are obtained.

[0054] In some embodiments, the first speech feature of the speech to be recognized can be determined by the following method: acquiring the speech to be recognized, extracting features from the speech to be recognized, and obtaining the first speech feature of the speech to be recognized.

[0055] In some embodiments, the aforementioned plurality of subject objects refers to at least two subject objects, and the second speech feature corresponds one-to-one with the subject object.

[0056] In some embodiments, obtaining the second speech features of the registered speech of multiple subject objects during speech registration can be achieved as follows: Perform the following processing for each subject object: obtain at least one registered speech of the subject object, extract features from each registered speech to obtain the registered speech features of each registered speech, and perform a weighted average of the registered speech features to obtain the second speech features of the registered speech of the subject object during speech registration.

[0057] In some embodiments, the above-mentioned feature extraction of each registered speech to obtain the registered speech features of each registered speech can be achieved by calling a speech coding model to extract features from each registered speech to obtain the registered speech features of each registered speech.

[0058] In some embodiments, the speech to be recognized may be speech emitted by any one of a plurality of subject objects.

[0059] As an example, see Figure 4 , Figure 4 This is a schematic diagram illustrating the principle of the speech-based information recognition method provided in this application embodiment. The method involves acquiring a speech 1 to be recognized, extracting features from the speech 1 to obtain a first speech feature 11. For each subject object, the following processing is performed: acquiring at least one registered speech 2 of the subject object, extracting features from each registered speech 2 to obtain a registered speech feature for each registered speech, and weighting and averaging the registered speech features to obtain a second speech feature 21 of the registered speech of the subject object during speech registration.

[0060] In step 102, the first feature similarity between the first speech feature and each of the second speech features is determined.

[0061] In some embodiments, the first feature similarity between the first speech feature and each of the second speech features may be the cosine similarity between the first speech feature and each of the second speech features. The cosine similarity between the first speech feature and each of the second speech features represents the degree of similarity between the first speech feature and the second speech feature, and the degree of similarity between the first speech feature and the second speech feature is positively correlated with the cosine similarity.

[0062] The expression for the cosine similarity between the first and second speech features can be: (1) in, The cosine similarity between the first and second speech features is used to characterize the speech features. Characterizing the first speech feature, Characterizing the second speech feature, The norm representing the first speech feature, The norm representing the second speech feature. The number of feature elements representing the first and second speech features. Feature elements that characterize the first speech feature Feature elements that characterize the second speech features.

[0063] As an example, see Figure 4 The similarity between the first speech feature 11 and each of the second speech features 21 is determined.

[0064] In this way, by accurately determining the first feature similarity between the first speech feature and each of the second speech features, it is easier to determine the target subject of the speech to be recognized from the subject objects corresponding to each first feature similarity, thus effectively improving the accuracy of the determined target subject of the speech to be recognized.

[0065] In step 103, the similarity of each first feature is compared with a similarity threshold, and based on the comparison results, a first number of second speech features whose first feature similarity is greater than the similarity threshold is determined.

[0066] In some embodiments, step 103 above can be implemented in the following manner: performing the following processing on the second speech features of each subject object respectively: comparing the first feature similarity corresponding to the second speech feature with a similarity threshold to obtain the comparison result corresponding to the second speech feature, the comparison result representing whether the first feature similarity of the second speech feature is greater than the similarity threshold; based on the comparison result, determining the first number of second speech features whose first feature similarity is greater than the similarity threshold.

[0067] In some embodiments, the similarity threshold can be specifically set according to actual needs under different recognition scenarios and recognition accuracy requirements.

[0068] In some embodiments, when the comparison result represents the first feature similarity of the second speech feature greater than the similarity threshold, it represents that the first feature similarity between the first speech feature and the second speech feature is high, that is, the probability that the first speech feature and the second speech feature belong to the same subject is high. In other words, the magnitude of the first feature similarity between the first speech feature and the second speech feature is directly proportional to the probability that the first speech feature and the second speech feature belong to the same subject.

[0069] After step 103 above, the following processing can be performed for different first quantities of second speech features: when the first quantity is one, the subject object corresponding to the second speech feature whose first feature similarity is greater than the similarity threshold is determined as the target subject object of the speech to be recognized; when the first quantity is multiple or zero, and there is no subject object with derived speech features among the multiple subject objects, the subject object corresponding to the second target speech feature is determined as the target subject object of the speech to be recognized; wherein, the second target speech feature is the second speech feature corresponding to the largest first feature similarity.

[0070] In some embodiments, when the first quantity is one, that is, when the number of second speech features with a first feature similarity greater than the similarity threshold is one, it indicates that among the first feature similarities corresponding to each subject object, there is only one subject object whose first feature similarity is greater than the similarity threshold. That is, the subject object corresponding to this second speech feature with a first feature similarity greater than the similarity threshold is most likely to belong to the same subject object as the first speech feature. In other words, the subject object corresponding to the second speech feature with a first feature similarity greater than the similarity threshold can be determined as the target subject object of the speech to be recognized.

[0071] In some embodiments, when the first quantity is multiple or zero, it indicates that among the first feature similarities corresponding to each subject object, at least two or no subject objects have first feature similarities greater than the similarity threshold. When there are at least two subject objects with first feature similarities greater than the similarity threshold, then the target subject object of the speech to be recognized needs to be selected from these at least two subject objects. When there are no target objects with first feature similarities greater than the similarity threshold, then the target subject object of the speech to be recognized needs to be selected from all subject objects.

[0072] In some embodiments, under the premise that the first number is multiple or zero, it can be further determined whether there is a subject object with derived speech features among the multiple subject objects. When there is no subject object with derived speech features among the multiple subject objects, it is impossible to determine the target subject object of the speech to be recognized by means of the derived speech features. In this case, the subject object with the highest first similarity can be selected from the multiple subject objects as the target subject object of the speech to be recognized. That is, the subject object corresponding to the second speech feature with the highest first feature similarity is directly determined as the target subject object of the speech to be recognized. When there is a subject object with derived speech features among the multiple subject objects, the target subject object of the speech to be recognized can be determined by means of the derived speech features, that is, step 104 is performed below.

[0073] In some embodiments, the derived speech features are obtained by updating the historical speech features of the subject object, which are obtained by extracting features from the identified speech of the subject object.

[0074] In some embodiments, derived speech features are speech features evolved and updated from the historical speech features of a subject object. These derived speech features reflect the speech characteristics of the subject object during speech recognition at a historical stage. The historical speech features of the subject object include the first speech features of the subject object during speech recognition at that historical stage. When the subject object has not yet performed speech recognition, it does not have corresponding historical speech features, and therefore does not possess derived speech features. After the subject object performs speech recognition once, the first speech features of the subject object during its first speech recognition become the subject object's historical speech features. At this point, the subject object possesses historical speech features, which are then used to evolve into derived speech features. The generation of derived speech features occurs after each speech recognition operation, when the subject object's historical speech features are updated. Derived speech features are obtained by updating the subject object's historical speech features, which are obtained by extracting features from the subject object's already recognized speech. For example, when a subject undergoes speech recognition for the first time, it has no corresponding historical speech features. After the first speech recognition is completed, the first speech features of the subject evolve into the derived speech features of the subject. When the subject undergoes speech recognition for the Nth time (… At this point, the subject object has corresponding historical speech features (i.e., the first speech features of the speech to be recognized during the previous N-1 speech recognitions). After the subject object completes the Nth speech recognition, the first speech features of the speech to be recognized during the Nth speech recognition are transformed into historical speech features. At this point, the historical speech features of the subject object (i.e., the first speech features of the speech to be recognized during the previous N speech recognitions) are then used to evolve into the derived speech features of the subject object.

[0075] In some embodiments, derived speech features can serve as sub-center templates for the subject object, and second speech features of the subject object can serve as the central template for the subject object. During speech recognition, the first speech features of the speech to be recognized are used to continuously update the corresponding sub-center templates and central templates of the subject object. This ensures that the sub-center templates and central templates retain the features of the speech to be recognized in previous speech recognition processes, making the central templates and sub-center templates more up-to-date and promptly incorporating the latest speech features of the subject object. This provides accuracy assurance for subsequent speech recognition, thereby significantly improving the accuracy of speech recognition.

[0076] As an example, see Figure 4 The multiple subject objects include: subject object 1, subject object 2, subject object 3, and subject object 4. Taking subject object 1 as an example, the derived speech features are explained as follows: The registration speech of subject object 1 at 13:00 on December 12th is received, and feature extraction is performed on the registration speech to obtain the second speech feature corresponding to the registration speech of subject object 1. In response to the speech to be recognized 11 performed by subject object 1 at 14:00 on December 13th, feature extraction is performed on the speech to be recognized 11 to obtain the first speech feature corresponding to the speech to be recognized 11. The target subject object of the speech to be recognized 11 is determined as subject object 1, and the subject object... The first speech feature corresponding to the speech to be recognized 11 is determined as the derived speech feature of subject object 1; in response to the speech to be recognized 12 when subject object 1 performs speech recognition at 15:00 on December 14, feature extraction is performed on the speech to be recognized 12 to obtain the first speech feature corresponding to the speech to be recognized 12, the target subject object of the speech to be recognized 12 is determined as subject object 1, the first speech feature corresponding to the speech to be recognized 11 of subject object 1 is determined as the historical speech feature of subject object 1, and the derived speech feature of subject object 1 is updated based on the historical speech feature of subject object 1 and the first speech feature corresponding to the speech to be recognized 12 of subject object 1.

[0077] Thus, by adopting different methods to determine the target subject of the speech to be recognized for different initial quantities, the most accurate and time-saving method can be adopted to determine the target subject of the speech to be recognized under different initial quantities, thereby effectively improving the accuracy of the determined target subject of the speech to be recognized and effectively saving the computation time of determining the target subject of the speech to be recognized.

[0078] In step 104, when the first quantity is multiple or zero, and there is a subject object with derived speech features among the multiple subject objects, the second feature similarity between the first speech feature and the derived speech feature is determined.

[0079] In some embodiments, derived speech features are obtained by updating the historical speech features of the subject object.

[0080] In some embodiments, when the first quantity is multiple or zero, it indicates that among the first feature similarities corresponding to each subject object, at least two or no subject objects have first feature similarities greater than the similarity threshold. If at least two subject objects have first feature similarities greater than the similarity threshold, then the target subject object for the speech to be recognized needs to be selected from these at least two subject objects. If no target object has first feature similarities greater than the similarity threshold, then the target subject object for the speech to be recognized needs to be selected from all subject objects. Further, when the first quantity is multiple or zero, and among the multiple subject objects there exists a subject object with derived speech features, the second feature similarity between the first speech feature and the derived speech feature is further determined. Based on the second feature similarity and the first feature similarity, the target subject object for the speech to be recognized is further selected from the multiple subject objects.

[0081] As an example, see Figure 5 , Figure 5 This is a schematic diagram illustrating the principle of the speech-based information recognition method provided in the embodiments of this application. Figure 5The diagram shows subject objects 1, 2, 3, 4, and 5, and the second speech feature 1-1 of the registered speech of subject object 1 during speech registration, the second speech feature 2-1 of the registered speech of subject object 2 during speech registration, the second speech feature 3-1 of the registered speech of subject object 3 during speech registration, the second speech feature 4-1 of the registered speech of subject object 4 during speech registration, and the second speech feature 5-1 of the registered speech of subject object 5 during speech registration. Subject object 2 has derived speech feature 2-2, and subject object 3 has derived speech feature 3-2. When the first quantity is multiple or zero, and among the multiple subject objects (subject objects 1 to subject objects 5), there exists a subject object (subject object 2 and subject object 3) with derived speech features, the similarity between the first speech feature and the second feature of the derived speech features (derived speech feature 2-2 and derived speech feature 3-2) is determined.

[0082] In some embodiments, the second feature similarity between the first speech feature and each derived speech feature can be the cosine similarity between the first speech feature and each derived speech feature. The cosine similarity between the first speech feature and each derived speech feature represents the degree of similarity between the first speech feature and the derived speech feature, and the degree of similarity between the first speech feature and the derived speech feature is positively correlated with the cosine similarity.

[0083] In some embodiments, the expression for the cosine similarity between the first speech feature and the derived speech feature can be: (2) in, The cosine similarity between the first speech feature and the derived speech feature is used to represent the speech feature. Characterizing the first speech feature, Characterize derived speech features, The norm representing the first speech feature, The norm representing the features of derived speech. The number of feature elements representing the first speech feature and the derived speech feature. Feature elements that characterize the first speech feature Feature elements that characterize the features of derived speech.

[0084] As an example, see Figure 4 The similarity between the first speech feature and the second features of derived speech feature 2-2 and derived speech feature 3-2 is determined respectively.

[0085] Thus, by adopting different methods to determine the target subject of the speech to be recognized for different initial quantities, the most accurate and time-saving method can be adopted to determine the target subject of the speech to be recognized under different initial quantities, thereby effectively improving the accuracy of the determined target subject of the speech to be recognized and effectively saving the computation time of determining the target subject of the speech to be recognized.

[0086] In step 105, based on the second feature similarity and the first feature similarity, the target subject of the speech to be recognized is determined from multiple subject objects.

[0087] As an example, see Figure 5 Based on the second feature similarity and the first feature similarity, the target subject of the speech to be recognized is determined from subject 1, subject 2, subject 3, subject 4 and subject 5.

[0088] In some embodiments, see Figure 6 , Figure 6 This is a flowchart illustrating the speech-based information recognition method provided in an embodiment of this application. Figure 6 Step 105 shown can be implemented by steps 1051 to 1052.

[0089] In step 1051, the maximum value of the second feature similarity and the first feature similarity is determined.

[0090] In some embodiments, step 1051 above can be implemented as follows: sort the first feature similarity between the first speech feature and each second speech feature, and the second feature similarity between the first speech feature and each derived speech feature in descending order, to obtain the maximum similarity among the first feature similarity between the first speech feature and each second speech feature, and the second feature similarity between the first speech feature and each derived speech feature.

[0091] In step 1052, the subject object corresponding to the maximum value is determined as the target subject object of the speech to be recognized.

[0092] In some embodiments, step 1052 above can be implemented in the following way: the subject object corresponding to the derived speech feature or the second speech feature corresponding to the maximum value is determined as the target subject object of the speech to be recognized.

[0093] As an example, see Figure 5 When the similarity between the first speech feature of the speech to be recognized and the second feature of the derived speech feature 2-2 of the subject object 2 is the maximum value between the second feature similarity and the first feature similarity, the subject object 2 is determined as the target subject object of the speech to be recognized.

[0094] In some embodiments, see Figure 7 , Figure 7 This is a flowchart illustrating the speech-based information recognition method provided in an embodiment of this application. Figure 7 After step 105 shown, the speech features of the target subject can be updated through steps 106 to 108.

[0095] In step 106, the derived information of the target subject object is obtained. The derived information indicates whether the target subject object has derived speech features.

[0096] In some embodiments, the derived information of the subject object characterizes whether the target subject object has derived speech features.

[0097] As an example, see Figure 5 The derived information of subject object 1 indicates that subject object 1 does not have derived speech features. The derived information of subject object 2 indicates that subject object 2 has derived speech features. The derived information of subject object 5 indicates that subject object 5 does not have derived speech features.

[0098] In step 107, when the derived information characterizes the target subject object without derived speech features, the second speech feature of the target subject object is updated, and the derived speech feature of the target subject object is generated.

[0099] As an example, see Figure 5 When the target subject is subject 4, and the derived information of subject 4 indicates that subject 4 does not have derived speech features, the second speech feature of subject 4 is updated, and the derived speech feature of subject 4 is generated.

[0100] In some embodiments, see Figure 8 , Figure 8 This is a flowchart illustrating the speech-based information recognition method provided in an embodiment of this application. Figure 8 Step 107 shown can be achieved through steps 1071 to 1072.

[0101] In step 1071, when the derived information characterizes the target subject object without derived speech features, the second speech feature of the target subject object is updated based on the first speech feature to obtain the updated second speech feature of the target subject object.

[0102] In some embodiments, when the derived information characterizes the target subject object without derived speech features, that is, when the target subject object does not have a sub-center template but only a center template, the second speech feature of the target subject object is updated based on the first speech feature of the speech to be identified, that is, the center template of the target subject object is updated, and the updated second speech feature of the target subject object (i.e., the updated center template) is obtained.

[0103] As an example, see Figure 5 When the target subject is subject 5, the derived information of subject 5 indicates that subject 5 does not have derived speech features. Based on the first speech feature 5-3 of the speech to be identified, the second speech feature 5-1 of subject 5 is updated to obtain the updated second speech feature 5-2 of subject 5.

[0104] In some embodiments, step 1071 above can be implemented as follows: obtain the first weight of the first speech feature and the second weight of the second speech feature of the target subject, wherein the sum of the first weight and the second weight is equal to 1; and perform a weighted summation of the first speech feature and the second speech feature of the target subject according to the first weight and the second weight, respectively, to obtain the updated second speech feature of the target subject.

[0105] As an example, the expression for the updated second speech feature of the target subject object can be: (3) in, The updated second speech features representing the target subject object Characterizes the first weight, Characterizing the second weight, Characterizing the first speech feature, The second speech feature that represents the target subject. .

[0106] In step 1072, the first speech feature is determined as a derived speech feature of the target subject.

[0107] As an example, when the derived information characterizes the target subject object without derived speech features, that is, when the target subject object does not have a sub-center template but only a center template, the first speech feature of the speech to be identified is determined as the derived speech feature of the target subject object (that is, the first speech feature of the speech to be identified is determined as the sub-center template of the target subject object).

[0108] As an example, see Figure 5 The first speech feature 5-3 is determined as the derived speech feature of the subject object 5.

[0109] Thus, when the derived information characterizes the target subject object without derived speech features, the second speech feature of the target subject object is updated based on the first speech feature to obtain the updated second speech feature of the target subject object. The first speech feature is then identified as the derived speech feature of the target subject object. This achieves the updating of the second speech feature and the generation of the derived speech feature of the target subject object when the target subject object does not have derived speech features. This facilitates subsequent speech recognition by using the updated second speech feature and the derived speech feature, significantly improving the accuracy of speech recognition.

[0110] In step 108, when the derived information indicates that the target subject object has derived speech features, the second speech feature of the target subject object and the derived speech feature of the target subject object are updated.

[0111] As an example, see Figure 5 When the target subject is subject 3, subject 3 has derived speech feature 3-2, then update the second speech feature 3-1 of subject 3 and the derived speech feature 3-2 of subject 3.

[0112] In some embodiments, see Figure 9 , Figure 9 This is a flowchart illustrating the speech-based information recognition method provided in an embodiment of this application. Figure 9 Step 108 shown can be implemented by steps 1081 to 1082.

[0113] In step 1081, when the derived information characterizes the target subject object as having derived speech features, the derived speech features of the target subject object are updated based on the first speech features to obtain the updated derived speech features.

[0114] As an example, see Figure 5 When the target subject is subject 3, the derived information of subject 3 represents that subject 3 has derived speech features. Based on the first speech feature of the speech to be identified, the derived speech feature 3-2 of subject 3 is updated to obtain the updated derived speech feature 3-3.

[0115] In some embodiments, step 1081 above can be implemented as follows: obtaining a second number of derived speech features of the target subject object; when the second number is one, determining the derived speech features of the target subject object as speech features to be updated, and updating the speech features to be updated based on the first speech features to obtain the updated derived speech features; when the second number is multiple, selecting target derived speech features from the derived speech features of the target subject object based on the second number, and updating the target derived speech features based on the first speech features to obtain the updated derived speech features.

[0116] In some embodiments, when the second number of derived speech features of the target subject is one, then this one derived speech feature needs to be updated. When the second number of derived speech features of the target subject is multiple, then at least a portion of the derived speech features of the target subject need to be selected and the selected derived speech features updated.

[0117] As an example, see Figure 5 When the target subject is subject 3, the derived information of subject 3 indicates that subject 3 has derived speech features, and the second number of derived speech features of subject 3 is one. The derived speech feature 3-2 of subject 3 is determined as the speech feature to be updated, and the speech feature to be updated is updated based on the first speech feature of the speech to be recognized, so as to obtain the updated derived speech feature 3-3.

[0118] In some embodiments, when there are multiple second quantities, selecting target derived speech features from the derived speech features of the target subject object based on the second quantity can be achieved as follows: obtaining the third feature similarity between the first speech feature and each derived speech feature of the target subject object; when the second quantity is less than a quantity threshold, determining the derived speech feature corresponding to the largest third feature similarity as the target derived speech feature; when the second quantity is greater than or equal to the quantity threshold, performing the following processing on each derived speech feature of the target subject object: determining the derived speech feature as a reference speech feature, and determining the fourth feature similarity between the reference speech feature and each other speech feature, and determining the target derived speech feature based on the third feature similarity and the fourth feature similarity; wherein, the other speech features are the derived speech features of the target subject object other than the reference speech feature.

[0119] In some embodiments, the determination of target derived speech features based on third feature similarity and fourth feature similarity can be achieved as follows: when the third feature similarity is greater than the fourth feature similarity, the derived speech feature corresponding to the largest third feature similarity is determined as the speech feature to be updated; when the third feature similarity is less than the fourth feature similarity, the derived speech features of the target subject are sorted in descending order of third feature similarity to obtain a sorted queue, and the first and second derived speech features starting from the head of the sorted queue are merged to obtain a merged speech feature; the merged speech feature is determined as the speech feature to be updated.

[0120] In some embodiments, the third feature similarity between the first speech feature and the derived speech feature of the target subject can be the cosine similarity between the first speech feature and the derived speech feature of the target subject. The cosine similarity between the first speech feature and the derived speech feature of the target subject represents the degree of similarity between the first speech feature and the derived speech feature, and the degree of similarity between the first speech feature and the derived speech feature is positively correlated with the cosine similarity.

[0121] In some embodiments, the expression for the similarity between the first speech feature and the third feature of the derived speech feature of the target subject can be: (4) in, The cosine similarity represents the relationship between the first speech feature and the derived speech features of the target subject. Characterizing the first speech feature, Derived speech features that characterize the target subject. The norm representing the first speech feature, The norm representing the derived speech features of the target subject object. The number of feature elements representing the primary speech features and the derived speech features of the target subject. Feature elements that characterize the first speech feature Feature elements that characterize the derived speech features of the target subject.

[0122] In some embodiments, the expression for the fourth feature similarity between the above-mentioned reference speech features and other speech features can be: (5) in, The cosine similarity (fourth feature similarity) between the reference speech features and other speech features is used to represent the similarity between the reference speech features and other speech features. Characterize the reference speech features, Characterizing other speech features, The norm that characterizes the features of the reference speech. Norms that characterize other speech features The number of feature elements representing the reference speech features and other speech features. Feature elements that represent the features of the reference speech. Feature elements that characterize other speech features.

[0123] Thus, when the derived information represents that the target subject has derived speech features, the derived speech features of the target subject are updated based on the first speech feature to obtain the updated derived speech features. Based on the updated derived speech features, the second speech features of the target subject are updated to obtain the updated second speech features of the target subject. This realizes the updating of the derived speech features and the second speech features of the target subject when the target subject has derived speech features. This facilitates subsequent recognition in the speech recognition process by using the updated second speech features and derived speech features, significantly improving the accuracy of speech recognition.

[0124] In some embodiments, the above-mentioned updating of the speech features to be updated based on the first speech features to obtain the updated derived speech features can be achieved in the following way: obtaining the third weight of the first speech features and the fourth weight of the speech features to be updated, wherein the sum of the third weight and the fourth weight is 1; and weighting the first speech features and the speech features to be updated according to the third weight and the fourth weight respectively to obtain the updated derived speech features of the target subject.

[0125] In some embodiments, the expression for the updated derived speech features of the target subject object can be: (6) in, The updated derived speech features representing the target subject object Characterizing the third weight, Characterizing the fourth weight, Characterizing the first speech feature, Characterize the speech features to be updated. .

[0126] In step 1082, the second speech features of the target subject are updated based on the updated derived speech features to obtain the updated second speech features of the target subject.

[0127] In some embodiments, step 1082 can be implemented as follows: obtaining the fifth weight of the updated derived speech feature, the sixth weight of the second speech feature of the target subject, and the seventh weight of the first speech feature; wherein the sum of the fifth weight, the sixth weight, and the seventh weight equals 1; and weighting the updated derived speech feature, the second speech feature of the target subject, and the first speech feature according to the fifth weight, the sixth weight, and the seventh weight, respectively, to obtain the updated second speech feature of the target subject.

[0128] In some embodiments, the expression for the updated second speech feature of the target subject object can be: (7) in, The updated second speech features representing the target subject object Characterizing the fifth weight, Characterizing the sixth weight, Characterizing the seventh weight, The derived speech features are represented after the update. The second speech feature that represents the target subject. Characterizing the first speech feature, .

[0129] Thus, by using the first feature similarity between the first speech feature of the speech to be identified and the second speech feature of each subject object during speech registration, and when the number of second speech features with a first feature similarity greater than a similarity threshold is multiple or zero, and there is a subject object with derived speech features among the multiple subject objects, the second feature similarity between the first speech feature of the speech to be identified and the derived speech feature is determined. Based on the first feature similarity and the second feature similarity, the target subject object of the speech to be identified is determined from the multiple subject objects. In this way, by comparing the first speech feature of the speech to be identified with the second speech feature of multiple subject objects during speech registration, as well as the derived speech feature, the first feature similarity and the second feature similarity are obtained. The target subject object of the speech to be identified is then identified using the first feature similarity and the second feature similarity. Since the derived speech features of the subject are obtained by updating the historical speech features of the subject, the derived speech features of the subject continuously absorb the speech characteristics of the historical speech features during the update process. This allows them to more accurately reflect the changes in the speech of the subject at different times. As a result, when identifying the target subject of the speech to be identified, the similarity between the first speech feature of the speech to be identified and the second feature of the derived speech feature is comprehensively considered, which significantly improves the accuracy of speech recognition.

[0130] The following will describe an exemplary application of the embodiments of this application in a real-world speech recognition scenario.

[0131] Speech recognition is an interdisciplinary field. Over the past two decades, speech recognition technology has made significant progress, moving from the laboratory to the market. It will penetrate various fields, including industry, home appliances, communications, automotive electronics, healthcare, home services, and consumer electronics. The areas involved in speech recognition include signal processing, pattern recognition, probability theory and information theory, vocalization and auditory mechanisms, and artificial intelligence, among others.

[0132] With the development of machine learning and deep learning, biometric technologies based on embedding vector inner product methods for similarity comparison have been increasingly applied, such as face recognition, voiceprint recognition, and gait recognition. Speech signals are one of the important carriers of identity information. Compared with other biometric features such as faces and fingerprints, each person's speech signal is unique due to the influence of the vocal organs and environment. Furthermore, speech signals are proactive, easily accepted, and low-cost; therefore, speaker verification technology is also an important automatic identity verification technology. In deep learning-based speaker verification systems, such as... Figure 5 As shown, registered audio is typically processed by a model trained on a large amount of audio data to extract fixed-dimensional vectors, which are called the speech features of the registered audio. The speech features of registered audio are usually discriminative and capable of representing the speaker's identity information.

[0133] While speaker verification technology is a crucial identity verification method, the characteristics of the speech signal also affect its performance. A person's voice changes with environment and time; for example, the timbre changes throughout the day, such as upon waking and after sleep, and the voice signal may "fatigue" between morning and evening, causing changes in timbre. During registration, audio is often read clearly and spoken precisely, while test audio is more casual, with less fixed time and location. Furthermore, over time, due to the time-varying characteristics of the speech signal, the gap between registration and verification audio widens, leading to a decline in the performance of the speaker verification system and a worse user experience. This is one reason why speaker verification systems cannot be widely used. Some studies have used attention-based mechanisms to assign weights to large amounts of historical test data to select the optimal template; however, in practical systems, due to information security and user privacy concerns, service providers cannot save or upload historical data.

[0134] The speech-based information recognition method provided in this application uses multiple sub-centers as templates and achieves the characteristic that the speaker confirmation system becomes more accurate with use by updating the templates in real time.

[0135] The speech-based information recognition method provided in this application updates the registered speech of the speaker in real time and splits into multiple sub-center templates (i.e., the derived speech features described above) as the test progresses. This makes the registered speech template (center template) increasingly accurate with user use, improving interaction quality. The advantages of this application mainly include: one-time user registration, with real-time template updates based on test speech; speaker verification using the mean center template (i.e., the second speech feature described above) and sub-center templates; real-time updates of the center template and sub-center templates during test speech; no additional calculation modules are introduced, and no historical data is required.

[0136] This application embodiment is divided into a registration and verification stage and a template update stage. The first step: In the registration stage, the speech features of all registered speech samples are weighted and averaged to obtain the mean center speech features as the initial speech features. The second step: In the verification stage, the verification speech samples are processed... The cosine similarity score is calculated between the mean center template and all sub-center templates, and a scoring threshold is set. Speech instances exceeding the threshold are considered to be from the same speaker, and the template is then updated. In the third step, during the template update phase, the score result is categorized into the following cases; except for the following cases, the template is not updated: When there are no sub-centers initially, the score is calculated based on the given sub-center threshold (…). When the score is higher than the decision threshold ( and lower than At that time, it was assumed that the speech could generate a subcenter. The test speech is used as the sub-center template embedding and saved, then the mean center template is... It will be updated according to formula (1). When there are subcenters, but the upper limit of the number of subcenter classes has not been reached: when the score of the test speech is higher than and arbitrary subcenter If the test speech data belongs to the highest-scoring sub-center template, then the highest-scoring sub-center template is updated first. (8) Then update the central template: (9) This ensures that the distance between the mean center template and the sub-center templates does not increase, and also improves the generalization by increasing the coverage of the registration templates through the addition of sub-center templates. When the test speech score is higher than But below the threshold for any sub-center template If a new subcenter is generated, then a new subcenter is generated. When there are subcenters, but the upper limit for the number of subcenter classes is reached: since subcenter templates cannot be added indefinitely, an upper limit for the number of subcenter template classes needs to be set. When the upper limit for subcenter template classes is reached, the cosine similarity matrix between all subcenter templates and the test speech needs to be calculated. If the similarity between the test speech and all subcenter templates is greater than the similarity between any two templates, then the test speech is assigned to the subcenter with the highest score, and the steps for when there are no subcenters are performed; if the similarity between the test speech and all subcenter templates is less than the similarity between any two templates, then the two subcenter templates with the highest similarity are merged into one subcenter, and the test speech is used as the new subcenter template, and the steps for when there are no subcenters are performed again.

[0137] In some embodiments, see Figure 10 , Figure 10 This is a schematic diagram illustrating the effect of the speech-based information recognition method provided in this application embodiment. The horizontal axis represents the number of days, and the vertical axis represents the cosine similarity. Method 7 indicates no template update operation; Method 1 indicates template update for all test speech; Method 2 indicates template update only for speech exceeding a given threshold. The template updates are categorized into several methods: Method 1, Method 2, and Method 3. Method 3 represents the multi-center template update method proposed in this application embodiment, but the number of sub-centers is fixed at 4 based on the test time. Method 4 represents the multi-center template update method we propose, which divides the number of sub-centers according to a threshold, with an upper limit of 5 categories. Method 5 is the same as Method 4, with an upper limit of 10 categories. Method 6 is the same as Method 4, with an upper limit of 20 categories. It can be observed that without template updates, the model performance gradually decreases and the score gets lower as the interval between test and registration voices lengthens. When conventional template update methods (Methods 1 and 2) are used, the performance remains stable and does not gradually deteriorate over time, but the trend is relatively flat and does not gradually improve with continuous user use. However, after adopting the multi-center template update method proposed in this application embodiment, there is a 10% performance improvement compared to conventional methods, and it can be observed that as user interaction becomes more frequent over time, the similarity gradually increases, thus achieving the characteristic that the template gets better with use. Furthermore, by setting an upper limit on the number of classes, the optimal upper limit is 10 classes. After comparing methods 5 and 6, it was found that the computational workload doubled when the number of classes exceeded 10, but the effect was not significant.

[0138] The first feature similarity between the first speech feature of the speech to be identified and the second speech feature of each subject object during speech registration is determined. If the number of second speech features with a first feature similarity greater than a similarity threshold is multiple or zero, and among the multiple subject objects, there exists a subject object with derived speech features, the second feature similarity between the first speech feature of the speech to be identified and the derived speech feature is determined. Based on the first feature similarity and the second feature similarity, the target subject object of the speech to be identified is determined from among the multiple subject objects. Thus, by comparing the first speech feature of the speech to be identified with the second speech feature of multiple subject objects during speech registration, as well as the derived speech feature, the first feature similarity and the second feature similarity are obtained. The target subject object of the speech to be identified is then identified using the first feature similarity and the second feature similarity. Since the derived speech features of the subject are obtained by updating the historical speech features of the subject, the derived speech features of the subject continuously absorb the speech characteristics of the historical speech features during the update process. This allows them to more accurately reflect the changes in the speech of the subject at different times. As a result, when identifying the target subject of the speech to be identified, the similarity between the first speech feature of the speech to be identified and the second feature of the derived speech feature is comprehensively considered, which significantly improves the accuracy of speech recognition.

[0139] It is understood that in the embodiments of this application, data related to the first voice feature and the second voice feature are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0140] The following description continues to illustrate the exemplary structure of the voice-based information recognition device 455 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2As shown, the software modules stored in the speech-based information recognition device 455 in the memory 450 may include: an acquisition module 4551, used to acquire a first speech feature of the speech to be recognized, and a second speech feature of the registered speech of multiple subject objects during speech registration; a first feature similarity determination module 4552, used to determine the first feature similarity between the first speech feature and each second speech feature respectively; a comparison module 4553, used to compare each first feature similarity with a similarity threshold, and based on the comparison result, determine a first number of second speech features whose first feature similarity is greater than the similarity threshold; a second feature similarity determination module 4554, used to determine the second feature similarity between the first speech feature and the derived speech feature when the first number is multiple or zero, and there is a subject object with derived speech features among the multiple subject objects; wherein, the derived speech feature is obtained by updating the historical speech features of the subject object; and a target subject object determination module 4555, used to determine the target subject object of the speech to be recognized from the multiple subject objects based on the second feature similarity and the first feature similarity.

[0141] In some embodiments, the above-described speech-based information recognition device 455 further includes: a supplementary determination module, configured to determine the subject object corresponding to the second speech feature whose first feature similarity is greater than a similarity threshold as the target subject object of the speech to be recognized when the first quantity is one; and to determine the subject object corresponding to the second target speech feature as the target subject object of the speech to be recognized when the first quantity is multiple or zero, and there is no subject object with derived speech features among the multiple subject objects; wherein the second target speech feature is the second speech feature corresponding to the largest first feature similarity.

[0142] In some embodiments, the target subject object determination module 4555 is further configured to determine the maximum value between the second feature similarity and the first feature similarity; and determine the subject object corresponding to the maximum value as the target subject object of the speech to be recognized.

[0143] In some embodiments, the above-described speech-based information recognition device 455 further includes: an update module, configured to acquire derived information of a target subject object, wherein the derived information indicates whether the target subject object has derived speech features; when the derived information indicates that the target subject object does not have derived speech features, update the second speech features of the target subject object and generate derived speech features of the target subject object; when the derived information indicates that the target subject object has derived speech features, update the second speech features of the target subject object and the derived speech features of the target subject object.

[0144] In some embodiments, the above-mentioned updating module is further configured to update the second speech feature of the target subject object based on the first speech feature when the derived information characterizes the target subject object as not having derived speech features, thereby obtaining the updated second speech feature of the target subject object; and to determine the first speech feature as the derived speech feature of the target subject object.

[0145] In some embodiments, the update module is further configured to obtain a first weight of the first speech feature and a second weight of the second speech feature of the target subject, wherein the sum of the first weight and the second weight is equal to 1; and to perform a weighted summation of the first speech feature and the second speech feature of the target subject according to the first weight and the second weight, respectively, to obtain the updated second speech feature of the target subject.

[0146] In some embodiments, the above-mentioned updating module is further configured to update the derived speech features of the target subject object based on the first speech features when the derived information represents that the target subject object has derived speech features, to obtain the updated derived speech features; and update the second speech features of the target subject object based on the updated derived speech features, to obtain the updated second speech features of the target subject object.

[0147] In some embodiments, the update module is further configured to obtain a second number of derived speech features of the target subject object; when the second number is one, the derived speech features of the target subject object are determined as speech features to be updated, and the speech features to be updated are updated based on the first speech features to obtain the updated derived speech features; when the second number is multiple, a target derived speech feature is selected from the derived speech features of the target subject object based on the second number, and the target derived speech feature is updated based on the first speech feature to obtain the updated derived speech features.

[0148] In some embodiments, the update module is further configured to obtain the third feature similarity between the first speech feature and each derived speech feature of the target subject object; when the second number is less than the number threshold, the derived speech feature corresponding to the largest third feature similarity is determined as the target derived speech feature; when the second number is greater than or equal to the number threshold, the following processing is performed on each derived speech feature of the target subject object: the derived speech feature is determined as the reference speech feature, and the fourth feature similarity between the reference speech feature and each other speech feature is determined, and the target derived speech feature is determined based on the third feature similarity and the fourth feature similarity; wherein, the other speech features are the derived speech features of the target subject object other than the reference speech feature.

[0149] In some embodiments, the above-mentioned updating module is further configured to obtain the third weight of the first speech feature and the fourth weight of the speech feature to be updated, wherein the sum of the third weight and the fourth weight is 1; and to perform weighted summation of the first speech feature and the speech feature to be updated according to the third weight and the fourth weight, respectively, to obtain the updated derived speech feature of the target subject object.

[0150] In some embodiments, the update module is further configured to obtain the fifth weight of the updated derived speech feature, the sixth weight of the second speech feature of the target subject, and the seventh weight of the first speech feature; wherein the sum of the fifth weight, the sixth weight, and the seventh weight is equal to 1; the updated derived speech feature, the second speech feature of the target subject, and the first speech feature are weighted and summed according to the fifth weight, the sixth weight, and the seventh weight, respectively, to obtain the updated second speech feature of the target subject.

[0151] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the speech-based information recognition method described above in this application.

[0152] This application provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by a processor, they cause the processor to execute the speech-based information recognition method provided in this application. For example... Figure 3 The speech-based information recognition method is shown.

[0153] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of electronic devices including one or any combination of the above-mentioned memories.

[0154] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0155] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0156] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0157] In summary, the embodiments of this application have the following beneficial effects: (1) By comparing the first speech feature of the speech to be identified with the second speech feature of each subject object during speech registration, and when the first feature similarity is greater than the similarity threshold for the first number of second speech features, and there are multiple subjects with derived speech features among the multiple subjects, the second feature similarity between the first speech feature of the speech to be identified and the derived speech feature is determined. Based on the first feature similarity and the second feature similarity, the target subject object of the speech to be identified is determined from the multiple subjects. Thus, by comparing the first speech feature of the speech to be identified with the second speech feature of multiple subjects during speech registration, as well as the derived speech feature, the first feature similarity and the second feature similarity are obtained. The target subject object of the speech to be identified is identified by comparing the first speech feature of the speech to be identified with the second speech feature of multiple subjects during speech registration, as well as the derived speech feature, the target subject object of the speech to be identified is identified. Since the derived speech features of the subject are obtained by updating the historical speech features of the subject, the derived speech features of the subject continuously absorb the speech characteristics of the historical speech features during the update process. This allows them to more accurately reflect the changes in the speech of the subject at different times. As a result, when identifying the target subject of the speech to be identified, the similarity between the first speech feature of the speech to be identified and the second feature of the derived speech feature is comprehensively considered, which significantly improves the accuracy of speech recognition.

[0158] (2) By accurately determining the first feature similarity between the first speech feature and each second speech feature, it is easier to determine the target subject of the speech to be recognized from the subject objects corresponding to each first feature similarity, thus effectively improving the accuracy of the target subject of the recognized speech.

[0159] (3) By adopting different methods to determine the target subject of the speech to be recognized for different first quantities, the most accurate and time-saving method can be adopted to determine the target subject of the speech to be recognized under different first quantities, thereby effectively improving the accuracy of the target subject of the speech to be recognized and effectively saving the computation time of determining the target subject of the speech to be recognized.

[0160] (4) When the derived information characterizes the target subject object without derived speech features, the second speech feature of the target subject object is updated based on the first speech feature to obtain the updated second speech feature of the target subject object, and the first speech feature is determined as the derived speech feature of the target subject object. This realizes the updating of the second speech feature of the target subject object and the generation of the derived speech feature when the target subject object does not have the derived speech feature. This facilitates the recognition of speech by the updated second speech feature and the derived speech feature in the subsequent speech recognition process, which significantly improves the accuracy of speech recognition.

[0161] (5) When the derived information represents that the target subject has derived speech features, the derived speech features of the target subject are updated based on the first speech features to obtain the updated derived speech features. Based on the updated derived speech features, the second speech features of the target subject are updated to obtain the updated second speech features of the target subject. This realizes the updating of the derived speech features and the second speech features of the target subject when the target subject has derived speech features. This facilitates the recognition of speech by the updated second speech features and derived speech features in the subsequent speech recognition process, which significantly improves the accuracy of speech recognition.

[0162] (6) Derived speech features can be used as sub-center templates of the subject object, and the second speech features of the subject object can be used as the center template of the subject object. In the process of speech recognition, the first speech features of the speech to be recognized are used to continuously update the corresponding sub-center templates and center templates of the subject object, thereby ensuring that the sub-center templates and center templates can retain the features of the speech to be recognized in the previous speech recognition when performing speech recognition next time. This allows the center templates and sub-center templates to keep up with the times and absorb the latest speech features of the subject object in a timely manner, providing accuracy assurance for subsequent speech recognition, thereby significantly improving the accuracy of speech recognition.

[0163] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A speech-based information recognition method, characterized in that, The method includes: The first speech feature of the speech to be recognized and the second speech feature of the registered speech of multiple subjects during speech registration are obtained. Determine the first feature similarity between the first speech feature and each of the second speech features; The similarity of each of the first features is compared with a similarity threshold, and based on the comparison results, a first number of second speech features whose first feature similarity is greater than the similarity threshold is determined; When the first quantity is multiple or zero, and there is a subject object with derived speech features among the multiple subject objects, the second feature similarity between the first speech feature and the derived speech feature is determined; The derived speech features are obtained by updating the historical speech features of the subject object; Based on the second feature similarity and the first feature similarity, the target subject of the speech to be recognized is determined from the plurality of subject objects; Obtain the derived information of the target subject object, wherein the derived information characterizes whether the target subject object has the derived speech features; When the derived information indicates that the target subject does not have the derived speech feature, the second speech feature of the target subject is updated based on the first speech feature to obtain the updated second speech feature of the target subject. The first speech feature is determined as a derived speech feature of the target subject.

2. The method according to claim 1, characterized in that, After determining the first number of second speech features whose first feature similarity is greater than the similarity threshold based on the comparison results, the method further includes: When the first quantity is one, the subject object corresponding to the second speech feature whose first feature similarity is greater than the similarity threshold is determined as the target subject object of the speech to be recognized; When the first quantity is multiple or zero, and there is no subject object with the derived speech feature among the multiple subject objects, the subject object corresponding to the second target speech feature is determined as the target subject object of the speech to be identified; wherein, the second target speech feature is the second speech feature corresponding to the largest first feature similarity.

3. The method according to claim 1, characterized in that, The step of determining the target subject of the speech to be recognized from the plurality of subject objects based on the second feature similarity and the first feature similarity includes: Determine the maximum value between the second feature similarity and the first feature similarity; The subject object corresponding to the maximum value is determined as the target subject object of the speech to be recognized.

4. The method according to claim 1, characterized in that, After obtaining the derived information of the target subject object, the method further includes: When the derived information indicates that the target subject object has the derived speech features, the second speech features of the target subject object and the derived speech features of the target subject object are updated.

5. The method according to claim 1, characterized in that, The step of updating the second speech features of the target subject object based on the first speech features to obtain the updated second speech features of the target subject object includes: Obtain the first weight of the first speech feature and the second weight of the second speech feature of the target subject object, wherein the sum of the first weight and the second weight is equal to 1; The first speech feature and the second speech feature of the target subject are weighted and summed according to the first weight and the second weight, respectively, to obtain the updated second speech feature of the target subject.

6. The method according to claim 4, characterized in that, When the derived information indicates that the target subject object has the derived speech features, updating the second speech features of the target subject object and the derived speech features of the target subject object includes: When the derived information indicates that the target subject object has the derived speech features, the derived speech features of the target subject object are updated based on the first speech features to obtain the updated derived speech features; Based on the updated derived speech features, the second speech features of the target subject are updated to obtain the updated second speech features of the target subject.

7. The method according to claim 6, characterized in that, The step of updating the derived speech features of the target subject based on the first speech feature to obtain the updated derived speech features includes: Obtain a second number of the derived speech features of the target subject object; When the second quantity is one, the derived speech features of the target subject are determined as the speech features to be updated, and the speech features to be updated are updated based on the first speech features to obtain the updated derived speech features; When there are multiple second quantities, based on the second quantity, target derived speech features are selected from the derived speech features of the target subject object, and based on the first speech features, the target derived speech features are updated to obtain the updated derived speech features.

8. The method according to claim 7, characterized in that, When the second quantity is multiple, based on the second quantity, selecting target derived speech features from the derived speech features of the target subject object includes: Obtain the third feature similarity between the first speech feature and each derived speech feature of the target subject object; When the second quantity is less than the quantity threshold, the derived speech feature corresponding to the largest third feature similarity is determined as the target derived speech feature; When the second quantity is greater than or equal to the quantity threshold, the following processing is performed on each derived speech feature of the target subject object: the derived speech feature is determined as a reference speech feature, and the fourth feature similarity between the reference speech feature and each other speech feature is determined. Based on the third feature similarity and the fourth feature similarity, the target derived speech feature is determined; wherein, the other speech features are the derived speech features of the target subject object other than the reference speech feature.

9. The method according to claim 7, characterized in that, The step of updating the speech features to be updated based on the first speech features to obtain the updated derived speech features includes: Obtain the third weight of the first speech feature and the fourth weight of the speech feature to be updated, wherein the sum of the third weight and the fourth weight is 1; The first speech feature and the speech feature to be updated are weighted and summed according to the third weight and the fourth weight, respectively, to obtain the updated derived speech feature of the target subject.

10. The method according to claim 6, characterized in that, The step of updating the second speech features of the target subject object based on the updated derived speech features to obtain the updated second speech features of the target subject object includes: Obtain the fifth weight of the updated derived speech feature, the sixth weight of the second speech feature of the target subject, and the seventh weight of the first speech feature; Wherein, the sum of the fifth weight, the sixth weight, and the seventh weight is equal to 1; The updated derived speech features, the second speech features of the target subject, and the first speech features are weighted and summed according to the fifth weight, the sixth weight, and the seventh weight, respectively, to obtain the updated second speech features of the target subject.

11. A voice-based information recognition device, characterized in that, The device includes: The acquisition module is used to acquire the first speech features of the speech to be recognized, and the second speech features of the registered speech of multiple subject objects when they are registering their speech. The first feature similarity determination module is used to determine the first feature similarity between the first speech feature and each of the second speech features respectively; The comparison module is used to compare the similarity of each of the first features with a similarity threshold, and based on the comparison result, determine a first number of second speech features whose first feature similarity is greater than the similarity threshold; The second feature similarity determination module is used to determine the second feature similarity between the first speech feature and the derived speech feature when the first quantity is multiple or zero, and there is a subject object with derived speech features among the multiple subject objects; wherein, the derived speech feature is obtained by updating the historical speech features of the subject object; The target subject object determination module is used to determine the target subject object of the speech to be recognized from the plurality of subject objects based on the second feature similarity and the first feature similarity; An update module is used to obtain derived information of the target subject object, wherein the derived information indicates whether the target subject object has the derived speech feature; when the derived information indicates that the target subject object does not have the derived speech feature, the second speech feature of the target subject object is updated based on the first speech feature to obtain the updated second speech feature of the target subject object; and the first speech feature is determined as the derived speech feature of the target subject object.

12. The apparatus according to claim 11, characterized in that, The device further includes: The supplementary determination module is used to determine the subject object corresponding to the second speech feature whose first feature similarity is greater than the similarity threshold as the target subject object of the speech to be recognized when the first number is one, after determining the first number of the second speech feature whose first feature similarity is greater than the similarity threshold based on the comparison result. When the first quantity is multiple or zero, and there is no subject object with the derived speech feature among the multiple subject objects, the subject object corresponding to the second target speech feature is determined as the target subject object of the speech to be identified; wherein, the second target speech feature is the second speech feature corresponding to the largest first feature similarity.

13. The apparatus according to claim 11, characterized in that, The target subject object determination module is also used for: Determine the maximum value between the second feature similarity and the first feature similarity; The subject object corresponding to the maximum value is determined as the target subject object of the speech to be recognized.

14. The apparatus according to claim 11, characterized in that, The update module is also used for: After obtaining the derived information of the target subject object, when the derived information indicates that the target subject object has the derived speech features, the second speech features of the target subject object and the derived speech features of the target subject object are updated.

15. The apparatus according to claim 11, characterized in that, The update module is also used for: Obtain the first weight of the first speech feature and the second weight of the second speech feature of the target subject object, wherein the sum of the first weight and the second weight is equal to 1; The first speech feature and the second speech feature of the target subject are weighted and summed according to the first weight and the second weight, respectively, to obtain the updated second speech feature of the target subject.

16. The apparatus according to claim 14, characterized in that, The update module is also used for: When the derived information indicates that the target subject object has the derived speech features, the derived speech features of the target subject object are updated based on the first speech features to obtain the updated derived speech features; Based on the updated derived speech features, the second speech features of the target subject are updated to obtain the updated second speech features of the target subject.

17. The apparatus according to claim 16, characterized in that, The update module is also used for: Obtain a second number of the derived speech features of the target subject object; When the second quantity is one, the derived speech features of the target subject are determined as the speech features to be updated, and the speech features to be updated are updated based on the first speech features to obtain the updated derived speech features; When there are multiple second quantities, based on the second quantity, target derived speech features are selected from the derived speech features of the target subject object, and based on the first speech features, the target derived speech features are updated to obtain the updated derived speech features.

18. The apparatus according to claim 17, characterized in that, The update module is also used for: Obtain the third feature similarity between the first speech feature and each derived speech feature of the target subject object; When the second quantity is less than the quantity threshold, the derived speech feature corresponding to the largest third feature similarity is determined as the target derived speech feature; When the second quantity is greater than or equal to the quantity threshold, the following processing is performed on each derived speech feature of the target subject object: the derived speech feature is determined as a reference speech feature, and the fourth feature similarity between the reference speech feature and each other speech feature is determined. Based on the third feature similarity and the fourth feature similarity, the target derived speech feature is determined; wherein, the other speech features are the derived speech features of the target subject object other than the reference speech feature.

19. The apparatus according to claim 17, characterized in that, The update module is also used for: Obtain the third weight of the first speech feature and the fourth weight of the speech feature to be updated, wherein the sum of the third weight and the fourth weight is 1; The first speech feature and the speech feature to be updated are weighted and summed according to the third weight and the fourth weight, respectively, to obtain the updated derived speech feature of the target subject.

20. The apparatus according to claim 16, characterized in that, The update module is also used for: Obtain the fifth weight of the updated derived speech feature, the sixth weight of the second speech feature of the target subject, and the seventh weight of the first speech feature; Wherein, the sum of the fifth weight, the sixth weight, and the seventh weight is equal to 1; The updated derived speech features, the second speech features of the target subject, and the first speech features are weighted and summed according to the fifth weight, the sixth weight, and the seventh weight, respectively, to obtain the updated second speech features of the target subject.

21. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the speech-based information recognition method according to any one of claims 1 to 10.

22. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, they implement the speech-based information recognition method according to any one of claims 1 to 10.

23. A computer program product, comprising a computer program or computer-executable instructions, characterized in that, When the computer program or computer-executable instructions are executed by a processor, the speech-based information recognition method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Audio processing method and device, storage medium and computer equipment

    CN113763962A

  • Electronic apparatus and controlling method thereof

    US20210390959A1