A keyword detection method and device, electronic equipment and storage medium

By determining the probability of audio frames and character units in keyword detection and utilizing a dynamic programming algorithm, the high computational complexity of existing technologies is solved, achieving more efficient and accurate keyword detection.

CN116416981BActive Publication Date: 2025-10-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111664577.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-10-24
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

Existing keyword detection technologies using the sliding window method suffer from high computational complexity, resulting in large computational loads and low efficiency.

Method used

By determining the first probability of a target audio frame corresponding to a target character unit in a target audio segment, and determining the second probability of a target audio segment corresponding to a preset keyword based on the first probability, a dynamic programming algorithm is used to reduce the computational load and improve detection efficiency and accuracy.

Benefits of technology

This reduces the computational load of keyword detection and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116416981B_ABST
    Figure CN116416981B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a keyword detection method and device, electronic equipment and storage medium. The method comprises: determining, for a target audio segment in target audio, a first probability that a target audio frame in the target audio segment corresponds to a target character unit, the target character unit being a character unit included in a preset keyword, a position of the target audio frame in the target audio segment corresponding to a position of the target character unit in the preset keyword; determining, according to the first probability, a second probability that the target audio segment corresponds to the preset keyword, the second probability representing probabilities that respective audio frames in the target audio segment are respective character units in the preset keyword in sequence; and determining, according to the second probability, whether the target audio segment is a voice segment of the preset keyword. The detection of the preset keyword in the audio segment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of information technology, and in particular, to a keyword detection method and device, an electronic device, and a storage medium. BACKGROUND

[0002] With the development of speech recognition technology and the continuous popularization of intelligent voice devices, recognizing preset keywords contained in audio has become an operation that needs to be performed in many scenarios.

[0003] The current keyword detection technology basically detects a piece of audio through a sliding window. This method has the problem of repeated calculation and high computational complexity for some audio. SUMMARY

[0004] To solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a keyword detection method, device, electronic device and storage medium, which reduce the amount of calculation and improve the detection efficiency and accuracy.

[0005] In a first aspect, the embodiments of the present disclosure provide a keyword detection method, which comprises:

[0006] For a target audio segment in a target audio, a first probability that a target audio frame in the target audio segment corresponds to a target character unit is determined, wherein the first probability represents a probability that the target audio frame is a speech frame of the target character unit, the target character unit is a character unit included in a preset keyword, and a position of the target audio frame in the target audio segment corresponds to a position of the target character unit in the preset keyword;

[0007] A second probability that the target audio segment corresponds to the preset keyword is determined according to the first probability, and the second probability represents a probability that each audio frame in the target audio segment is a character unit in the preset keyword in sequence;

[0008] It is determined whether the target audio segment is a speech segment of the preset keyword according to the second probability.

[0009] In a second aspect, the embodiments of the present disclosure also provide a keyword detection device, which comprises:

[0010] The first determining module is configured to determine, for a target audio segment in target audio, a first probability that a target audio frame in the target audio segment corresponds to a target character unit, wherein the first probability represents a probability that the target audio frame is a speech frame of the target character unit, the target character unit is a character unit included in a preset keyword, and a position of the target audio frame in the target audio segment corresponds to a position of the target character unit in the preset keyword.

[0011] The second determining module is configured to determine, according to the first probability, a second probability that the target audio segment corresponds to the preset keyword, the second probability representing probabilities that respective audio frames in the target audio segment are, in sequence, respective character units in the preset keyword.

[0012] The third determining module is configured to determine, according to the second probability, whether the target audio segment is a speech segment of the preset keyword.

[0013] In a third aspect, the embodiments of the present disclosure further provide an electronic device, which comprises:

[0014] one or more processors;

[0015] a storage device configured to store one or more programs;

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the keyword detection method as described above.

[0017] In a fourth aspect, the embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the keyword detection method as described above.

[0018] Compared with the prior art, the technical solutions provided by the embodiments of the present disclosure have at least the following advantages:

[0019] The keyword detection method provided by the embodiments of the present disclosure determines a first probability that a target audio frame in a target audio segment corresponds to a target character unit, wherein the target character unit is a character unit included in a preset keyword; determines, according to the first probability, a second probability that the target audio segment corresponds to the preset keyword; and determines, according to the second probability, whether the target audio segment is a speech segment of the preset keyword. Thus, the purposes of reducing computation, improving detection efficiency, and improving detection accuracy are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0020] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings. The same or similar components have the same or similar reference numbers regardless of the drawing number. It should be understood that the drawings are not necessarily to scale, with emphasis being placed upon illustrating the principles of the present disclosure.

[0021] Figure 1 A flow chart of a keyword detection method in an embodiment of the present disclosure;

[0022] Figure 2 A structural schematic diagram of a keyword detection apparatus in an embodiment of the present disclosure;

[0023] Figure 3 A structural schematic diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.

[0025] It should be understood that each step recited in the method embodiments of the present disclosure can be executed in different orders and in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.

[0026] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising, but not limited to." The term "based on" is "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related definitions of other terms will be given in the description below.

[0027] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.

[0028] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that "one" or "multiple" should be understood as "one or more" unless otherwise explicitly indicated in the context.

[0029] Names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0030] Figure 1 A flowchart of a keyword detection method in an embodiment of the present disclosure. The method is suitable for detecting a preset keyword in audio, and can be executed by a keyword detection device. The device can be implemented in software and / or hardware, and can be configured in an electronic device, such as a terminal or a server. The terminal specifically includes but is not limited to a smartphone, a palm computer, a tablet computer, a portable wearable device, a smart home device (such as a table lamp), and the like.

[0031] As shown in FIG. 1, the method specifically can include the following steps: Figure 1

[0032] In step 110, for a target audio segment in target audio, a first probability that a target audio frame in the target audio segment corresponds to a target character unit is determined.

[0033] The target audio segment is a segment of audio in the target audio. The first probability represents a probability that the target audio frame is a speech frame of the target character unit. The target character unit is a character unit included in a preset keyword. A position of the target audio frame in the target audio segment corresponds to a position of the target character unit in the preset keyword.

[0034] For example, the preset keyword is “open the door”, which includes three character units, namely “open”, “open”, and “door”. The first character unit is marked as “open”, the second character unit is marked as “open”, and the third character unit is marked as “door”. The target character unit is a specific one of the three character units. Suppose the target audio segment includes three audio frames, namely a first audio frame, a second audio frame, and a third audio frame. The target audio frame is a specific one of the three audio frames. When the first audio frame corresponds to the speech frame of the first character unit “open”, the second audio frame corresponds to the speech frame of the second character unit “open”, and the third audio frame corresponds to the speech frame of the third character unit “door”, it can be determined that the target audio segment is a speech segment of the preset keyword “open the door”. In other words, the position of the target audio frame in the target audio segment corresponds to the position of the target character unit in the preset keyword.

[0035] ​If the audio language in the target audio frame is Chinese, the character unit can be a Chinese character, specifically any one of the 5000 commonly used Chinese characters.

[0036] If the audio language in the audio frame is English, the character unit can be a syllable.

[0037] Optionally, the determining the first probability of the target audio frame in the target audio segment corresponding to the target character unit comprises: determining an audio feature of the target audio frame; inputting the audio feature into a trained neural network model to obtain the first probability of the target audio frame corresponding to the target character unit.

[0038] Optionally, taking the example that the audio language in the target audio frame is Chinese, the preset keyword is “open the light off”, and the target character unit is “hit”, “open”, “light” or “off”, the output of the neural network model is the probability of each target audio frame corresponding to “hit”, “open”, “light” and “off” respectively, that is, there are four results for one target audio frame, which are the first probability p(hit) corresponding to “hit”, the first probability p(open) corresponding to “open”, the first probability p(light) corresponding to “light” and the first probability p(off) corresponding to “off”.

[0039] In some optional embodiments, the output of the neural network model can also be the first probability of each target audio frame corresponding to the 5000 commonly used Chinese characters in the dictionary, that is, there are 5000 results for one target audio frame, which can be represented by P(t, i) to represent the first probability of the target audio frame t corresponding to the i-th Chinese character, where t represents the t-th audio frame in the target audio segment, and i represents the i-th Chinese character.

[0040] The audio feature specifically comprises fbank feature and mfcc feature, wherein the mfcc feature can be obtained by performing discrete cosine transform on the fbank feature. The fbank feature can be obtained by windowing-Fourier transform-filtering on the audio frame.

[0041] Step 120, determining a second probability of the target audio segment corresponding to the preset keyword according to the first probability, the second probability representing the probability of each audio frame in the target audio segment being the character unit in the preset keyword in order.

[0042] It is assumed that the target audio segment includes three audio frames, which are respectively a first audio frame, a second audio frame and a third audio frame. The preset keyword is "open the door", which includes three character units, which are respectively "open", "open" and "door". "Open" is marked as the first character unit, "open" is marked as the second character unit, and "door" is marked as the third character unit. The second probability represents the probability that the first audio frame corresponds to the voice frame of the first character unit "open", the second audio frame corresponds to the voice frame of the second character unit "open", and the third audio frame corresponds to the voice frame of the third character unit "door". That is, the second probability represents the probability that each audio frame in the target audio segment is in order respectively in each character unit in the preset keyword.

[0043] It can be understood that if the audio frame acquisition frequency is high, the preset keyword is long, that is, the preset keyword includes more target character units, and a single audio frame is not enough to include the preset keyword, then the preset keyword is detected according to the continuous multiple audio frames; if the audio frame acquisition frequency is low, the preset keyword is short, and a single audio frame may include a complete preset keyword, then the detection of the preset keyword is realized according to one audio frame.

[0044] In some embodiments, the second probability that the target audio segment corresponds to the preset keyword is determined according to the first probability, comprising:

[0045] (1) determining a first probability that a target audio frame corresponds to a last target character unit in the preset keyword; (2) determining a maximum value in a credibility that a target audio frame before the target audio frame in the target audio segment appears a second last target character unit in the preset keyword; (3) determining a sum of the maximum value and the first probability that the target audio frame corresponds to the last target character unit in the preset keyword as the second probability; wherein the target audio frame is any one of the audio frames in the target audio segment.

[0046] For example, the preset keyword is "open the light off", and the target character units constituting the preset keyword are four, which are respectively "open", "open", "light" and "off", that is, there are four target character units, wherein the target character unit "open" is the first target character unit of the preset keyword, the target character unit "open" is the second target character unit of the preset keyword, the target character unit "light" is the third target character unit of the preset keyword, and the target character unit "off" is the fourth target character unit of the preset keyword.

[0047] In other words, the target character unit "guan" is the last target character unit of the preset keyword, the target character unit "deng" is the second last target character unit of the preset keyword, or in other words, the target character unit "deng" is a neighbor target character unit before the target character unit "guan", the target character unit "deng" is located before the target character unit "guan" and adjacent to the target character unit "guan", so the target character unit "deng" is a neighbor target character unit before the target character unit "guan".

[0048] For example, it is assumed that there are T continuous audio frames in the target audio segment, the target audio frame is the tth audio frame, the preset keyword is "open the light off", and the target character units constituting the preset keyword are "dian", "kai", "deng" and "guan", that is, there are four target character units, and the target character unit "guan" is the last target character unit of the preset keyword. The array w[i] is used to represent each target character unit constituting the preset keyword, in this example scenario, the value range of i is 1-4, when i=1, w[i] represents the first target character unit "dian", when i=2, w[i] represents the second target character unit "kai", when i=3, w[i] represents the third target character unit "deng", and when i=4, w[i] represents the fourth target character unit "guan". In the above example scenario, the first probability that the target audio frame corresponds to the last target character unit of the preset keyword is P(t, w[4]). Specifically, the audio feature of the tth audio frame is input into the trained neural network model to obtain the first probability P(t, w[4]) that the tth audio frame corresponds to the target character unit "guan". That is, the first probability that the target audio frame corresponds to the last target character unit of the preset keyword is P(t, w[4]) in the above (1).

[0049] The maximum value of the confidence that the last-but-one target character unit of the preset keyword appears in the audio frame before the target audio frame is determined according to the above (2): the essence of this step is that it is assumed that the last target character unit (for example, “off”) of the preset keyword appears in the target audio frame (i.e., the tth audio frame), and it is necessary to determine in which audio frame the last-but-one target character unit (for example, “light”) of the preset keyword appears. After determining in which audio frame the last-but-one target character unit of the preset keyword appears, it is further necessary to determine in which audio frame the last-but-three target character unit (for example, “on”) of the preset keyword appears, until each target character unit in the preset keyword is determined to appear in which audio frame. Specifically, the confidence that the last-but-one target character unit of the preset keyword appears in each audio frame before the target audio frame is the sum of a target maximum value and a target probability. The target maximum value is the maximum value of the confidence that the last-but-three target character unit (i.e., the neighbor target character unit before the last-but-one target character unit) of the preset keyword appears in each audio frame before the target audio frame; and the target probability is the probability that the target audio frame corresponds to the last-but-one target character unit of the preset keyword.

[0050] Generally, the confidence that a target character unit of the preset keyword appears in a audio frame is the sum of a target maximum value and a target probability; the target maximum value is the maximum value of the confidence that the neighbor target character unit before the target character unit of the preset keyword appears in each audio frame before the audio frame; and the target probability is the first probability that the audio frame corresponds to the target character unit. In this way, the confidence that the target character unit appears in the first audio frame in the target audio segment is the first probability that the first audio frame corresponds to the target character unit.

[0051] It is assumed that the preset keyword is “turn on the light off”, and the plurality of target character units that constitute the preset keyword are “turn on”, “light”, and “off”, i.e., four target character units. An array w[i] is used to represent each target character unit that constitutes the preset keyword, and the value range of i is 1-4. When i=1, w[i] represents the first target character unit “turn on”, when i=2, w[i] represents the second target character unit “light”, when i=3, w[i] represents the third target character unit “off”, and when i=4, w[i] represents the fourth target character unit “off”.

[0052] The determination process of the confidence that a target character unit of the preset keyword appears in an audio frame in the target audio segment can be expressed by the following program:

[0053]

[0054]

[0055] Wherein, t represents the t-th audio frame, T represents the total number of audio frames in the target audio segment. i represents the i-th target character unit, and 4 represents the total number of target character units included in the preset keyword. Score(t,i) represents the credibility of the i-th target character unit in the preset keyword appearing in the t-th audio frame, P(t,w[i]) represents the first probability that the t-th audio frame corresponds to the i-th target character unit in the preset keyword; max(Score(tj,i-1)+P(t,w[i])) represents the maximum value of the credibility of the (i-1)-th target character unit in the preset keyword appearing in the audio frame before the t-th audio frame and P(t,w[i]), where j is less than t, and is usually taken as (t-1), (t-2), (t-3), (t-4), (t-5), (t-6) ... and other numbers closest to t, which can be specifically determined according to the total number of target character units included in the preset keyword.

[0056] Step 130: Determine whether the target audio segment is a voice segment of the preset keyword based on the second probability.

[0057] Specifically, determining whether the target audio segment is the voice segment of the preset keyword according to the second probability includes: if the second probability is greater than a preset threshold, determining that the target audio segment is the voice segment of the preset keyword.

[0058] The keyword detection method provided by the embodiment of the present disclosure, after determining the first probability that the target audio frame corresponds to the target character unit in the preset keyword, calculates the second probability of the preset keyword appearing at each moment through a dynamic programming algorithm, and determines whether the preset keyword is detected based on whether the second probability reaches a preset threshold. This can reduce the amount of calculation when detecting the preset keyword and improve the efficiency of keyword detection. Specifically, the first probability of the last target character unit of the preset keyword corresponding to the target audio frame is determined, and the maximum value of the credibility of the second to last target character unit of the preset keyword appearing in the audio frame before the target audio frame is determined; the sum of the maximum value and the first probability is determined as the second probability of the preset keyword appearing in the target audio frame, and based on the second probability, it is determined whether the target audio segment is a voice segment of the preset keyword.

[0059] Figure 2 FIG. 1 is a schematic diagram of the structure of a keyword detection device in an embodiment of the present disclosure. Figure 2 As shown, the keyword detection device specifically includes: a first determination module 210 , a second determination module 220 and a third determination module 230 .

[0060] The first determining module 210 is configured to determine a first probability that a target audio frame in a target audio segment corresponds to a target character unit in the target audio, wherein the target audio segment is in a target audio, the first probability represents a probability that the target audio frame is a speech frame of the target character unit, the target character unit is a character unit included in a preset keyword, and a position of the target audio frame in the target audio segment corresponds to a position of the target character unit in the preset keyword. The second determining module 220 is configured to determine a second probability that the target audio segment corresponds to the preset keyword according to the first probability, wherein the second probability represents probabilities that respective audio frames in the target audio segment are character units in the preset keyword in sequence. The third determining module 230 is configured to determine whether the target audio segment is a speech segment of the preset keyword according to the second probability.

[0061] Optionally, the first determining module 210 includes a first determining unit configured to determine an audio feature of the target audio frame, and an input unit configured to input the audio feature into a trained neural network model to obtain the first probability that the target audio frame corresponds to the target character unit.

[0062] Optionally, the second determining module 220 includes a second determining unit configured to determine a first probability that a target audio frame corresponds to a last target character unit in the preset keyword, a third determining unit configured to determine a maximum value in a confidence degree that an audio frame before the target audio frame in the target audio segment appears a second last target character unit in the preset keyword, and a fourth determining unit configured to determine a sum of the maximum value and the first probability that the target audio frame corresponds to the last target character unit in the preset keyword as the second probability, wherein the target audio frame is any one of the audio frames in the target audio segment.

[0063] Optionally, a confidence degree that a target audio frame appears a target character unit in a preset keyword is a sum of a target maximum value and a target probability, the target maximum value is a maximum value in a confidence degree that an audio frame before the target audio frame appears a neighbor target character unit before the target character unit in the preset keyword, and the target probability is a first probability that the target audio frame corresponds to the target character unit.

[0064] Optionally, the third determining module 230 is specifically configured to determine that the target audio segment is the speech segment of the preset keyword if the second probability is greater than a preset threshold.

[0065] Optionally, the character unit includes Chinese characters.

[0066] The keyword detection device provided by the embodiment of the present disclosure, after determining the first probability that the target audio frame corresponds to the target character unit in the preset keyword, calculates the second probability of the preset keyword appearing at each moment through a dynamic programming algorithm, and determines whether the preset keyword is detected based on whether the second probability reaches a preset threshold. This can reduce the amount of calculation when detecting the preset keyword and improve the efficiency of keyword detection. Specifically, the first probability of the last target character unit of the preset keyword corresponding to the target audio frame is determined, and the maximum value of the credibility of the second to last target character unit of the preset keyword appearing in the audio frame before the target audio frame is determined; the sum of the maximum value and the first probability is determined as the second probability of the preset keyword appearing in the target audio frame, and based on the second probability, it is determined whether the target audio segment is a voice segment of the preset keyword.

[0067] The keyword detection device provided in the embodiment of the present disclosure can execute the steps of the keyword detection method provided in the method embodiment of the present disclosure, and the execution steps and beneficial effects are not repeated here.

[0068] Figure 3 This is a schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present disclosure. The electronic device 300 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable electronic devices, and fixed terminals such as digital TVs, desktop computers, smart home devices, and the like. Figure 3 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0069] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes to implement the methods of the embodiments described in the present disclosure according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage device 308 into the random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0070] In general, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 308 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 309. The communication devices 309 can allow the electronic device 300 to communicate wirelessly or wired with other devices to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or less devices can alternatively be implemented or present.

[0071] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts, thereby implementing the methods as described above. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 309, or installed from the storage devices 308, or installed from the ROM 302. When the computer program is executed by the processing devices 301, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0072] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, in which the computer-readable program code is contained. Such a propagated data signal can take any of a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the foregoing.

[0073] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0074] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device described above, and can be accessed via the electronic device described above.

[0075] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device described above, cause the electronic device to:

[0076] For a target audio segment in the target audio, a first probability that a target audio frame in the target audio segment corresponds to a target character unit is determined, where the first probability represents a probability that the target audio frame is a speech frame of the target character unit, the target character unit is a character unit included in a preset keyword, and a position of the target audio frame in the target audio segment corresponds to a position of the target character unit in the preset keyword; a second probability that the target audio segment corresponds to the preset keyword is determined according to the first probability, the second probability representing probabilities that respective audio frames in the target audio segment are character units in the preset keyword in sequence; and whether the target audio segment is a speech segment of the preset keyword is determined according to the second probability.

[0077] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0078] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or combinations of hardware and software.

[0079] The units described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0080] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0081] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0082] According to one or more embodiments of the present disclosure, the present disclosure provides a keyword detection method, which comprises: for a target audio segment in target audio, determining a first probability that a target audio frame in the target audio segment corresponds to a target character unit, wherein the first probability represents a probability that the target audio frame is a speech frame of the target character unit, the target character unit is a character unit included in a preset keyword, and a position of the target audio frame in the target audio segment corresponds to a position of the target character unit in the preset keyword; determining a second probability that the target audio segment corresponds to the preset keyword according to the first probability, the second probability representing a probability that each audio frame in the target audio segment is a character unit in the preset keyword in sequence; and determining whether the target audio segment is a speech segment of the preset keyword according to the second probability.

[0083] According to one or more embodiments of the present disclosure, in the keyword detection method provided by the present disclosure, optionally, the determining the first probability that the target audio frame in the target audio segment corresponds to a target character unit comprises: determining an audio feature of the target audio frame; inputting the audio feature into a trained neural network model to obtain the first probability that the target audio frame corresponds to the target character unit.

[0084] According to one or more embodiments of the present disclosure, in the keyword detection method provided by the present disclosure, optionally, the determining the second probability that the target audio segment corresponds to the preset keyword according to the first probability comprises: determining a first probability that a target audio frame corresponds to a last target character unit in the preset keyword; determining a maximum value in a confidence that an audio frame before the target audio frame in the target audio segment appears a second last target character unit in the preset keyword; determining a sum of the maximum value and the first probability that the target audio frame corresponds to the last target character unit in the preset keyword as the second probability; wherein the target audio frame is any one of the audio frames in the target audio segment.

[0085] According to one or more embodiments of the present disclosure, in the keyword detection method provided by the present disclosure, optionally, a confidence that the one audio frame appears a target character unit in the preset keyword is a sum of a target maximum value and a target probability; the target maximum value is a maximum value in a confidence that an audio frame before the one audio frame appears a neighbor target character unit before the target character unit in the preset keyword; and the target probability is a first probability that the one audio frame corresponds to the target character unit.

[0086] According to one or more embodiments of the present disclosure, in the keyword detection method provided by the present disclosure, optionally, the determining whether the target audio segment is a voice segment of the preset keyword according to the second probability comprises: if the second probability is greater than a preset threshold, determining that the target audio segment is the voice segment of the preset keyword.

[0087] According to one or more embodiments of the present disclosure, in the keyword detection method provided by the present disclosure, optionally, the character unit comprises a Chinese character.

[0088] According to one or more embodiments of the present disclosure, the present disclosure provides a keyword detection apparatus, comprising: a first determination module configured to determine, for a target audio segment in target audio, a first probability that a target audio frame in the target audio segment corresponds to a target character unit, wherein the first probability represents a probability that the target audio frame is a speech frame of the target character unit, the target character unit is a character unit included in a preset keyword, and a position of the target audio frame in the target audio segment corresponds to a position of the target character unit in the preset keyword; a second determination module configured to determine, according to the first probability, a second probability that the target audio segment corresponds to the preset keyword, the second probability representing probabilities that respective audio frames in the target audio segment sequentially correspond to respective character units in the preset keyword; and a third determination module configured to determine, according to the second probability, whether the target audio segment is a speech segment of the preset keyword. According to one or more embodiments of the present disclosure, in the keyword detection apparatus provided by the present disclosure, optionally, the first determination module comprises a first determination unit configured to determine an audio feature of the target audio frame; and an input unit configured to input the audio feature into a trained neural network model to obtain the first probability that the target audio frame corresponds to the target character unit.

[0089] According to one or more embodiments of the present disclosure, in the keyword detection apparatus provided by the present disclosure, optionally, the second determination module comprises: a second determination unit configured to determine a first probability that a target audio frame corresponds to a last target character unit in the preset keyword; a third determination unit configured to determine a maximum value in a credibility that audio frames before the target audio frame in the target audio segment appear a second last target character unit in the preset keyword; a fourth determination unit configured to determine, as the second probability, a sum of the maximum value and the first probability that the target audio frame corresponds to the last target character unit in the preset keyword; and wherein the target audio frame is any one of the audio frames in the target audio segment.

[0090] According to one or more embodiments of the present disclosure, in the keyword detection apparatus provided by the present disclosure, optionally, a credibility that a target audio frame appears a target character unit in the preset keyword is a sum of a target maximum value and a target probability; the target maximum value is a maximum value in a credibility that audio frames before the target audio frame appear a neighbor target character unit before the target character unit in the preset keyword; and the target probability is a first probability that the target audio frame corresponds to the target character unit.

[0091] According to one or more embodiments of the present disclosure, in the keyword detection apparatus provided by the present disclosure, optionally, the third determination module is specifically configured to: if the second probability is greater than a preset threshold, determining that the target audio segment is a voice segment of the preset keyword.

[0092] According to one or more embodiments of the present disclosure, in the keyword detection apparatus provided by the present disclosure, optionally, the character unit includes Chinese characters.

[0093] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, comprising:

[0094] one or more processors;

[0095] a memory for storing one or more programs;

[0096] When the one or more programs are executed by the one or more processors, the one or more processors implement the keyword detection method according to any one of the present disclosure.

[0097] According to one or more embodiments of the present disclosure, the present disclosure provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the keyword detection method according to any one of the present disclosure.

[0098] The present disclosure also provides a computer program product, which includes a computer program or instructions, and the computer program or instructions are executed by a processor to implement the keyword detection method as described above.

[0099] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.

[0100] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0101] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A keyword detection method characterized by, The method comprises: For a target audio segment in target audio, determining a first probability that a target audio frame in the target audio segment corresponds to a target character unit, wherein the first probability represents a probability that the target audio frame is a speech frame of the target character unit, the target character unit being a character unit included in a preset keyword, and a position of the target audio frame in the target audio segment corresponding to a position of the target character unit in the preset keyword; determining a second probability that the target audio segment corresponds to the preset keyword according to the first probability, the second probability representing probabilities that respective audio frames in the target audio segment are respective character units in the preset keyword in sequence; determining whether the target audio segment is a speech segment of the preset keyword according to the second probability; The determining of the second probability that the target audio segment corresponds to the preset keyword according to the first probability comprises: determining a first probability that a target audio frame corresponds to a last target character unit in the preset keyword; determining a maximum value of a confidence level that an audio frame before the target audio frame in the target audio segment appears a second last target character unit in the preset keyword; determining a sum of the maximum value and the first probability that the target audio frame corresponds to the last target character unit in the preset keyword as the second probability. The target audio frame is any one of the audio frames in the target audio segment.

2. The method of claim 1, wherein, The determining of the first probability that the target audio frame in the target audio segment corresponds to the target character unit comprises: determining an audio feature of the target audio frame; inputting the audio feature into a trained neural network model to obtain the first probability that the target audio frame corresponds to the target character unit.

3. The method of claim 1, wherein, The confidence level that the one audio frame appears the one target character unit in the preset keyword is a sum of a target maximum value and a target probability; The target maximum value is a maximum value of confidence levels that audio frames before the one audio frame appear neighbor target character units before the one target character unit in the preset keyword; The target probability is a first probability that the one audio frame corresponds to the one target character unit.

4. The method according to any one of claims 1 to 3, characterized in that, The determining of whether the target audio segment is the speech segment of the preset keyword according to the second probability comprises: if the second probability is greater than a preset threshold, determining that the target audio segment is the speech segment of the preset keyword.

5. The method according to any one of claims 1 to 3, characterized in that, The character unit comprises Chinese characters.

6. A keyword detection apparatus characterized by comprising: The apparatus comprises: a first determining module configured to, for a target audio segment in target audio, determine a first probability that a target audio frame in the target audio segment corresponds to a target character unit, wherein the first probability represents a probability that the target audio frame is a speech frame of the target character unit, the target character unit being a character unit included in a preset keyword, and a position of the target audio frame in the target audio segment corresponding to a position of the target character unit in the preset keyword; The second determining module is configured to determine a second probability that the target audio segment corresponds to the preset keyword according to the first probability, the second probability representing probabilities that respective audio frames in the target audio segment are respective character units in the preset keyword in sequence; The third determining module is configured to determine whether the target audio segment is a voice segment of the preset keyword according to the second probability; The second determining module includes: A second determining unit configured to determine a first probability that a target audio frame corresponds to a last target character unit in the preset keyword; A third determining unit configured to determine a maximum value in a confidence degree that a target audio frame before the target audio frame in the target audio segment appears a second last target character unit in the preset keyword; A fourth determining unit configured to determine a sum of the maximum value and the first probability that the target audio frame corresponds to the last target character unit in the preset keyword as the second probability; The target audio frame is any one of audio frames in the target audio segment.

7. The apparatus of claim 6, wherein, The first determining module includes: A first determining unit configured to determine an audio feature of the target audio frame; An input unit configured to input the audio feature to a trained neural network model to obtain the first probability that the target audio frame corresponds to the target character unit.

8. An electronic device, comprising: The electronic device includes: One or more processors; A storage device configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method in any one of claims 1-5.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method in any one of claims 1-5.

Citation Information

Patent Citations

  • Keyword detection method in voice signal, device, terminal and storage medium

    CN108615526A

  • Keyword detection method, device, storage medium and electronic equipment

    CN110414450A