A speech recognition method, device, electronic device, system and medium

The microphone array detects the similarity between the wake-up word and the lyrics information to prevent the intelligent voice system from being mistakenly woken up when the user is listening to or singing. This solves the problem of false wake-up caused by the similarity between the lyrics and the wake-up word and improves the user experience.

CN114694653BActive Publication Date: 2025-09-19SHENZHEN HORIZON ROBOTICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210350012.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-02
Publication Date
2025-09-19
Estimated Expiration
2042-04-02

AI Technical Summary

Technical Problem

When users are listening to or singing songs, if the lyrics of the songs are similar to the wake-up words of the intelligent voice system, it may cause false wake-up and affect the user experience.

Method used

The voice signal is obtained through the microphone array, the similarity score between the wake-up word and the lyrics information is detected, the validity of the wake-up word is judged, and the wake-up word is not responded to when the similarity is higher than the threshold. The similarity score of the wake-up sentence is obtained to further judge the validity of the wake-up word to ensure that the intelligent voice system is not woken up by mistake.

Benefits of technology

Effectively prevent the intelligent voice system from being woken up by mistake, improve the user's listening or singing experience, and maintain smoothness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694653B_ABST
    Figure CN114694653B_ABST
Patent Text Reader

Abstract

Disclosed are a speech recognition method, apparatus, electronic device, system, and medium. The method includes: obtaining lyrics information corresponding to a song in response to a song playback instruction; obtaining a voice signal based on a microphone array, and detecting a wake-up word based on the voice signal; determining a first similarity score between the wake-up word and the lyrics information; when the first similarity score between the wake-up word and the lyrics information is greater than or equal to a first threshold, determining that a first detection type of the wake-up word is invalid and not responding to the wake-up word; if the first similarity score between the wake-up word and the lyrics information is less than the first threshold, obtaining a wake-up sentence containing the wake-up word; and determining, based on a second similarity score between the wake-up sentence and the lyrics information, that a second detection type of the wake-up word is invalid and not responding to the wake-up word. This method can effectively prevent the intelligent voice system from being mistakenly awakened while listening to or singing, maintaining the user's listening or singing fluency, and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech recognition, and in particular to a speech recognition method, apparatus, electronic device, system, and computer-readable storage medium. Background Art

[0002] With the continuous development of science and technology and people's pursuit of a higher quality of life, more and more devices have voice control functions, such as cars, home appliances, and smart homes. One scenario is when a user is listening to or singing a song. If the lyrics of the song being played or sung are the same or similar to the wake-up word preset in the intelligent voice system, there is a possibility that the intelligent voice system will be mistakenly woken up. For example, if the lyrics include "next song" or "next", the intelligent voice system may be triggered to wake up in response to the current wake-up word, causing the song to switch or play the next song, thus affecting the user experience. Summary of the Invention

[0003] In order to solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a speech recognition method, apparatus, electronic device, system, and medium.

[0004] According to one aspect of an embodiment of the present disclosure, a speech recognition method is provided, the method comprising:

[0005] In response to a song play instruction, obtaining lyrics information corresponding to the song;

[0006] Acquire a voice signal based on a microphone array, and detect a wake-up word based on the voice signal;

[0007] Determining a first similarity score between the wake-up word and the lyrics information;

[0008] Based on the first similarity score being greater than or equal to a first threshold, determining that the first detection type of the wake-up word is invalid, and not responding to the wake-up word;

[0009] Based on the first similarity score being less than the first threshold, obtaining a wake-up sentence including the wake-up word;

[0010] determining a second detection type of the wake-up word based on a second similarity score between the wake-up sentence and the lyrics information;

[0011] The second detection type based on the wake-up word is invalid, and the wake-up word is not responded to;

[0012] Based on the second detection type of the wake-up word being valid, respond to the wake-up word.

[0013] According to another aspect of the embodiments of the present disclosure, a speech recognition device is provided, the device comprising:

[0014] Lyrics acquisition module, used for obtaining lyrics information corresponding to a song in response to a song playing instruction;

[0015] A voice acquisition module is used to acquire voice signals based on a microphone array;

[0016] a speech detection module, configured to detect a wake-up word based on the speech signal acquired by the speech acquisition module, determine a first similarity score between the wake-up word and the lyric information acquired by the lyric acquisition module, and, based on the first similarity score being greater than or equal to a first threshold, determine that a first detection type of the wake-up word is invalid; based on the first similarity score being less than the first threshold, acquire a wake-up statement containing the wake-up word, and, based on a second similarity score between the wake-up statement and the lyric information, determine that a second detection type of the wake-up word is invalid;

[0017] A voice processing module is used to not respond to the wake-up word based on the first detection type of the wake-up word determined by the voice detection module to be invalid, or the second detection type of the wake-up word to be invalid, and to respond to the wake-up word based on the second detection type of the wake-up word determined to be valid.

[0018] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0019] processor;

[0020] a memory for storing instructions executable by the processor;

[0021] The processor is configured to read executable instructions from the memory and execute the instructions to implement the above-mentioned speech recognition method.

[0022] Optionally, the processor includes but is not limited to a vehicle computer, a vehicle computer processor, a CPU, a processing module, etc. The memory includes but is not limited to a random access memory RAM, a read-only memory ROM, an erasable programmable read-only memory EPROM, etc.

[0023] According to another aspect of the embodiment of the present disclosure, a speech recognition system is further provided, the system comprising a speech playback device, a speech collection device, and a speech recognition device; wherein,

[0024] The voice playback device is used to respond to a song playback instruction and play the audio and video corresponding to the instruction;

[0025] The voice collection device is used to collect the voice signal input by the user;

[0026] The speech recognition device is used to call computer program instructions stored in a memory based on the audio and video played by the speech playback device and the speech signal collected by the speech collection device, and execute the instructions to implement the above-mentioned speech recognition method.

[0027] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is used to execute the above-mentioned speech recognition method.

[0028] Based on the voice recognition method, device electronic device, system and medium provided by the above embodiments of the present disclosure, when a wake-up word is detected in the user's voice signal, the wake-up word is compared with the lyrics information of the currently playing song or the sung song for similarity. When the first similarity score is greater than or equal to the first threshold, it is determined that the detected wake-up word is the lyrics of the currently playing song or the sung song, and the wake-up word is invalid, so the wake-up word is not responded to, and the song continues to play or the accompaniment music continues to play, thereby effectively preventing the intelligent voice system from being woken up by mistake and improving the user experience.

[0029] In addition, when the first similarity score is less than the first threshold, a wake-up statement containing the wake-up word is further obtained, and the similarity between the wake-up statement and the lyrics information is compared to obtain a second similarity score. When the second similarity score is greater than or equal to the second threshold, it is determined that the wake-up statement is a line of lyrics in the lyrics information, the wake-up word is invalid, and the wake-up word is not responded to. The song or the accompaniment music continues to be played to prevent the system from being woken up by mistake.

[0030] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The above and other purposes, features, and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0032] Figure 1 is a system diagram of speech recognition to which the present disclosure is applicable;

[0033] Figure 2 is a flowchart of a speech recognition method provided by an exemplary embodiment of the present disclosure;

[0034] Figure 3 This is a schematic diagram of detecting the similarity between a wake-up word and lyrics information disclosed in the present invention;

[0035] Figure 4 is a flow chart of a speech recognition method provided by another exemplary embodiment of the present disclosure;

[0036] Figure 5 is a structural diagram of a speech recognition device provided by an exemplary embodiment of the present disclosure;

[0037] Figure 6 is a structural diagram of a voice detection module provided by an exemplary embodiment of the present disclosure;

[0038] Figure 7 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0040] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.

[0041] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.

[0042] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.

[0043] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0044] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.

[0045] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.

[0046] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0047] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0048] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0049] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0050] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems.

[0051] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.

[0052] Application Overview

[0053] Artificial Intelligence (AI) enables machines to perform complex tasks that typically require human intelligence. Efficient and accurate human-machine interaction is essential for executing human commands. In recent years, with the continuous development of AI technology, the application of speech recognition technology in smart devices has attracted increasing attention within the industry.

[0054] Especially in driving scenarios, in order to avoid the inconvenience and insecurity of users manually operating and controlling in-vehicle equipment, in related technologies, in-vehicle equipment is controlled through voice commands.

[0055] However, the inventors have discovered through research that when a user is listening to or singing a song, if the lyrics of the song being played or the lyrics of the song being sung are the same as or similar to a wake-up word preset by the voice recognition system, it may trigger a false wake-up, causing the voice recognition system to respond to the wake-up word content, affecting the user's listening or singing experience.

[0056] In view of this, the embodiments of the present disclosure propose a voice recognition method, device, electronic device, system and medium to reduce false wake-up of voice recognition and enhance the user's experience of listening to or singing music.

[0057] The disclosed embodiment acquires a voice signal through a microphone array to detect a wake-up word, determines a first similarity score between the wake-up word and the lyrics information, that is, judges the similarity between the wake-up word and the lyrics information of the currently playing song or the sung song, and determines a first detection type of the wake-up word based on the first similarity score to judge whether the wake-up word is valid. When the first similarity score is greater than or equal to a first threshold, it is determined that the wake-up word is invalid, and the wake-up word is not responded to, and the song or the accompaniment music continues to be played, thereby effectively solving the problem of the intelligent voice system being mistakenly awakened. This method improves the user experience.

[0058] In addition, when the first similarity score is less than the first threshold, a wake-up statement containing the wake-up word is obtained, and a second similarity score between the wake-up statement and the lyrics information is determined. The second detection type of the wake-up word is determined based on the second similarity score. When it is determined that the second detection type of the wake-up word is invalid, the wake-up word is not responded to, and the song continues to be played, thereby maintaining the fluency of the user's listening or singing, thereby improving the user experience.

[0059] Exemplary Systems

[0060] Figure 1 This is a system diagram to which the present disclosure applies. Figure 1 As shown, the system architecture 100 may include a terminal device 101 , a microphone array 102 and an audio playback device 103 .

[0061] The terminal device 101 may be installed with various application software, such as multimedia applications, search applications, web browser applications, instant messaging tools, etc. In this embodiment, the terminal device 101 mainly installs multimedia applications for playing media streams, such as audio and video of songs.

[0062] The microphone array 102 can collect audio signals emitted in the target space, such as the audio signal of a user singing. The audio playback device 103 can play the audio signals collected by the microphone array 102, and can also play songs in multimedia applications.

[0063] The terminal device 101 can be various electronic devices, including but not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.

[0064] in addition, Figure 1 The number of terminal devices 101, microphone arrays 102, and audio playback devices 103 in the above description is merely illustrative. Any number of terminal devices 101, microphone arrays 102, and audio playback devices 103 may be provided as needed. For example, if audio signals do not require remote processing, the above system architecture may include only the microphone array, terminal devices, and audio playback devices, without including a network or server.

[0065] Exemplary Methods

[0066] See also Figure 2 This is a flow chart of a speech recognition method provided by an exemplary embodiment of the present disclosure. This method can be applied to electronic devices such as Figure 2 As shown, the method includes the following steps:

[0067] Step 201: In response to a song play instruction, obtain lyrics information corresponding to the song.

[0068] In this embodiment, the electronic device can receive a song playback instruction input by the user. The song playback instruction can be a voice instruction or a gesture instruction. For example, the user clicks on the display screen of the electronic device to launch a multimedia application and clicks on the song playback. In response to the user's click operation, the electronic device starts and plays the specified song or the song accompaniment.

[0069] It should be noted that the electronic device can directly play a designated song or the song's accompaniment music through an audio playback device. For example, in response to a user's song playback instruction, the electronic device activates song playback software to play the accompaniment music of the designated song. The user can then sing along with the accompaniment music and output a voice signal. In this embodiment, when the user sings, the accompaniment music of the designated song is generally played, excluding the original soundtrack of the song. The original soundtrack can be enabled as desired by the user.

[0070] Lyrics information is the lyrics corresponding to the currently playing song and is bound to the currently playing song. This lyric information includes lyrics files, such as LRC files. LRC files are created by editing the lyrics according to their appearance time, and then displaying them synchronously during song playback.

[0071] In this embodiment, the lyrics information includes a time tag (Time-tag) and an identification tag (ID-tags). Among them, the time tag is in the form of "[mm:ss]" (minutes:seconds) or "[mm:ss.ff]". Optionally, the time tag can be located at the beginning of a sentence in a line of lyrics. When the song playback reaches a certain time point, the electronic device will look for the corresponding time tag and display the lyrics text after the tag, thereby achieving the functional effect of "lyrics synchronization". In addition, the identification tag, whose format is "[identification name: value]", can include the following predefined tags: [ar: singer name], [ti: song name], [al: album name], [by: editor (referring to the producer of lrc lyrics)], [offset: time compensation value].

[0072] In step 201, the electronic device retrieves the corresponding lyrics information while playing the specified song based on the user's song playback instruction. This retrieval method includes retrieving the lyrics from a local resource library or downloading them from the cloud. This embodiment does not limit the specific implementation method for retrieving the lyrics information.

[0073] In one example, a song that a user in a vehicle wants to sing is called "Spring Mud", and a song play instruction is issued to the electronic device. The electronic device responds to the song play instruction and plays the user-specified song "Spring Mud" or the accompaniment music of "Spring Mud" through the audio playback device, and at the same time obtains the lyrics information of "Spring Mud", that is, the LRC file of the song "Spring Mud", which includes the time tag and identification tag of the song "Spring Mud".

[0074] In addition, when the electronic device plays the song "Spring Mud" or the accompaniment music of "Spring Mud", the microphone array also collects the audio signal output by the user when singing. After audio processing, the audio playback device replays the sound of the user's singing.

[0075] Step 202: Acquire a voice signal based on a microphone array, and detect a wake-up word based on the voice signal.

[0076] Specifically, the electronic device can obtain at least one audio signal of the user singing from the microphone array. Figure 1The microphone array 102 shown is used to collect sound from a target space, generating at least one audio signal, with each audio signal corresponding to a microphone. The microphone array converts and processes the at least one audio signal to generate a corresponding voice signal, including the audio signal output by a user singing. The target space can be any space, such as a car or a room.

[0077] In this embodiment, the process and implementation method of converting at least one audio signal into a voice signal are not limited.

[0078] It should be noted that the at least one audio signal captured by the microphone array also includes the audio signal played by the audio playback device. For example, if the multimedia application software of the electronic device plays the song "Spring Mud" or the accompaniment music of the song "Spring Mud", the microphone array also captures the audio signal of the song "Spring Mud" played by the multimedia application.

[0079] The above-mentioned wake-up word detection based on voice signals means that at least one word in the information content based on voice signal analysis is the same as the wake-up word of the intelligent voice function in the electronic device, that is, one or more words in the voice signal can wake up the intelligent voice function of the electronic device, and the one or more words are called "wake-up words".

[0080] In one example, the voice signal collected by the electronic device when the user is singing includes "next", which is the same as the wake-up word "next" of the intelligent voice function, then the "next" in the voice signal is used as the currently detected wake-up word.

[0081] It should be understood that the above wake-up words can also be other words, such as next song, previous song, previous song, etc., and this embodiment does not limit this.

[0082] Step 203: Determine a first similarity score between the wake-up word and the lyrics information.

[0083] The first similarity score is used to measure the similarity between the wake-up word detected in the previous step (step 202) and the lyrics of the currently playing song. In this embodiment, the first similarity score can be used to detect whether the detection type of the wake-up word is valid. A higher first similarity score indicates that the detected wake-up word is the lyrics of the currently playing song, and the first detection type of the current wake-up word is invalid. Conversely, a lower first similarity score indicates that the wake-up word is not the lyrics of the song, and the first detection type is valid.

[0084] In this embodiment, a scoring tool, such as a text similarity scoring tool, may be used to obtain the first similarity score.

[0085] In one example, when an electronic device plays the accompaniment music for the song "Spring Mud," and collects a voice signal of the user singing "Blooming the next season of flowers," it detects that the wake-up word currently present is "next," then compares the similarity between the wake-up word "next" and the lyrics of "Spring Mud" to obtain a first similarity score. Optionally, the first similarity score can range from 0 to 100.

[0086] In this embodiment, a text similarity scoring tool can be used to obtain the first similarity score. Figure 3 As shown, the text similarity scoring tool is used to obtain the first similarity scores of the following three wake-up words "next", "next song", and "previous song" with the lyrics information "open the next flower season", and three return results are obtained:

[0087] ("next", "bloom the next flower season"), the first similarity score of the returned result is 100 points;

[0088] ("next song", "bloom the next flower season"), the first similarity score of the returned result is 75 points;

[0089] ("Previous song", "Blooming the next flower season"), the first similarity score of the returned result is 33 points.

[0090] Step 204: Determine whether the first similarity score is greater than or equal to a first threshold value, wherein the first threshold value can be customized.

[0091] If the first similarity score is greater than or equal to the first threshold, the first detection type of the currently detected wake-up word is determined to be invalid, and the following step 205 is executed. If the first similarity score is less than the first threshold, the first detection type of the wake-up word is determined to be valid, and the following step 206 is executed.

[0092] Step 205: Do not respond to the wake-up word.

[0093] Step 206: Obtain a wake-up sentence containing the wake-up word.

[0094] In this example, the first threshold is set to 60 points. When the wake-up word is "next", the first similarity score is 100, which is greater than the first threshold of 60. In this case, the first detection type of the wake-up word "next" is determined to be invalid, and the wake-up word is not responded to. When the wake-up word is "next song", the first similarity score is 75, which is greater than the first threshold of 60. In this case, the first detection type of the wake-up word "next song" is determined to be invalid, and the wake-up word is not responded to.

[0095] When the wake-up word is “previous song”, the first similarity score of 33 is obtained, which is less than the first threshold of 60, and the first detection type is determined to be valid, and step 206 is executed.

[0096] The wake-up statement is a sentence containing the currently detected wake-up word, or a paragraph containing the wake-up word and the upper and lower sentences of the wake-up word. Because there may be contextual sentences in the process of the user singing, and there may be a time difference between the audio signal when the user sings and the lyrics of the audio and video being played by the electronic device, it is necessary to obtain the sentence or paragraph corresponding to the wake-up word (referred to as the wake-up statement in this embodiment) in order to further determine whether the wake-up word is valid.

[0097] In this example, the audio playback device 103 plays the accompaniment music of the song "Spring Mud". When the electronic device detects that the wake-up word is "next", it obtains the wake-up sentence containing the wake-up word "next", that is, "bloom the next flower season", providing a basis for subsequent verification of the validity of the wake-up word.

[0098] Step 207: Determine a second similarity score between the wake-up sentence and the lyrics information.

[0099] The second similarity score is used to verify whether the detection type of the wake-up word is valid. Furthermore, the method for determining the second similarity score can be the same as the method for determining the first similarity score in step 203. For example, a text similarity scoring tool can be used to obtain the second similarity score of the wake-up sentence. In addition, the lyrics information can be all or part of the lyrics of the currently playing song.

[0100] In this example, the wake-up sentence "Bloom in the next flower season" is compared with all the lyrics in the song "Spring Mud" for similarity, and a second similarity score is obtained by comprehensive scoring, for example, the second similarity score is 99 points.

[0101] Step 208: Determine a second detection type of the wake-up word based on a second similarity score between the wake-up sentence and the lyrics information.

[0102] In this embodiment, to distinguish the first detection type in step 203, the detection type of the wake-up word is defined as the second detection type. Step 208 determines the second detection type based on the second similarity score, specifically including: comparing the second similarity score with the second threshold. If the second similarity score is greater than or equal to the second threshold, the second detection type of the current wake-up word is determined to be invalid, and the wake-up word is not responded to, similar to step 205 above. If the second similarity score is less than the second threshold, the second detection type of the wake-up word is determined to be valid, and step 209 is executed.

[0103] The second threshold may be the same as or different from the first threshold, and this embodiment does not impose any limitation on this.

[0104] In this example, the second threshold is set to 60, which is the same as the aforementioned first threshold. Then, the second similarity score 99 is greater than the second threshold 60, and it is determined that the second detection type of the wake-up word is invalid, and the wake-up word is not responded to.

[0105] Step 209: Respond to the wake-up word and execute the instruction of the wake-up word.

[0106] The method provided in this embodiment uses the first similarity score during the wake-up word detection process. When it is determined that the first detection type of the wake-up word is invalid, it is determined that the wake-up word is highly similar to the lyrics of the currently playing song or the song being sung, indicating that the song the user is singing is an invalid wake-up word. At this time, the wake-up word is not responded to, and the song or the accompaniment music continues to be played, thereby preventing the system from being mistakenly awakened and affecting the user's singing experience.

[0107] In addition, when it is determined that the first detection type of the wake-up word is valid, continue to search for the wake-up sentence (context sentence) where the wake-up word is located, and use the second similarity score to perform similarity judgment with the song lyrics. When it is determined that the second detection type of the wake-up word is invalid, do not respond to the wake-up word, continue to play the song or continue to play the accompaniment music to prevent the system from being woken up by mistake and affecting the user's singing experience.

[0108] In some optional implementations, the above step 203, determining the first similarity score, specifically includes the following sub-steps:

[0109] The similarity between the wake-up word and the lyrics information is scored based on a preset algorithm to obtain the first similarity score.

[0110] The preset algorithm is used to score the similarity between the wake-up word and the lyrics information. Based on the preset algorithm, the similarity score of the currently detected wake-up word and the full text of the lyrics information is scored, and then the scores are combined to obtain a first similarity score. The preset algorithm includes but is not limited to simple matching, incomplete matching, order-ignoring matching, duplicate-removing subset matching, fuzzy matching algorithms, etc. It should be understood that the preset algorithm can also be based on a combination of one or more of the above matching methods.

[0111] Similarly, when determining the second similarity score in the above step 207, the similarity comparison can be performed on the full text of the wake-up sentence and the lyrics information based on the above preset algorithm, and then the similarity scores of each word are combined to obtain the second similarity score.

[0112] This implementation uses a preset algorithm to score the similarity between the wake-up word and the lyrics, obtaining a corresponding similarity score, thereby determining the effectiveness of the current wake-up word. Furthermore, the preset algorithm can be predefined, increasing the accuracy and flexibility of the similarity comparison. This method solves the problem of the intelligent voice system being mistakenly awakened when the user sings.

[0113] In some optional implementations, after step 204 and before step 206, the following steps may be further included:

[0114] The context of the wake-up word is determined based on voice activity detection, that is, whether the wake-up word contains a context sentence. If it contains a context sentence, the above step 206 also includes: based on the detected context of the wake-up word, obtaining a wake-up sentence containing the wake-up word.

[0115] Specifically, one method for detecting whether a context sentence is included is to use activation sound detection technology, such as Voice Activity Detection (VAD), also known as voice endpoint detection, to detect whether the VAD detection results are all 0 within a preset time period (e.g., within 1 second) before and after the current wake-up word. If not, it is determined that the wake-up word includes context, and then based on the context of the wake-up word, a wake-up sentence containing the wake-up word is obtained.

[0116] The preset time period is within 1 second, which can be understood as the time period 1 second before and 1 second after the detected wake-up word. In this embodiment, setting the preset time for VAD to detect the wake-up word within 1 second can reduce latency and improve the accuracy of speech recognition. It should be understood that the preset time period can also be set to a longer or shorter time interval, and this embodiment is not limited to this.

[0117] In one example, based on VAD detection, the current wake-up word "next" has the context sentences "bloom" and "flower season", the VAD detection result is not 0, and the wake-up sentence containing "next" is "bloom the next flower season".

[0118] This implementation method can accurately detect whether the wake-up word contains a context sentence through the VAD method, which can prepare for the similarity judgment of subsequent wake-up sentences and provide a basis for the effectiveness judgment of subsequent wake-up words.

[0119] In other optional embodiments, during the above-mentioned VAD detection process, the method further includes: if the VAD detection results of the current wake-up word are all 0, the context of the wake-up word is not detected, and the wake-up statement that does not include the wake-up word is determined to be only a command statement, thereby determining that the third detection type of the wake-up word is valid. At this time, the system is awakened and responds to the wake-up word.

[0120] For example, when performing VAD detection on the wake-up word "next song", the VAD result of the wake-up word "next song" is detected to be 0, that is, it does not include context, then the third detection type of the wake-up word is determined to be valid, the system is woken up, and in response to the voice command of the wake-up word "next song", the "next song" is played.

[0121] This implementation can accurately detect whether the wake-up word contains a context sentence through the VAD method, realize the verification of the wake-up word context, and wake up the voice system to respond to the wake-up word when the detection does not include a context sentence.

[0122] In some alternative embodiments, see Figure 4 , the above step 207 can be performed as follows:

[0123] Step 207 - 1 : Score the similarity between the wake-up statement and the lyrics information based on a preset algorithm to obtain a second similarity score.

[0124] The preset algorithm may be the same as the preset algorithm of step 203. One possible implementation of step 207-1 is to compare the similarity of each word in the current wake-up sentence with the lyrics information based on the preset algorithm, and then comprehensively score the similarities of each word to obtain a second similarity score.

[0125] In addition, the above step 208 specifically further includes:

[0126] Step 208 - 1 : Based on the second similarity score being greater than or equal to a second threshold, determine that the second detection type of the wake-up word is invalid.

[0127] The second detection type is used to detect the validity of the wake-up word contained in the wake-up statement. This can be determined by the second similarity score and the second threshold. When the second detection type of the wake-up word is determined to be invalid, step 205 is executed above, and the wake-up word is not responded to. The song or the song accompaniment music continues to play, and the user continues to sing or listen to the song.

[0128] Step 208 - 2 : Based on the second similarity score being less than a second threshold, determining that the second detection type of the wake-up word is valid.

[0129] At this time, the above step 209 is executed to respond to the wake-up word.

[0130] For the specific detection process, please refer to the description of steps 203 and 204 in the above embodiment, which will not be repeated here in this embodiment. In this embodiment, a preset algorithm is used to determine the second similarity score of the current wake-up word, and the second similarity score is compared with the second threshold value to implement the validity detection of the second detection type of the wake-up word. When the second detection type of the wake-up word is detected to be invalid, the wake-up word is not responded to, maintaining the fluency of the user listening to or singing, and improving the user experience.

[0131] In some further optional implementations, based on the above step 207 - 1 , the method further includes: scoring the similarity between the wake-up sentence and the full text of the lyrics information based on a preset algorithm to determine a second similarity score.

[0132] In one example, the wake-up sentence is "Bloom in the next flower season", and the wake-up sentence is compared and scored with the full lyrics of the song "Spring Mud", and finally the second similarity score of the wake-up word "next" is determined.

[0133] In this embodiment, a preset algorithm is used to compare the similarity between the wake-up sentence and the full text of the lyrics, thereby preventing the system from being mistakenly awakened due to differences between the lyrics content and the wake-up word. This method can improve the accuracy of wake-up word detection.

[0134] In some optional implementations, based on the above step 207-1, the following steps may be further performed:

[0135] First, based on the acquisition time of the wake-up sentence and the lyrics information, the lyrics of the song to be played are determined.

[0136] Then, based on the lyrics of the song played, the preset range of the lyrics information is determined.

[0137] Finally, a similarity score is performed on a preset range of the wake-up sentence and the lyrics information based on a preset algorithm to determine a second similarity score.

[0138] Specifically, the wake-up sentence is acquired during the time period corresponding to the context sentences before and after the wake-up word is detected. For example, in one example, the wake-up word "next" is detected at [01:08] (1 minute and 8 seconds) of the song "Spring Mud." The electronic device determines that the wake-up sentence containing the wake-up word is "Bloom the next flower season." The electronic device then determines that the acquisition time of the wake-up sentence "Bloom the next flower season" is from [01:07] to [01:09].

[0139] In addition, based on the lyrics information, the time period in which the lyrics sentence "Blooming the next flower season" in the current song is played can also be determined. For example, the song "Spring Mud" plays "Blooming the next flower season" from [01:06] to [01:08]. In the lyrics information, each lyrics sentence corresponds to a timestamp, and the timestamp is fixed.

[0140] The preset range of lyrics information includes the range of the previous and next verses of the current verse. For example, based on the lyrics information, the corresponding context of the line "Blooming the next flower season" in the song "Spring Mud" is determined to be "Nourishing the earth" and "Your tears in the wind", and the corresponding timestamps are [01:04] and [01:09].

[0141] [01:04]Nourish the earth

[0142] [01:06]The next flower season begins

[0143] [01:09]Your tears in the wind

[0144] In this example, based on the lyrics, the preset range of the lyrics information is determined to be [01:04] to [01:09] seconds, containing a total of three lyrics. Based on the preset algorithm, the wake-up phrase "Blooming the next flower season" acquired between [01:07] and [01:09] seconds is compared and scored for similarity with the three lyrics in the lyrics information at [01:04] and [01:09] seconds, to obtain a second similarity score.

[0145] In this embodiment, in the process of determining the second similarity score, the wake-up sentence containing the wake-up word is compared with the preset range of lyrics information (such as the upper and lower sentences of the lyrics sentence). Since the time period of the preset range includes the upper and lower sentences of the wake-up sentence, it can avoid the system misjudgment caused by the time difference between the rhythm of the lyrics sung by the user and the beat of the song played by the actual audio playback device. This method overcomes the problem that the voice signal of the user singing and the lyrics information of the song are not synchronized, and the time difference causes the system to be mistakenly awakened.

[0146] In addition, this method detects part of the lyrics information, which reduces the detection information, improves the detection efficiency, and saves detection time compared to checking the entire lyrics information.

[0147] In some other optional implementations, based on the above step 207-1, it further includes: based on the lyrics of the song played, determining that the lyrics paragraph containing the lyrics is the preset range of the lyrics information.

[0148] The lyric paragraph is the paragraph containing the awakening sentence, and the lyric paragraph can be pre-divided. For example, lyrics with the same melody are divided into one paragraph. For example, in this example, based on the lyric information, the paragraph containing the awakening sentence "Blooming the next flower season" is determined to be:

[0149] [00:28]Those painful memories

[0150] [01:02]Falling in the spring soil

[0151] [01:04]Nourish the earth

[0152] [01:06]The next flower season begins

[0153] [01:09]Your tears in the wind

[0154] [01:11]Drop by drop, falling into memories

[0155] [01:13] Let's name it "Cherish"

[0156] The electronic device determines that the acquisition time of the wake-up sentence "Bloom out the next flower season" is from [01:07] to [01:09] s, and then compares the similarity with the paragraph where the above "Bloom out the next flower season" is located to obtain a second similarity score.

[0157] Optionally, a similarity comparison and scoring method is to compare and score the wake-up sentence "Bloom the next flower season" with each lyric sentence in the paragraph within the above preset range to obtain a similarity score, and then filter out the one with the largest score as the second similarity score of the sentence, which is used for comparison with the second threshold.

[0158] In this embodiment, the wake-up sentence containing the wake-up word is compared with a preset range of lyrics information (such as the paragraph where the lyrics are located) for similarity, thereby avoiding system misjudgment caused by the time difference between the rhythm of the user singing the lyrics and the beat of the song played by the actual audio playback device. This method improves the accuracy of wake-up word detection.

[0159] In some further optional implementations, based on the above step 207-1, the following steps are further included:

[0160] First, based on the lyrics of the song being played, the lyrics paragraph containing the lyrics is determined.

[0161] Then, similar lyrics paragraphs of the lyrics paragraphs are determined based on semantic analysis, and the similar lyrics paragraphs are determined as a preset range of lyrics information.

[0162] The similar lyrics segment contains the aforementioned wake-up sentence, or includes the aforementioned wake-up sentence and the paragraph containing the preceding and following sentences. Furthermore, similar lyrics segments can be determined by time tags in the lyrics file, and the aforementioned wake-up sentence can be compared and scored with each similar lyrics segment to obtain a second similarity score.

[0163] In this example, based on the wake-up sentence "Bloom in the next flower season", the full text of the lyrics of "Spring Mud" is searched for identical or similar paragraphs. After searching the full text, there are 4 paragraphs similar to "Bloom in the next flower season". These 4 similar paragraphs are used as the preset range of lyrics information, and are compared and scored with "Bloom in the next flower season" respectively, and 4 similarity scores are obtained. The largest one is then selected from these 4 scores as the second similarity score and compared with the second threshold to obtain the detection result.

[0164] In this embodiment, all similar lyrics paragraphs in the song are compared with the wake-up sentences and scored, thereby preventing the system from being woken up by mistake and improving the accuracy of wake-up word detection.

[0165] Exemplary devices

[0166] See also Figure 5 , which is a schematic diagram of the structure of a speech recognition device provided in an embodiment of the present disclosure, which is used to implement all or part of the functions of the aforementioned method embodiment. Specifically, the speech recognition device 500 includes: a lyrics acquisition module 510, a speech acquisition module 520, a speech detection module 530, and a speech processing module 540.

[0167] In addition, the device may also include other modules or units, such as a storage module, a sending module, etc., which are not limited in this embodiment. Furthermore, the speech recognition device may be an electronic device, such as a terminal device.

[0168] The lyrics acquisition module 510 is used to acquire lyrics information corresponding to a song in response to a song playing instruction.

[0169] The voice acquisition module 520 is configured to acquire a voice signal based on a microphone array and detect a wake-up word based on the voice signal.

[0170] The speech detection module 530 is used to detect the wake-up word based on the speech signal obtained by the speech acquisition module 520, determine a first similarity score between the wake-up word and the lyrics information obtained by the lyrics acquisition module 510, and determine that the first detection type of the wake-up word is invalid based on the first similarity score being greater than or equal to a first threshold; and based on the first similarity score being less than the first threshold.

[0171] In addition, the voice detection module 530 is further configured to obtain a wake-up sentence including a wake-up word, and determine that the second detection type of the wake-up word is invalid based on the second similarity score between the wake-up sentence and the lyrics information.

[0172] The voice processing module 540 is used to not respond to the wake-up word based on the first detection type of the wake-up word determined by the voice detection module 530 being invalid, or the second detection type of the wake-up word being invalid, and to respond to the wake-up word based on the second detection type of the wake-up word determined by the voice detection module 530 being valid.

[0173] The voice signal input by the microphone array includes a voice signal corresponding to the audio signal output when the user sings, or also includes: an audio signal played by an audio playback device.

[0174] The wake-up word detected by the voice acquisition module 520 refers to one or more words in the voice signal that can wake up the intelligent voice function of the electronic device. The wake-up word includes but is not limited to the next song, previous song, and previous song. The first similarity score is used to measure the similarity between the wake-up word detected by the voice acquisition module 520 in the previous step and the lyrics of the currently playing song. In this embodiment, the first similarity score can be used to detect whether the first detection type of the wake-up word is valid. In addition, a second similarity score is included to verify whether the second detection type of the wake-up word is valid.

[0175] In one example, the lyrics acquisition module 510 responds to a user's instruction to play the song "Spring Mud" and acquires the lyrics for the song. The voice acquisition module 520 uses a microphone array to capture the user's singing voice signal and detects that the wake-up word is "next." The voice detection module 530 compares and scores the wake-up word "next" with the lyrics acquired by the lyrics acquisition module 510, obtaining a first similarity score. If the first similarity score is greater than or equal to a first threshold, the voice detection module 530 determines that the first detection type for the wake-up word "next" is invalid.

[0176] The voice detection module 530 obtains the wake-up sentence containing the wake-up word "Bloom in the next flower season", compares and scores the wake-up sentence with the lyrics information of "Spring Mud", and obtains a second similarity score. Assuming that the second similarity score is 99 points and the second threshold is 60 points, 99>60, it is determined that the second detection type of the wake-up word "next" is invalid.

[0177] Based on the result that the “next” second detection type detected by the voice detection module 530 is invalid, the voice processing module 540 does not respond to the wake-up word and continues to play the song “Spring Mud”.

[0178] The device provided in this embodiment uses the first similarity score during the wake-up word detection process. When it is determined that the first detection type of the wake-up word is invalid, it is determined that the wake-up word is a lyric of the song currently sung by the user or is highly similar to the lyrics, indicating that the user is singing. At this time, the system does not respond to the wake-up word and continues to play the song, thereby avoiding the system being mistakenly awakened and affecting the user's singing experience.

[0179] In addition, when it is determined that the first detection type of the wake-up word is valid, continue to search for the wake-up sentence (context sentence) where the wake-up word is located, and use the second similarity score to perform similarity judgment with the song lyrics. When it is determined that the second detection type of the wake-up word is invalid, do not respond to the wake-up word and continue to play the song to prevent the system from being woken up by mistake and affecting the user's singing experience.

[0180] In some optional implementations, the speech detection module 530 is further configured to determine the context of the wake-up word based on voice activity detection (VAD) before obtaining the wake-up sentence containing the wake-up word. Furthermore, the speech detection module 530 is specifically configured to obtain the wake-up sentence containing the wake-up word based on the detected context of the wake-up word.

[0181] In one example, based on VAD, it is detected that the current wake-up word "next" has the context sentences "bloom" and "flower season", and the wake-up sentence containing "next" is "bloom the next flower season".

[0182] This implementation method can accurately detect whether the wake-up word contains a context sentence through the VAD method, which can prepare for the similarity judgment of subsequent wake-up sentences and provide a basis for the effectiveness judgment of subsequent wake-up words.

[0183] In some optional implementations, the voice detection module 530 is further used to determine a wake-up statement that does not include a wake-up word based on a context in which the wake-up word is not detected, and to determine that the third detection type of the wake-up word is valid, and respond to the wake-up word.

[0184] In one example, when performing VAD detection on the wake-up word "next song", the VAD result of the wake-up word "next song" is detected to be 0, that is, it does not include context. The voice detection module 530 determines that the third detection type of the wake-up word is valid, wakes up the system, and responds to the voice command of the wake-up word "next song" to play the "next song".

[0185] This implementation can accurately detect whether the wake-up word contains a context sentence through the VAD method, realize the verification of the wake-up word context, and wake up the voice system to respond to the wake-up word when the detection does not include a context sentence.

[0186] In some optional implementations, the voice detection module 530 is specifically configured to score the similarity between the wake-up word and the lyrics information based on a preset algorithm to obtain the first similarity score.

[0187] The preset algorithm is used to score the similarity between the wake-up word and the lyrics information. Based on the preset algorithm, the similarity score of the currently detected wake-up word and the full text of the lyrics information is scored, and then the scores are combined to obtain a first similarity score. The preset algorithm includes but is not limited to simple matching, incomplete matching, order-ignoring matching, duplicate-removing subset matching, fuzzy matching algorithms, etc. It should be understood that the preset algorithm can also be based on a combination of one or more of the above matching methods.

[0188] This implementation uses a preset algorithm to score the similarity between the wake-up word and the lyrics information, obtaining a corresponding similarity score, thereby determining the effectiveness of the current wake-up word. In addition, the preset algorithm can be predefined, increasing the accuracy and flexibility of the similarity comparison.

[0189] In some optional implementations, see Figure 6 As shown, the voice detection module 530 specifically includes:

[0190] Scoring module 5301, configured to score the similarity between the wake-up sentence and the lyrics information based on a preset algorithm to obtain a second similarity score;

[0191] The comparison module 5302 is configured to compare the second similarity score determined by the scoring module 5301 with a preset second threshold.

[0192] Determination module 5303 is used to determine that the second detection type of the wake-up word is invalid when the second similarity score compared by comparison module 5302 is greater than or equal to the second threshold; and to determine that the second detection type of the wake-up word is valid when the second similarity score compared by comparison module 5302 is less than the second threshold.

[0193] In some optional implementations, the comparison module 5302 in the speech detection module 530 is further configured to perform a similarity score on the wake-up sentence and the full text of the lyrics information based on a preset algorithm to determine a second similarity score. The preset algorithm can be described in the aforementioned embodiment and will not be further described here.

[0194] In some optional implementations, the determination module 5303 in the speech detection module 530 is further configured to determine the lyrics of the song being played based on the time the wake-up statement was acquired and the lyrics information; and to determine a preset range of lyrics information based on the lyrics of the song being played. The comparison module 5302 is further configured to score the similarity between the wake-up statement and the preset range of lyrics information based on a preset algorithm to determine a second similarity score.

[0195] In some optional implementations, the determination module 5303 in the voice detection module 530 is also used to determine, based on the lyrics of the song played, the lyrics paragraph containing the lyrics as the preset range of the lyrics information; or, based on the lyrics of the song played, determine the upper and lower sentences of the lyrics as the preset range of the lyrics information.

[0196] In some optional implementations, the determination module 5303 in the voice detection module 530 is also used to determine the lyrics paragraph containing the lyrics sentence based on the lyrics sentence played in the song; determine the similar lyrics paragraphs of the lyrics paragraph based on semantic analysis, and determine the lyrics paragraphs and similar lyrics paragraphs as the preset range of lyrics information.

[0197] The speech recognition device provided by the above-mentioned embodiment of the present disclosure, in the process of determining the second similarity score, performs a similarity comparison between the wake-up sentence containing the wake-up word and a preset range of lyrics information (such as the upper and lower sentences of the lyrics sentence, the paragraph where the lyrics are located, and similar lyrics paragraphs). Since the time period of the preset range includes the upper and lower sentences of the wake-up sentence, it can avoid the system misjudgment caused by the time difference between the rhythm of the lyrics sung by the user and the beat of the song played by the actual audio playback device, thereby overcoming the problem that the voice signal of the user singing and the lyrics information of the song are not synchronized in time, and the time difference causes the system to be mistakenly woken up.

[0198] In addition, in the above implementation, in the process of determining the similarity score, part of the lyrics information (such as the context, the paragraph where the lyrics are located, and similar lyrics paragraphs) is detected. Compared with checking the full text of the lyrics information, this reduces the detection information, improves the detection efficiency, and saves detection time.

[0199] Exemplary electronic devices

[0200] Below, reference Figure 7 The electronic device according to the embodiment of the present disclosure is described below. The electronic device may be any terminal device in the aforementioned embodiment, such as a vehicle-mounted terminal, and is configured to implement the speech recognition method in the aforementioned embodiment.

[0201] Figure 7 A structural diagram of an electronic device according to an embodiment of the present disclosure is illustrated.

[0202] like Figure 7 As shown, the electronic device includes one or more processors 701 and a memory 702 .

[0203] The processor 701 may be a central processing unit (CPU) or other processing unit with data processing capability and / or instruction execution capability, and may control other components in the electronic device to perform desired functions. In addition, the processor 701 may be equipped with multimedia playback application software, such as audio and video applications.

[0204] The memory 702 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 701 may execute the program instructions to implement the speech recognition method of the various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, instructions for controlling the play / pause of media streams, song lyrics databases, etc. may also be stored in the computer-readable storage medium.

[0205] In one example, the electronic device may further include an input device 703 and an output device 704 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0206] For example, when the electronic device is a vehicle-mounted terminal or a vehicle-mounted processor, the input device 703 can be a microphone or a microphone array for collecting input signals from a sound source. In addition, the input device 703 can also be connected to the processor 701 to receive audio and video input signals played by multimedia playback application software.

[0207] In addition, the input device 703 may also include, for example, a keyboard, a mouse, and the like.

[0208] The output device 704 can output various information to the outside. Further, the output device 704 can include, for example, a display / screen, a speaker, a communication network and a remote output device connected thereto, and the like.

[0209] Of course, to simplify, Figure 7 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.

[0210] Exemplary computer program products and computer-readable storage media

[0211] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the speech recognition method according to various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.

[0212] The computer program product may be written in any combination of one or more programming languages ​​to implement the operations of the disclosed embodiments, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, or partially on the user's computing device.

[0213] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enable the processor to execute the steps of the speech recognition method according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0214] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0215] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0216] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0217] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0218] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0219] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A speech recognition method, comprising: In response to a song play instruction, obtaining lyrics information corresponding to the song; Acquire a voice signal based on a microphone array, and detect a wake-up word based on the voice signal; Determining a first similarity score between the wake-up word and the lyrics information; Based on the first similarity score being greater than or equal to a first threshold, determining that the first detection type of the wake-up word is invalid, and not responding to the wake-up word; Based on the first similarity score being less than the first threshold, obtaining a wake-up sentence including the wake-up word; Scoring the similarity between the wake-up statement and the lyrics information based on a preset algorithm to obtain a second similarity score between the wake-up statement and the lyrics information; In response to the second similarity score being greater than or equal to a second threshold, determining that the second detection type of the wake-up word is invalid, and not responding to the wake-up word; In response to the second similarity score being less than the second threshold, determining that the second detection type of the wake-up word is valid, and responding to the wake-up word.

2. The method according to claim 1, wherein Based on the first similarity score being less than a first threshold, before obtaining the wake-up sentence including the wake-up word, the method further includes: determining a context of the wake word based on voice activity detection; Obtaining a wake-up sentence containing the wake-up word includes: Based on the detected context of the wake-up word, a wake-up sentence including the wake-up word is obtained.

3. The method according to claim 2, wherein: The determining the context of the wake-up word based on voice activity detection further includes: Based on the context in which the wake-up word is not detected, determining a wake-up statement that does not include the wake-up word, determining that the third detection type of the wake-up word is valid, and responding to the wake-up word.

4. The method according to any one of claims 1 to 3, wherein: Determining a first similarity score between the wake-up word and the lyrics information includes: The similarity between the wake-up word and the lyrics information is scored based on a preset algorithm to obtain the first similarity score.

5. The method according to any one of claims 1 to 3, wherein: The performing similarity scoring on the wake-up sentence and the lyrics information based on a preset algorithm to obtain the second similarity score includes: The wake-up statement and the full text of the lyrics information are scored for similarity based on a preset algorithm to determine the second similarity score.

6. The method according to any one of claims 1 to 3, wherein: The performing similarity scoring on the wake-up sentence and the lyrics information based on a preset algorithm to obtain the second similarity score includes: Determining the lyrics of the song to be played based on the acquisition time of the wake-up sentence and the lyrics information; Determining a preset range of the lyrics information based on the lyrics of the song played; The similarity score is performed on the wake-up statement and a preset range of the lyrics information based on a preset algorithm to determine the second similarity score.

7. The method according to claim 6, wherein: The determining of the preset range of the lyrics information based on the lyrics of the song played includes: Based on the lyrics of the song played, determining the lyrics paragraph containing the lyrics as the preset range of the lyrics information; Alternatively, based on the lyrics of the song being played, the upper and lower sentences of the lyrics are determined to be the preset range of the lyrics information.

8. The method according to claim 6, wherein: The determining of the preset range of the lyrics information based on the lyrics of the song played includes: Based on the lyrics of the song played, determining the lyrics paragraph containing the lyrics; Based on semantic analysis, similar lyrics paragraphs of the lyrics paragraph are determined, and the lyrics paragraph and the similar lyrics paragraphs are determined to be a preset range of the lyrics information.

9. A speech recognition device, comprising: Lyrics acquisition module, used for obtaining lyrics information corresponding to a song in response to a song playing instruction; A voice acquisition module is used to acquire voice signals based on a microphone array; a speech detection module, configured to detect a wake-up word based on the speech signal acquired by the speech acquisition module, determine a first similarity score between the wake-up word and the lyrics information acquired by the lyrics acquisition module, and, based on the first similarity score being greater than or equal to a first threshold, determine that a first detection type of the wake-up word is invalid; and based on the first similarity score being less than the first threshold, acquire a wake-up sentence containing the wake-up word; Scoring the similarity between the wake-up sentence and the lyrics information based on a preset algorithm to obtain a second similarity score between the wake-up sentence and the lyrics; and determining that the second detection type of the wake-up word is invalid in response to the second similarity score being greater than or equal to a second threshold; In response to the second similarity score being less than the second threshold, determining that the second detection type of the wake-up word is valid; A voice processing module is used to not respond to the wake-up word in response to the voice detection module determining that the first detection type of the wake-up word is invalid, or that the second detection type of the wake-up word is invalid, and to respond to the wake-up word in response to determining that the second detection type of the wake-up word is valid.

10. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the speech recognition method according to any one of claims 1 to 8.

11. A speech recognition system, comprising a speech playback device, a speech collection device, and a speech recognition device; in, The voice playback device is used to respond to a song playback instruction and play the audio and video corresponding to the instruction; The voice collection device is used to collect the voice signal input by the user; The speech recognition device is used to call computer program instructions stored in a memory based on the audio and video played by the speech playback device and the speech signal collected by the speech collection device, and execute the instructions to implement the speech recognition method described in any one of claims 1 to 8 above.

12. A computer-readable storage medium, wherein a computer program is stored in the storage medium, and the computer program is used to execute the speech recognition method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and apparatus for preventing voice command misidentification

    CN106409294A

  • Voice wake-up method, device and system, equipment, server and storage medium

    CN109378000A