Voice signal response method and apparatus, device, medium, and product
By distinguishing between human and non-human voices, the response type is determined, thus solving the problem of false wake-up of voice assistants and improving the reliability of voice signal response and user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-03-26
AI Technical Summary
In existing technologies, voice assistants or voice robots are easily disturbed by non-human voices in the surrounding environment, leading to false wake-ups, wasting computing resources and causing unnecessary disturbance to users. How can we improve the reliability of voice signal response?
By acquiring speech signals and determining their vocal type, the system distinguishes between human and non-human voices, determines the response type based on the vocal type (including positive or negative responses), utilizes the acoustic characteristics of the low-frequency portion and speech processing models for identification, and shields against interference from non-human voices by dividing the sound source region.
It improves the reliability of response to speech signals containing human voice components, reduces false responses, saves computing resources, and enhances the user experience.
Smart Images

Figure CN2025122749_26032026_PF_FP_ABST
Abstract
Description
Response method, device, equipment, medium and product of voice signal
[0001] Cross-reference to related applications
[0002] The present application is based on and claims priority to Chinese patent application No. 202411312014.7, filed on September 19, 2024, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0003] Embodiments of the present application relate to the technical field of voice processing, in particular to a response method, device, equipment, medium and product of a voice signal. BACKGROUND
[0004] Artificial intelligence (AI) voice assistants or voice robots are increasingly widely used in computers, mobile phones, and vehicle-mounted fields. In general, a user can wake up a device or interact with an AI voice assistant or voice robot by speaking a fixed wake-up word, and then the AI voice assistant or voice robot can respond to the user's voice instruction based on a deep learning algorithm application, accurately identify the user's intent, and thus realize automatic answering, self-service business handling, and intelligent control of the device, etc.
[0005] This way of interacting through a fixed wake-up word has the following disadvantages: the user himself / herself does not speak the fixed wake-up word and does not want to interact with the device, AI voice assistant or voice robot, but there are some disturbances in the surrounding environment, such as a nearby TV or mobile phone playing a voice containing the fixed wake-up word, or emitting a voice similar to human voice and relatively similar in pronunciation to the fixed wake-up word, which causes the device to be mistakenly woken up or the AI voice assistant or voice robot to perform unnecessary interactions, etc., making an incorrect response to the voice signal, wasting the computing resources of the device, and causing unnecessary disturbance to the user. In view of the above incorrect response situation, how to improve the reliability of the response to the voice signal has become a problem to be solved. SUMMARY
[0006] The present application provides a response method, device, equipment, medium and product of a voice signal to improve the reliability of the response to a voice signal containing a human voice component.
[0007] In a first aspect, the embodiments of the present application provide a response method of a voice signal, comprising:
[0008] obtaining a voice signal, wherein the voice signal includes a human voice component;
[0009] determine a vocalization type of the voice signal, the vocalization type comprising a real human vocalization or a non-real human vocalization;
[0010] determine a response type of the voice signal according to the vocalization type of the voice signal.
[0011] In a second aspect, an embodiment of the present application further provides a response device of a voice signal, comprising:
[0012] an obtaining module configured to obtain a voice signal, the voice signal comprising a human voice component;
[0013] a type determining module configured to determine a vocalization type of the voice signal, the vocalization type comprising a real human vocalization or a non-real human vocalization;
[0014] a response module configured to determine a response type of the voice signal according to the vocalization type of the voice signal.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0016] one or more processors;
[0017] a storage device configured to store one or more programs;
[0018] when the one or more programs are executed by the one or more processors, the one or more processors implement the response method of the voice signal according to the first aspect.
[0019] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the response method of the voice signal according to the first aspect.
[0020] In a fifth aspect, an embodiment of the present application further provides a computer program product, which comprises a computer program and / or instructions, and the computer program and / or instructions are executed by a processor to implement the response method of the voice signal according to any of the above embodiments.
[0021] The embodiments of the present application provide a response method, device, equipment, medium and product of a voice signal, the response method of the voice signal comprising: obtaining a voice signal, the voice signal comprising a human voice component; determining a vocalization type of the voice signal, the vocalization type comprising a real human vocalization or a non-real human vocalization; determining a response type of the voice signal according to the vocalization type of the voice signal. The above technical solution can make a corresponding type of response to the voice signal according to the discrimination of whether the human voice component in the voice signal is a real human vocalization, can make a suitable and flexible response to the voice signal of the real human vocalization and the non-real human vocalization, and improve the reliability of the response to the voice signal comprising the human voice component. BRIEF DESCRIPTION OF DRAWINGS
[0022] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which: like reference numerals refer to like elements throughout. The drawings are intended to be illustrative, and are not necessarily drawn to scale, but are intended to conceptually illustrate the features of the present disclosure.
[0023] FIG. 1 is a flowchart of a method for responding to a voice signal according to an embodiment of the present disclosure;
[0024] FIG. 2 is a flowchart of another method for responding to a voice signal according to an embodiment of the present disclosure;
[0025] FIG. 3 is a schematic diagram for determining a response type according to a position of a wake-up word according to an embodiment of the present disclosure;
[0026] FIG. 4 is a schematic diagram of a response device for a voice signal according to an embodiment of the present disclosure;
[0027] FIG. 5 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] The present application will be further described by examples in conjunction with the drawings. It is to be understood that the following examples are merely illustrative of the present application and should not be construed as limiting the present application in any way. In addition, it should be understood that the drawings are not necessarily to scale and in some instances proportions have been exaggerated in order to more clearly reveal details of a structure. Furthermore, elements of the drawings can be omitted or simplified in order not to obscure the concepts of the present application.
[0029] Before the exemplary embodiments are described in greater detail, it is noted that some exemplary embodiments are described as processes depicted as flowcharts or methods. Although the processes are described in a particular sequential order, many of the processes can be performed concurrently, in parallel, or simultaneously. In addition, the order of individual processes can be re-arranged. The processes can be terminated when their operations are completed, but the processes can also have additional steps not included in the figure(s) which can nevertheless be performed. The processes can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0030] It should be noted that the terms "first", "second", and the like, used in the description and in the claims of the present application are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances such that the descriptive terms "first", "second", etc. are to be interpreted, by those skilled in the art, as a structural or functional pertinence and not by their order or sequence.
[0031] For the convenience of understanding, the processing and responding process of the voice signal are described by taking Chinese as an example in the embodiments of the present application, and other languages can be used in actual application.
[0032] Figure 1 is a flowchart of a voice signal response method provided in an embodiment of this application. This embodiment is applicable to situations involving the processing and response to voice signals. Specifically, the voice signal response method can be executed by a voice signal response device, which can be implemented through software and / or hardware and integrated into an electronic device. The electronic device can refer to a user-end device or a server-end device, including but not limited to in-vehicle systems, electronic control units (ECUs), computers, in-vehicle intelligent terminals, smartphones, or cloud servers.
[0033] As shown in Figure 1, the method specifically includes the following steps:
[0034] S110, Acquire voice signal.
[0035] In the embodiments of this application, acquiring speech signals mainly refers to using a microphone or pickup to collect surrounding speech signals in real time. The speech signal contains human voice components, which may be the real-time voice of a person in the surrounding environment, the sound played by other devices in the surrounding environment, or noise similar to human voices in the surrounding environment.
[0036] S120. Determine the vocal type of the speech signal.
[0037] Optionally, the speech signal can be classified as either human voice or non-human voice. If the human voice component in the speech signal is a real sound emitted by a human in the surrounding environment, the speech type is human voice. If the human voice component in the speech signal is not a real sound emitted by a human in the surrounding environment, such as sound played by a device or noise, the speech type is non-human voice.
[0038] In some embodiments, acoustic features of a speech signal can be determined, and the speech type of the speech signal can be determined based on the acoustic features. The acoustic features may include at least one of features such as energy intensity, spectrum, audio fluctuation amplitude, pitch, and speech rate.
[0039] For example, determining the type of voice signal can be achieved by spectral analysis, such as analyzing the frequency spectrum of the voice signal using audio editing software or a spectrum analyzer, and identifying the type of voice signal as human voice or non-human voice according to the spectral characteristics; or it can also be achieved by analyzing the energy distribution of the voice signal, such as human voice usually has high energy, while non-human voice has low energy, and the type of voice signal can be identified as human voice or non-human voice according to the energy intensity of the voice signal; in addition, it can also be achieved according to the tone, audio, speed and / or noise of the vocal component, etc. to distinguish human voice and non-human voice; or it can also be achieved by AI model, machine learning (ML) or deep learning (DL) technology, etc.
[0040] S130, determining the response type of the voice signal according to the type of voice signal.
[0041] Optionally, the response type can be understood as the response mode of the voice signal, which can be a positive response or a negative response, and the positive response can specifically include performing one or more response operations or prompting the user that the voice signal is responded to, etc., and the negative response can specifically be no response (not performing any operation) or prompting the user that the voice signal is not responded to, etc.; the response type can also refer to fast response (such as immediate response) within a specified time (such as 1 second) or delayed response after a specified time (such as 1 second); the response type can also refer to positive response or faster response (faster response means responding within a specified time) when a predetermined condition is met, or negative response or slower response (slower response means responding after a specified time) when the predetermined condition is not met, etc., wherein the predetermined condition is, for example, determining that the voice signal should (or needs to be) responded to by analyzing the semantics of the voice signal, or the voice signal contains a specified wake-up word, or the specified wake-up word is in a specified position in the voice signal, or the interval between the wake-up word and the text before and / or after it does not exceed a predetermined threshold (such as 500 milliseconds), etc.
[0042] Generally, when the type of voice signal is human voice, the response type can be a positive response or a faster response, etc. to respond positively; when the type of voice signal is non-human voice, the response type can be a negative response or a slower response, etc. to respond negatively.
[0043] Correspondingly, in some embodiments, determining the response type of the voice signal can include:
[0044] determining the response type as a first response type when the preset condition is met;
[0045] determining the response type as a second response type when the preset condition is not met;
[0046] The first response type comprises at least one of the following:
[0047] responding within the specified duration;
[0048] positively responding;
[0049] The second response type comprises at least one of the following:
[0050] responding after the specified duration;
[0051] negatively responding.
[0052] In some embodiments, the preset condition comprises at least one of the following:
[0053] determining that the voice signal needs to be responded according to the semantics of the voice signal;
[0054] the voice signal contains a wake-up word;
[0055] the wake-up word is at a specified position in the voice signal;
[0056] the time interval between the wake-up word and the text before and / or after the wake-up word does not exceed a preset threshold.
[0057] The specific descriptions of the positive response and the negative response will be described in the following embodiments, and will not be repeated here.
[0058] The response method for the voice signal provided by the embodiments of the present application can determine whether the human voice component in the voice signal is real human voice, and make a response of a corresponding type accordingly. For real human voice, a positive response can be made. For voice signals of non-real human voice, although the voice signals contain human voice, the contained human voice is the voice broadcasted by other devices around, or the noise around is similar to the human voice, and the user actually has no intention of interaction. In this case, the false response to the non-real human voice can be avoided. By making appropriate and flexible responses to voice signals of real human voice and non-real human voice, the reliability of the response to voice signals containing human voice components can be improved.
[0059] In an embodiment, the voice signal type can be determined at a set frequency (such as every 300 milliseconds), so that real human voice and non-real human voice can be distinguished in time, and the response speed to the voice signal can be improved.
[0060] In an embodiment, determining the voice signal type comprises:
[0061] S1210, determine the acoustic feature of the low frequency part in the voice signal;
[0062] S1220, determine the sound type according to the acoustic feature.
[0063] For example, the acoustic features of real human voice and non-real human voice in the low frequency part are obviously different. Generally, the acoustic feature of non-real human voice in the low frequency part is very fuzzy, while the acoustic feature of real human voice in the low frequency part is relatively full. The acoustic feature of the low frequency part in the voice signal is obtained through spectrum analysis, so that the sound type can be determined according to the acoustic feature of the low frequency part.
[0064] For example, the acoustic feature of the low frequency part in the voice signal is determined, including: determining the energy intensity of the low frequency part in the voice signal; determining the sound type according to the acoustic feature, including: determining the sound type of the voice signal according to the energy intensity. For example, the energy distribution of the voice signal can be analyzed to obtain the energy intensity of the voice signal, and the sound type of the voice signal is determined according to the energy intensity of the voice signal. For example, in the case that the energy intensity of the voice signal is greater than a preset energy intensity threshold, the sound type of the voice signal is determined as real human voice, and in the case that the energy intensity of the voice signal is not greater than the preset energy intensity threshold, the sound type of the voice signal is determined as non-real human voice.
[0065] It can be understood that the low frequency part can refer to a frequency band in the frequency band involved in the voice signal, for example, a frequency band below the middle frequency of the frequency band involved, or a frequency band with the lowest 30% of the frequency in the frequency band involved, or the frequency band of the human voice is approximately between 80Hz (Hertz) and 8000Hz. The low frequency part can be a frequency band with a frequency range of 80Hz-200Hz, or other frequency band range. The embodiments of the present application do not limit the frequency band range of the low frequency part.
[0066] In the embodiments of the present application, the difference between real human voice and non-real human voice in the low frequency part is used to distinguish whether the voice signal is real human voice. Compared with using other frequency bands or the entire frequency band of the voice signal, it has higher accuracy, and provides a reliable basis for determining the response type according to the sound type.
[0067] In an embodiment, the sound type of the voice signal is determined, including:
[0068] S1230, input the voice signal into the trained voice processing model to determine the sound type of the voice signal through the voice processing model; wherein the voice processing model is trained based on real human voice sample data and non-real human voice sample data.
[0069] In an embodiment of the present application, the voice processing model can be an AI, machine learning (ML) or deep learning (DL) model. The voice processing model is trained based on real human voice sample data and non-real human voice sample data. The real human voice sample data and non-real human voice sample data can be obtained by collecting real data in various scenes in life. The voice processing model can be trained using the real human voice sample data and non-real human voice sample data, so that the voice processing model can analyze and infer the characteristics of a large amount of real human voice sample data and non-real human voice sample data, thereby having the ability to distinguish whether a voice signal is a real human voice or a non-real human voice. In actual application, the obtained voice signal (after vectorization) is input into the voice processing model, and the output obtained is the recognition result, for example, an output of 0 represents a non-real human voice, and an output of 1 represents a real human voice.
[0070] In an embodiment of the present application, the voice processing model determines the voice type, the voice processing model can automatically extract key features from the voice signal without manual design, and has strong data fitting ability, which can avoid errors in manual processing of features, reduce the time-consuming of processing features, and accurately and efficiently determine the voice type of the voice signal.
[0071] In an embodiment, the training process of the voice processing model includes:
[0072] Obtaining real human voice sample data and non-real human voice sample data;
[0073] According to the acoustic characteristics of the low frequency part in the real human voice sample data and the non-real human voice sample data, the real human voice sample data and the non-real human voice sample data are labeled;
[0074] Training the voice processing model based on the real human voice sample data and the non-real human voice sample data.
[0075] Optionally, the real human voice sample data and the non-real human voice sample data can be obtained by collecting real data in various scenes in life, for example, a large number of voice signals of real human voices such as drivers and passengers in vehicles, voice signals of playing human voices in vehicles using various devices, and voice signals containing noises in the environment around the vehicle, etc. can be collected. These voice signals are used as sample data, and according to the acoustic characteristics of the low frequency part, the sample data is labeled, that is, a label is added, which is used to identify which sample data is a real human voice and which sample data is a non-real human voice. Then, the voice processing model is trained using these sample data, so that it has the ability to distinguish whether a voice signal is a real human voice or a non-real human voice.
[0076] On this basis, the low-frequency part is used to distinguish between real human voice and non-real human voice by combining a voice processing model, which can avoid errors in manual feature processing and reduce the time required for processing features. In addition, the low-frequency part is used to distinguish between real human voice and non-real human voice, which has higher accuracy than using other frequency bands or the entire frequency band of the voice signal, thereby improving the efficiency and reliability of determining the voice type and response type.
[0077] In an embodiment, the voice signal is determined to have a voice type, comprising:
[0078] S1240, determining the sound source area corresponding to the voice signal;
[0079] S1250, in the case where there are at least two sound source areas corresponding to the voice signal, determining whether the voice type of the voice signal corresponding to each sound source area is real human voice;
[0080] S1260, shielding the voice signal of the sound source area corresponding to the voice type of non-real human voice, and retaining the voice signal of the sound source area corresponding to the voice type of real human voice.
[0081] S1270, determining the voice type of the voice signal based on the retained voice signal of the sound source area.
[0082] Optionally, the sound source area can be an area divided from the surrounding environment. The obtained voice signal can be one or a combination of multiple voice signals, and the multiple voice signals can come from one or more sound source areas. In some embodiments, in the case where the obtained voice signal is composed of multiple voice signals, for each voice signal, the sound source area corresponding to the voice signal can be determined.
[0083] Determining the sound source area corresponding to the voice signal can be, for example, by simultaneously collecting the voice signal by at least two microphones. For one of the voice signals, the sound source area of the voice signal can be determined according to the difference between the signals collected by the two microphones.
[0084] Alternatively, the sound source area corresponding to the voice signal can also be determined by using an AI / ML / DL model, etc. For example, a model is trained with sample data of different sound source areas, so that it can analyze and infer the features of the voice signal and output the sound source area corresponding to the voice signal.
[0085] The embodiments of the present application do not limit the specific ways and algorithms for determining the sound source area corresponding to the voice signal.
[0086] In the case that the voice signal comes from at least two sound source regions, for each sound source region, it can be determined (by spectrum analysis, energy analysis or voice processing model, etc.) whether the voice signal of the sound source region is human voice, and in the case that it is human voice, the voice signal of the sound source region is retained, and in the case that it is not human voice, the voice signal of the sound source region is shielded. Then, the voice signal can be determined based on the retained voice signal of the sound source region.
[0087] In some embodiments, for the retained voice signal of the sound source region, the determination of the voice type can be performed in the manner as described in the foregoing embodiments, which will not be repeated here.
[0088] For example, the interior of the vehicle can be divided into a main driver region, a co-driver region and a rear region. In the main driver region, the main driver can have the intention to interact and speak out a certain instruction, while in the co-driver region and the rear region, the passengers sitting in the co-driver region and the rear region can be playing audio with their mobile phones. In this case, the voice signals from the co-driver region and the rear region can be shielded, and only the voice signal from the main driver region can be retained for subsequent processing. On this basis, the noise interference of the sound source region with non-human voice can be excluded, the quality of the voice signal with human voice can be improved, and the accuracy of the semantic analysis of the voice signal with human voice can be further improved, and the reliability of the response to the voice signal containing human voice component can be improved.
[0089] It can be understood that in the case that there is only one sound source region, i.e. the surrounding environment is regarded as a whole, it is only necessary to determine whether each voice signal is human voice and the response type.
[0090] In the embodiments of the present application, by dividing the sound source region and shielding the voice signal of the sound source region with non-human voice, the interference of the sound source region with non-human voice can be excluded according to the spatial position, the influence of the non-human voice and the human voice on the determination of whether there is human voice can be avoided, the quality of the voice signal with human voice can be improved, the processing amount of the signals or features other than the human voice can be reduced, the accuracy of the determination of the voice type can be improved, and on this basis, the process of determining the response type can be focused on the region with human voice, and the reliability of the determination of the response type can be improved.
[0091] FIG. 2 is a flowchart of another voice signal response method provided by the embodiments of the present application. As shown in FIG. 2, the method considers the sound source region, and the response process of the voice signal includes:
[0092] S1, obtaining a voice signal.
[0093] S2, determine the sound source region corresponding to the voice signal.
[0094] S3, is there only one sound source region? If yes, then perform S4; otherwise, perform S5.
[0095] S4, determine whether the voice signal is real human voice? If yes, then perform S10 (e.g., determine as a positive response), or perform S11; otherwise, perform S10 (e.g., determine as a negative response).
[0096] S5, determine whether the current sound source region is real human voice? If yes, then perform S7; otherwise, perform S6.
[0097] S6, mask the voice signal of the sound source region.
[0098] S7, keep the voice signal of the sound source region.
[0099] S8, are there still undetermined sound source regions? If yes, then perform S9; otherwise, perform S11.
[0100] S9, take one undetermined sound source region as the current sound source region, and return to S5.
[0101] S10, determine the response type.
[0102] S11, determine the response type according to the position of the wake-up word, and / or the valid human voice component.
[0103] In an embodiment, determining the response type of the voice signal according to the voice type of the voice signal comprises: in a case where the voice type of the voice signal is real human voice, determining the response type of the voice signal according to the position of the wake-up word in the voice signal, and / or the valid human voice component in the voice signal except the wake-up word.
[0104] For example, in a case where it is determined that the voice signal is real human voice, it can be further determined how to respond to the voice signal.
[0105] For example, the response type of the voice signal can be determined according to the position of the wake-up word in the voice signal.
[0106] Taking a speech signal containing ten words as an example, the wake-up word can be represented as "ABC" (assuming that it contains three words), and the words other than the wake-up word can be represented as "x". The wake-up word can be located at the beginning of the speech signal or at the beginning of a sentence, such as "ABCxxxxxxx"; the wake-up word can also be located at the end of the speech signal or at the end of a sentence, such as "xxxxxxx ABC"; the wake-up word can also be located in the middle part of the speech signal (neither at the beginning nor at the end of a sentence), such as "xxABCxxxxx"; in addition, the wake-up word can also appear multiple times, for example, appearing at the beginning and the end of a sentence at the same time, such as "ABCxxxxABC", and the like; the multiple appearances of the wake-up word can be consecutive, such as "ABCABCxxxx", or non-consecutive, such as "xxABCxxABC". It can be understood that the position of the wake-up word can also be represented by the frame number in the speech signal or by the time corresponding to each word.
[0107] Different positions of the wake-up word can have different response types. For example, if the wake-up word is located in the middle part of the speech signal, a negative response can be made, such as rejecting the speech signal or indicating to the user that no operation will be performed, and the like; if the wake-up word is located at the beginning of the speech signal, a positive response can be made, such as waking up the device or performing an operation of opening a window according to the semantics of the speech signal, and the like; if the wake-up word is located at the end of the speech signal, a negative response can also be made. Optionally, in the case where the wake-up word appears multiple times in the speech signal, it can also be considered that the user has a high probability of interacting with the electronic device, and thus a positive response can be made.
[0108] Alternatively, the response type of the speech signal can also be determined according to the valid human voice component other than the wake-up word in the speech signal. The valid human voice component can be understood as the speech with substantive meaning spoken by the user when interacting with the electronic device, which can be an instruction, or a question, or a sentence with complete semantics and a short interval between the wake-up word, and the like. In general, it will not be an interjection, an exclamation, or a word with pronunciation not conforming to the language standard, and the like. In the case where the speech located before the wake-up word in the speech signal is a valid human voice component, the response type can be directly determined (such as a positive response), or the response type can be further determined according to the semantics thereof; in the case where the speech located before the wake-up word in the speech signal is an invalid human voice component (such as an exclamation, and the like), its effect is similar to that of the wake-up word located at the beginning of a sentence, and a positive response can be made.
[0109] In addition, the response type of the voice signal can also be determined in combination with the above two factors. FIG. 3 is a schematic diagram of determining the response type according to the position of the wake-up word provided in an embodiment of the present application. As shown in FIG. 3, for example, if the wake-up word is located in the middle part of the voice signal, a negative response can be made; if the wake-up word is located at the beginning of the voice signal, a positive response is made; and if the wake-up word is located at the end of the voice signal, the response type can be further determined according to whether the voice before the wake-up word is valid human voice component and the semantics of the valid human voice component.
[0110] Optionally, the response speed or response time can also be determined according to the position of the wake-up word in the voice signal and / or the valid human voice component in the voice signal except the wake-up word. For example, in the case where the wake-up word is located in the middle part of the voice signal, a negative response is made more quickly or earlier; in the case where the wake-up word is located at the beginning of the voice signal, a positive response is made more quickly or earlier; and in the case where the wake-up word is located at the end of the voice signal, the response can be made more slowly or later, and the response type can be further determined according to whether the voice before the wake-up word is valid human voice component and the semantics of the valid human voice component, or whether the user continues to give new voice content in the next short period of time, so as to make a more reliable response.
[0111] In the embodiments of the present application, for the voice signal of real human voice, from the perspective of sentence or sentence structure, the position of the wake-up word can be used to quickly determine whether the voice signal is interactive content, from the perspective of semantics, the valid human voice component can be analyzed to determine whether the user has interactive intention, and on this basis, the voice signal of real human voice can be further determined whether it has interactive intention, so as to determine the response type, and improve the flexibility and reliability of the response to the voice signal of real human voice.
[0112] In an embodiment, the response type of the voice signal is determined according to the voice type of the voice signal, comprising:
[0113] In the case where the voice type of the voice signal is real human voice, the voice signal is positively responded to;
[0114] In the case where the voice type of the voice signal is not real human voice, the voice signal is negatively responded to;
[0115] Optionally, the positive response comprises at least one of the following:
[0116] 1) waking up the electronic device;
[0117] 2) waking up the voice assistant;
[0118] 3) generating the first information in the form of a pop-up window by a message notification component (Toast), the first information being used to prompt the user that the voice signal is responded, for example, displaying "wakeup success", "wakeup", semantic analysis result and / or execution result of corresponding operation in the pop-up window, etc.
[0119] 4) performing corresponding operation according to the result of semantic analysis, for example, executing the instruction of the user (such as opening the window of the car or playing a song, etc.);
[0120] Optionally, the negative response includes at least one of the following:
[0121] 1) not waking up the electronic device;
[0122] 2) not waking up the voice assistant;
[0123] 3) generating the second information in the form of a pop-up window by a message notification component (Toast), the second information being used to prompt the user that the voice signal is not responded, for example, displaying "wakeup failure", "not wakeup", "no identifiable semantics" in the pop-up window, etc.
[0124] 4) refusing to recognize the voice signal, that is, not performing semantic analysis on the voice signal, in which case, the electronic device and / or the voice assistant can be already woken up or not.
[0125] For example, in the case that the electronic device is already woken up and the voice assistant is not woken up, the response device of the voice signal acquires the voice signal, the voice type of the voice signal is non-human voice, and the voice signal includes the wakeup word, then the voice assistant can not be woken up, and the second information is displayed in the form of a pop-up window by a message notification component (Toast).
[0126] It should be noted that in the embodiments of the present application, the response type mainly includes positive response and negative response. The response type can be determined for the wakeup service (including waking up the electronic device and / or the voice assistant) and the voice recognition service (mainly including recognizing the semantics in the voice signal), and the response type is determined independently for the two services and does not affect each other. For example, whether to wake up the electronic device can be determined according to whether the voice signal is human voice, the electronic device can be woken up when it is detected that the voice signal contains human voice, whether to perform voice recognition can be determined according to whether the voice signal is human voice, etc. It can also be understood that the step of determining the response type in the embodiments of the present application can be performed before the wakeup service, or can be performed after the wakeup service and before the voice recognition service.
[0127] In the embodiments of the present application, the response type is divided into positive response and negative response, and both the positive response and the negative response include one or more forms of response, on the basis of which, different responses can be flexibly made for the human voice signal and the non-human voice signal; in addition, the different forms of response can correspond to different services (mainly including wake-up service and semantic analysis service), that is, the response method of the voice signal can be flexibly applied to different services or stages, unnecessary wake-up and / or semantic analysis are avoided, the resources of the electronic device are saved, and unnecessary disturbance caused by false response to the user can also be avoided; and the lightweight prompt information can be generated through the toast, the occupation of the screen or the computing resources of the electronic device can be reduced, the method is flexible, simple and efficient, and especially for the scene of vehicle-mounted service and application, the user experience is improved.
[0128] The response method of the voice signal provided in the embodiments of the present application distinguishes human voice and non-human voice by using the acoustic characteristics of the low-frequency part, and can also distinguish them in combination with the voice processing model, which can improve the accuracy of distinguishing whether it is human voice, and further improve the reliability of the response to the voice signal; by distinguishing the voice signal of each sound source region, the non-human voice signal can be shielded, the interference of the sound source region of the non-human voice is excluded, and the quality of the human voice signal is improved; for the human voice signal, according to the position of the wake-up word and / or the effective human voice component, whether the user has an interactive intention can be flexibly and accurately judged, so as to reduce false wake-up and improve the reliability of the response to the voice signal; in addition, the response method of the voice signal is flexibly applied to different services or stages, and has the advantages of flexibility, simplicity and efficiency.
[0129] In some embodiments, determining the response type of the voice signal according to the voice type of the voice signal can include:
[0130] In the case where the voice type of the voice signal is non-human voice, detecting whether the voice assistant has been woken up;
[0131] In the case where the voice assistant has been woken up, obtaining a voice recognition result corresponding to the voice signal;
[0132] Displaying the voice recognition result in a preset form and obtaining a display duration of the voice recognition result;
[0133] Stopping displaying the voice recognition result after the display duration reaches a preset duration.
[0134] The preset duration can be set as needed, and the present application does not limit it. For example, it can be 0.5 seconds, 1 second, etc.
[0135] The preset form can include the color of the display, the size of the display, etc., and the present application does not limit it.
[0136] For example, in the case where the voice signal is acquired by the voice signal response apparatus after the voice assistant has been woken up, and the voice signal is non-human voice, the voice recognition result corresponding to the voice signal can be acquired, and the voice recognition result can be displayed in gray. After the display time of the voice recognition result reaches 0.5 seconds, the voice recognition result can be stopped from being displayed.
[0137] It should be noted that in the above case, the voice recognition result can be stopped from being displayed only after being displayed for a preset time, without subsequent semantic analysis and the like.
[0138] In this way, on the basis of avoiding the misresponse of the voice signal of non-human voice, the user can be informed that the voice signal of non-human voice is acquired by the voice signal response apparatus, and the user can be informed of the text content of the voice signal.
[0139] In some embodiments, a sound suppression switch can be preset in the electronic device, and the voice signal response apparatus can detect whether the sound suppression switch is in an open state. In the case where it is determined that the sound suppression switch is in the open state, after the voice signal is acquired, the voice signal response apparatus can determine the voice type of the voice signal in the manner described in the foregoing embodiments, and determine the response type of the voice signal according to the determined voice type. In the case where it is determined that the sound suppression switch is not in the open state, after the voice signal is acquired, the voice signal response apparatus can default the voice type of the voice signal to human voice, and perform the subsequent response process.
[0140] Correspondingly, determining the response type of the voice signal according to the voice type of the voice signal can include:
[0141] In the case where the sound suppression switch is in the open state, determining the response type of the voice signal according to the voice type of the voice signal.
[0142] In some embodiments, the user can set the sound suppression switch to be in the open state or not in the open state (i.e., in the closed state).
[0143] Correspondingly, the voice signal response method can further include:
[0144] Controlling the state of the sound suppression switch based on the control operation of the user on the sound suppression switch.
[0145] The control operation can include a touch type control operation, a voice type control operation, and the like, which are not limited in the present application.
[0146] Exemplarily, in a case where the control operation is a touch type control operation, a sound suppression switch can be displayed on the electronic device, and based on a touch type control operation of the user on the displayed sound suppression switch, the sound suppression switch is controlled to be in an open state or a closed state.
[0147] Exemplarily, in a case where the control operation is a voice type control operation, another voice signal of the user (which can be different from the voice signal in the foregoing embodiments, and for the sake of distinction, the voice signal is referred to as another voice signal in the embodiments of the present application) can be acquired, and it is detected whether the another voice signal includes a control instruction for controlling the sound suppression switch, and in a case where the another voice signal includes the control instruction, the sound suppression switch is controlled to be in an open state or a closed state based on the control instruction.
[0148] It should be noted that the manner of controlling the sound suppression switch based on the voice type control operation can be applicable to a case where the sound type is real person sound type, and can also be applicable to a case where the sound type is non-real person sound type.
[0149] Correspondingly, the state of the sound suppression switch can include:
[0150] Acquiring another voice signal of the non-real person sound type;
[0151] Detecting whether the another voice signal includes a control instruction for controlling the sound suppression switch;
[0152] In a case where the another voice signal includes the control instruction, controlling the state of the sound suppression switch based on the control instruction.
[0153] For example, in a case where the sound suppression switch is in a closed state, the voice signal responsive device acquires another voice signal, determines that the sound type of the another voice signal is a non-real person sound type, and the another voice signal includes an instruction for controlling the sound suppression switch to be open, and then the sound suppression switch can be controlled to be open based on the instruction.
[0154] In a case where the sound suppression switch is in an open state, the voice signal responsive device acquires another voice signal, determines that the sound type of the another voice signal is a non-real person sound type, and the another voice signal includes an instruction for controlling the sound suppression switch to be closed, and then the sound suppression switch can be controlled to be closed based on the instruction.
[0155] In some embodiments, in a case where the sound suppression switch is in an open state, then for a subsequently acquired voice signal, the response type of the voice signal can be determined according to the sound type of the voice signal.
[0156] In some embodiments, in the case that the sound suppression switch is in the off state, the voice signal can be directly regarded as a real human voice for the subsequent acquired voice signal without acquiring the voice signal type, and the voice signal is responded to, where the response can include at least one of a positive response, a response within a specified time period, a first response type response under a predetermined condition, a second response type response under a condition that the predetermined condition is not met, and the like.
[0157] Thus, the state of the sound suppression switch is flexibly controlled according to the control operation of the user, and the response type of the voice signal is flexibly determined according to the state of the sound suppression switch, meeting the individual needs of the user.
[0158] In some embodiments, in the case that the sound suppression switch is in the off state, the voice signal can be directly regarded as a real human voice for the subsequent acquired voice signal without acquiring the voice signal type, and the voice signal is responded to, where the response can include at least one of a positive response, a response within a specified time period, a first response type response under a predetermined condition, a second response type response under a condition that the predetermined condition is not met, and the like.
[0159] Correspondingly, determining the response type of the voice signal according to the voice signal type can include:
[0160] In the case that the voice signal type is a non-real human voice and the sound suppression switch is in the unopened state, determining the sound source region corresponding to the voice signal;
[0161] In the case that the sound source region is a first sound source region, responding to the voice signal;
[0162] In the case that the sound source region is a second sound source region, not responding to the voice signal.
[0163] For example, assuming that the first sound source region includes the co-pilot region and the rear region, and the second sound source region includes the main driver region, in the case that the sound suppression switch is in the unopened state, for the voice signal with a non-real human voice type in the main driver region, no response can be performed, and for the voice signal with a non-real human voice type in the co-pilot region and the rear region, a response can be performed, where the response can include at least one of a positive response, a response within a specified time period, a first response type response under a predetermined condition, a second response type response under a condition that the predetermined condition is not met, and the like.
[0164] In some embodiments, when responding to the voice signal of the first sound source region, the range of the response can also be set, such as responding only to specific control instructions.
[0165] Correspondingly, in a case where the sound source region is a first sound source region, responding to the voice signal can comprise:
[0166] In a case where the sound source region is a first sound source region, determining whether the voice signal comprises a control instruction of a non-device control type;
[0167] In a case where the voice signal comprises a control instruction of a non-device control type, responding to the control instruction of the non-device control type.
[0168] The control instruction of a device control type is an instruction for controlling a specific device, such as an instruction for controlling a device such as an air conditioner, a sound system, a seat, etc. in a car. The control instruction of a non-device control type is an instruction that is not for controlling a specific device, such as an instruction of a chat type, etc.
[0169] The response to the control instruction of the non-device control type can comprise at least one of an affirmative response, a response within a specified time length, a response of a first response type in a case where a preset condition is met, a response of a second response type in a case where the preset condition is not met, etc.
[0170] For example, assuming that the first sound source region comprises a co-driver region and a rear seat region, for a voice signal of a sound type of the co-driver region and the rear seat region that is not a real human voice, it can be determined whether the voice signal comprises a control instruction of a non-device control type, and in a case where the voice signal comprises the control instruction of the non-device control type, at least one of an affirmative response, a response within a specified time length, a response of a first response type in a case where a preset condition is met, a response of a second response type in a case where the preset condition is not met, etc. can be performed.
[0171] Thus, for a user who is inconvenient to make a sound, when a voice signal of a non-real human voice of the user is obtained, the voice signal can be flexibly responded to.
[0172] FIG. 4 is a structural schematic diagram of a voice signal response device provided by an embodiment of the present application. The voice signal response device provided by an embodiment of the present application comprises:
[0173] An acquisition module 210 is configured to acquire a voice signal, wherein the voice signal comprises a human voice component;
[0174] A type determination module 220 is configured to determine a sound type of the voice signal, wherein the sound type comprises a real human voice or a non-real human voice;
[0175] A response module 230 is configured to determine a response type of the voice signal according to the sound type of the voice signal.
[0176] The device can make appropriate and flexible responses to the voice signals of real human voice and non-real human voice by identifying whether the human voice component in the voice signal is real human voice, thereby improving the reliability of responses to voice signals containing human voice components.
[0177] Optionally, the type determination module 220 is specifically configured to:
[0178] determine the acoustic feature of the low-frequency part in the voice signal;
[0179] determine the sound type according to the acoustic feature.
[0180] Optionally, the type determination module 220 is specifically configured to:
[0181] determine the energy intensity of the low-frequency part in the voice signal;
[0182] determine the sound type according to the energy intensity.
[0183] Optionally, the type determination module 220 is specifically configured to:
[0184] input the voice signal into a trained voice processing model to determine the sound type of the voice signal through the voice processing model;
[0185] wherein the voice processing model is trained based on real human voice sample data and non-real human voice sample data.
[0186] Optionally, the training process of the voice processing model comprises:
[0187] obtain real human voice sample data and non-real human voice sample data;
[0188] label the real human voice sample data and the non-real human voice sample data according to the acoustic feature of the low-frequency part in the real human voice sample data and the non-real human voice sample data;
[0189] train a voice processing model based on the real human voice sample data and the non-real human voice sample data.
[0190] Optionally, the type determination module 220 is specifically configured to:
[0191] determine the sound source region corresponding to the voice signal;
[0192] if there are at least two sound source regions corresponding to the voice signal, then determine whether the sound type of the voice signal corresponding to each sound source region is real human voice;
[0193] screen the voice signal of a sound source region corresponding to a vocal type that is not real human voice to reserve the voice signal of a sound source region corresponding to a vocal type that is real human voice;
[0194] determine the vocal type of the voice signal based on the reserved voice signal of the sound source region.
[0195] Optionally, the response module 230 is specifically configured to:
[0196] in a case where the vocal type of the voice signal is real human voice, determine the response type of the voice signal according to the position of the wake-up word in the voice signal and / or the valid human voice component in the voice signal except the wake-up word.
[0197] Optionally, the response module 230 is specifically configured to:
[0198] in a case where the vocal type of the voice signal is real human voice, positively respond to the voice signal;
[0199] in a case where the vocal type of the voice signal is not real human voice, negatively respond to the voice signal;
[0200] wherein the positive response comprises at least one of the following:
[0201] wake up the electronic device;
[0202] wake up the voice assistant;
[0203] generate first information through a message notification component, the first information being used to prompt a user that the voice signal is responded to;
[0204] perform a corresponding operation according to the result of the semantic analysis;
[0205] the negative response comprises at least one of the following:
[0206] do not wake up the electronic device;
[0207] do not wake up the voice assistant;
[0208] generate second information through a message notification component, the second information being used to prompt a user that the voice signal is not responded to;
[0209] reject the voice signal.
[0210] Optionally, the response type comprises:
[0211] respond within a specified time length;
[0212] or, respond after a delay to the specified time length.
[0213] Optionally, the response module 230 is specifically configured to:
[0214] in the case where the preset condition is met, determining that the response type is a first response type;
[0215] in the case where the preset condition is not met, determining that the response type is a second response type;
[0216] the first response type comprises at least one of the following:
[0217] responding within the specified duration;
[0218] responding positively;
[0219] the second response type comprises at least one of the following:
[0220] responding after the specified duration;
[0221] responding negatively.
[0222] Optionally, the preset condition comprises at least one of the following:
[0223] determining that the voice signal needs to be responded to according to the semantics of the voice signal;
[0224] the voice signal contains a wake-up word;
[0225] the wake-up word is in a specified position in the voice signal;
[0226] the time interval between the wake-up word and the text before and / or after the wake-up word does not exceed a preset threshold.
[0227] Optionally, the response module 230 is specifically configured to:
[0228] in the case where the voice signal is not a real human voice, detecting whether the voice assistant has been woken up;
[0229] in the case where the voice assistant has been woken up, obtaining a voice recognition result corresponding to the voice signal;
[0230] displaying the voice recognition result in a preset form and obtaining a display duration of the voice recognition result;
[0231] after the display duration reaches a preset duration, stop displaying the voice recognition result.
[0232] Optionally, the response module 230 is specifically configured to:
[0233] in the case where the sound suppression switch is in an open state, determining the response type of the voice signal according to the voice type of the voice signal.
[0234] Optionally, the response device of the voice signal further comprises:
[0235] a control module, configured to control a state of the sound suppression switch based on a control operation of the user on the sound suppression switch.
[0236] Optionally, the control module is specifically configured to:
[0237] obtain another voice signal of a non-real person voice type;
[0238] detect whether the other voice signal comprises a control instruction for controlling the sound suppression switch;
[0239] in a case where the other voice signal comprises the control instruction, control a state of the sound suppression switch based on the control instruction.
[0240] Optionally, the response module 230 is specifically configured to:
[0241] in a case where the voice signal is of a non-real person voice type and the sound suppression switch is in an unopened state, determine a sound source area corresponding to the voice signal;
[0242] in a case where the sound source area is a first sound source area, respond to the voice signal;
[0243] in a case where the sound source area is a second sound source area, do not respond to the voice signal.
[0244] Optionally, the response module 230 is specifically configured to:
[0245] in a case where the sound source area is a first sound source area, determine whether the voice signal comprises a non-device control type control instruction;
[0246] in a case where the voice signal comprises a non-device control type control instruction, respond to the non-device control type control instruction.
[0247] The response device of the voice signal provided by the embodiments of the present application can be used to execute the voice signal response method provided by any of the above embodiments, and has the corresponding functions and beneficial effects.
[0248] FIG. 5 shows a structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application. The electronic device 10 can refer to a user-side device, and can also refer to a server-side device, including but not limited to a car machine, an electronic controller, a computer, a vehicle-mounted intelligent terminal, a smart phone, or a cloud server, etc. The electronic device 10 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device 10 can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, user equipment, and other similar computing devices. The components, their connections, and their functions, as described herein, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed herein.
[0249] As shown in FIG. 5, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0250] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks, wireless networks.
[0251] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above.
[0252] In some embodiments, the methods of the above-described embodiments can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., storage unit 18. In some embodiments, portions of the computer program, or all of the computer program, can be loaded onto the electronic device 10 via, e.g., ROM 12 and / or communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more of the steps of the above-described methods can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform any of the above-described methods by other means, e.g., with the aid of firmware.
[0253] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0254] Computer programs used to implement the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as part of a standalone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0255] In the context of this application, a computer readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer readable storage medium can be a machine readable signal medium. More specific examples of the machine readable storage medium will include a one or more lines of a electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0256] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device 10 having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device 10. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0257] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0258] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0259] The embodiments of the present application also provide a computer program product, comprising a computer program and / or instructions, which, when executed by a processor, implement the response method of the voice signal as described in any of the above embodiments.
[0260] It should be understood that the steps shown in the above various forms of flow can be reordered, added or deleted. For example, the steps described in the present application can be executed in parallel, sequentially or in different order, as long as the desired results of the technical solutions of the present application can be achieved, which are not limited herein.
[0261] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method of responding to a voice signal, wherein, The method comprises: acquiring a voice signal, wherein the voice signal comprises a human voice component; determining a voice type of the voice signal, wherein the voice type comprises real human voice or non-real human voice; determining a response type of the voice signal according to the voice type of the voice signal.
2. The method of claim 1, wherein, The method of determining the voice type of the voice signal comprises: determining an acoustic feature of a low-frequency part of the voice signal; determining the voice type according to the acoustic feature.
3. The method of claim 2, wherein, The method of determining the acoustic feature of the low-frequency part of the voice signal comprises: determining an energy intensity of the low-frequency part of the voice signal; The method of determining the voice type according to the acoustic feature comprises: determining the voice type according to the energy intensity.
4. The method of claim 1, wherein, The method of determining the voice type of the voice signal comprises: inputting the voice signal into a trained voice processing model to determine the voice type of the voice signal by the voice processing model; wherein the voice processing model is trained based on real human voice sample data and non-real human voice sample data.
5. The method of claim 4, wherein, The training process of the voice processing model comprises: acquiring real human voice sample data and non-real human voice sample data; annotating the real human voice sample data and the non-real human voice sample data according to acoustic features of low-frequency parts of the real human voice sample data and the non-real human voice sample data; training a voice processing model based on the real human voice sample data and the non-real human voice sample data.
6. The method according to any one of claims 1 to 5, wherein, The method of determining the voice type of the voice signal comprises: determining a sound source region corresponding to the voice signal; in a case where there are at least two sound source regions corresponding to the voice signal, determining whether the voice type of the voice signal corresponding to each sound source region is real human voice; masking the voice signal of the sound source region corresponding to the voice type of non-real human voice to retain the voice signal of the sound source region corresponding to the voice type of real human voice; determining the voice type of the voice signal based on the retained voice signal of the sound source region.
7. The method according to any one of claims 1 to 6, wherein, The method of determining the response type of the voice signal according to the voice type of the voice signal comprises: in a case where the voice type of the voice signal is real human voice, determining the response type of the voice signal according to a position of a wake-up word in the voice signal and / or a valid human voice component in the voice signal except the wake-up word.
8. The method according to any one of claims 1 to 7, wherein, The method of determining the response type of the voice signal according to the voice type of the voice signal comprises: in a case where the voice type of the voice signal is real human voice, positively responding to the voice signal; in a case where the voice type of the voice signal is non-real human voice, negatively responding to the voice signal; wherein the positive response comprises at least one of the following: waking up an electronic device; waking up a voice assistant; generating first information by a message notification component, wherein the first information is used to prompt a user that the voice signal is responded to; performing a corresponding operation according to a result of semantic analysis of the voice signal; the negative response comprises at least one of the following: not waking up an electronic device; not waking up a voice assistant; generating second information by a message notification component, wherein the second information is used to prompt a user that the voice signal is not responded to; reject the voice signal.
9. The method according to any one of claims 1-6, wherein, the response type, including: responding within a specified duration; or, responding after a delay to the specified duration.
10. The method according to any one of claims 1-6, wherein, the determination of the response type of the voice signal, including: in a case where a preset condition is met, determining the response type as a first response type; in a case where the preset condition is not met, determining the response type as a second response type; the first response type includes at least one of: responding within the specified duration; responding positively; the second response type includes at least one of: responding after a delay to the specified duration; responding negatively.
11. The method of claim 10, wherein, the preset condition includes at least one of: determining that the voice signal needs to be responded to according to the semantics of the voice signal; the voice signal contains a wake-up word; the wake-up word is in a specified position in the voice signal; the time interval between the wake-up word and the text before and / or after it does not exceed a preset threshold.
12. The method of any one of claims 1-6, wherein, the determination of the response type of the voice signal according to the voice type of the voice signal, including: in a case where the voice type of the voice signal is non-human voice, detecting whether the voice assistant has been woken up; in a case where the voice assistant has been woken up, obtaining a voice recognition result corresponding to the voice signal; displaying the voice recognition result in a preset form and obtaining a display duration of the voice recognition result; after the display duration reaches a preset duration, stop displaying the voice recognition result.
13. The method of any one of claims 1-12, wherein, the determination of the response type of the voice signal according to the voice type of the voice signal, including: in a case where the sound suppression switch is in an open state, determining the response type of the voice signal according to the voice type of the voice signal.
14. The method of claim 13, wherein, The method further includes: controlling the state in which the sound suppression switch is based on the user's control operation of the sound suppression switch.
15. The method of claim 14, wherein, controlling the state in which the sound suppression switch is, including: obtaining another voice signal with a non-human voice type; detecting whether the other voice signal includes a control instruction for controlling the sound suppression switch; in a case where the other voice signal includes the control instruction, controlling the state in which the sound suppression switch is based on the control instruction.
16. The method of any one of claims 1-6, wherein, determining the response type of the voice signal according to the voice type of the voice signal, including: in a case where the voice type of the voice signal is non-human voice and the sound suppression switch is in an unopened state, determining a sound source area corresponding to the voice signal; in a case where the sound source area is a first sound source area, responding to the voice signal; in a case where the sound source area is a second sound source area, not responding to the voice signal.
17. The method of claim 16, wherein, in a case where the sound source area is a first sound source area, responding to the voice signal, including: in a case where the sound source area is a first sound source area, determining whether the voice signal includes a control instruction of a non-device control type; in a case where the voice signal includes a control instruction of a non-device control type, responding to the control instruction of the non-device control type.
18. A voice signal response device, comprising: An acquisition module is configured to acquire a voice signal, wherein the voice signal comprises a human voice component; A type determination module is configured to determine a sound type of the voice signal, wherein the sound type comprises a real human sound or a non-real human sound; A response module is configured to determine a response type of the voice signal according to the sound type of the voice signal.
19. An electronic device comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the response method of the voice signal according to any one of claims 1-17.
20. A computer readable storage medium having stored thereon a computer program, wherein, The program is executed by the processor to implement the response method of the voice signal according to any one of claims 1-17.
21. A computer program product comprising computer programs / instructions, which, when executed by a processor, implement the response method of the voice signal according to any one of claims 1-17.
Citation Information
Patent Citations
Voice recognition method and device, storage medium and electronic equipment
CN110459204A
Method, device, equipment and medium for voice interaction control
CN110718223A
Device and method for simultaneously identifying human voice and non-human voice
CN112185357A
Vehicle-mounted voice interaction method, device and equipment and storage medium
CN114360527A
Method and device for waking up equipment by voice, storage medium and electronic device
CN116564285A