A false wake-up audio acquisition method, system, device and storage medium
By analyzing wake-up logs from multiple devices in real-world scenarios, the system accurately extracts false wake-up audio, solving the problem of low acquisition efficiency in existing technologies. This achieves efficient and accurate acquisition of false wake-up audio and incremental training data, thereby improving server stability.
Patent Information
- Application Number
- CN202111185730.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-12
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-10-12
AI Technical Summary
Existing technologies struggle to automate audio capture in real-world scenarios when testing for false wake-up, resulting in low capture efficiency and potential leakage of confidential or privacy information, leading to inefficient data analysis.
By receiving and storing ambient audio, multiple devices with speech recognition capabilities are used to perform wake-up operations under the same environment to generate wake-up logs. Based on a preset false wake-up strategy, the system analyzes and judges the data that caused the false wake-up, extracts the audio data, and uses a precise segmentation method to shorten the audio duration and improve recognition accuracy.
It enables efficient acquisition of false wake-up audio in real-world scenarios, shortens audio duration, improves the accuracy of recognition results and the amount of training data, and enhances server stability and data processing efficiency.
Smart Images

Figure CN113936645B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio data collection, and in particular to a false wake-up audio collection method, system, device and storage medium. BACKGROUND
[0002] In addition to the wake-up rate and the recognition rate, the performance indicators of an AI voice product (AI, Artificial Intelligence) also include the false wake-up rate. For product performance optimization, a large amount of environmental audio is usually collected in advance and input into a to-be-optimized model for training to improve the corresponding indicators. The environmental audio can be various noises and their superpositions that can occur in real scenes.
[0003] Usually, these performance indicators are tested in a sound laboratory, for example, by playing various noises. However, it is not possible to achieve fully automated collection of various noises, and only large-scale long-time simulation tests can be performed. After analyzing the test results, abnormal environmental audio is found and used for testing false wake-up rate training and optimization.
[0004] When performing performance indicator tests, the existing environmental audio collection technology at least has the following technical problems:
[0005] (1) For the purpose of controllable testing process and easy data acquisition, a simulation restoration test is usually performed in an audio laboratory. However, this test method has certain differences from real scenes and is difficult to reproduce one-to-one.
[0006] (2) In the testing process, a very long time is spent on recording audio. However, the audio time is too long, and technicians need to find abnormal audio for analysis and optimization after combining the test results, resulting in low efficiency.
[0007] (3) To collect real noise data, the preferred solution is to perform tests in a real use environment. However, uninterrupted audio recording may leak company secrets or personal privacy, and therefore, precise positioning of the audio is required. SUMMARY
[0008] To solve the above problems, the embodiments of the present application provide a false wake-up audio collection method, system, device and storage medium, which achieve the purposes of shortening the time length of false wake-up audio, improving the accuracy of recognition results, reducing the audio occupied space, and increasing the training data amount.
[0009] In a first aspect, the present application provides a false wake-up audio collection method, which comprises:
[0010] S100: receiving and storing environmental audio;
[0011] S200: receive a plurality of wake-up logs; the wake-up logs are configured to generate a plurality of wake-up logs by a plurality of same devices with voice recognition function according to a preset trigger wake-up strategy, and perform wake-up operation on the environmental audio of the same sound source under the same environment respectively;
[0012] S300: based on the same time period in the time sequence, according to the preset false wake-up strategy, the plurality of wake-up logs are combined as data sources for analysis and judgment, and according to the judgment result, each segment of audio data causing false wake-up is extracted from the environmental audio.
[0013] Further, the false wake-up strategy in the step S300 is configured to,
[0014] When the trigger wake-up results of all the wake-up logs in the data source are consistent, the environmental audio in the current time period is determined as normal audio, and when the environmental audio of the next time period is received, the environmental audio of the current time period is discarded;
[0015] When the trigger wake-up result of any of the wake-up logs in the data source is inconsistent with the results of other wake-up logs, the environmental audio in the current time period is determined as abnormal audio, and the time period to which the abnormal audio belongs is extracted, and the audio data is extracted from the environmental audio.
[0016] Further, the false wake-up strategy in the step S300 is configured to represent the trigger wake-up result of each device by 0 or 1 through the wake-up logs;
[0017] Wherein, 1 represents that the device triggers wake-up successfully, and 0 represents that the device triggers wake-up fails; so that when the trigger wake-up results of the wake-up logs in the composed data source are all 1 or all 0, it means that all the devices trigger wake-up successfully or trigger wake-up fails at the same time in the same time period, and the environmental audio in the current time period is determined as normal audio; otherwise, the environmental audio in the current time period is determined as abnormal audio.
[0018] Further, the method for extracting each segment of audio data causing false wake-up from the environmental audio in the step S300 is to cut the stored environmental audio; wherein when the judgment result indicates that there is abnormal audio in the environmental audio, the environmental audio in the time period to which the abnormal audio belongs is cut in front and back.
[0019] Further, in the step S100, the sound source of the received and stored environmental audio is the environmental audio generated in the actual scene, or the environmental audio of various places simulated by the audio playing device.
[0020] In a second aspect, the application provides a false wake-up audio acquisition system, which adopts the false wake-up audio acquisition method of the first aspect. The system comprises a host, an audio source, a sound collecting device, and a plurality of identical voice recognition devices with voice recognition function. The host is connected to each voice recognition device through a communication serial port and connected to the sound collecting device through an audio port.
[0021] The audio source is configured to output environmental audio.
[0022] The sound collecting device is configured to record the environmental audio in real time and transmit the environmental audio to the host.
[0023] The voice recognition device is configured to synchronously collect and recognize the environmental audio of the same audio source, perform a wake-up operation based on a preset trigger wake-up strategy, generate a plurality of wake-up logs, and transmit the wake-up logs to the host.
[0024] The host is configured to combine the plurality of wake-up logs into a data source for analysis and judgment based on the same time period in a time sequence according to a preset false wake-up strategy, and extract each segment of audio data that causes false wake-up from the environmental audio according to the judgment result.
[0025] In a third aspect, the application provides a false wake-up audio acquisition device, which adopts the false wake-up audio acquisition method of any one of the first aspect. The device comprises:
[0026] An environmental audio acquisition module configured to control a plurality of microphones to synchronously collect environmental audio of the same audio source in the same environment through corresponding sound pickup holes.
[0027] A wake-up log acquisition module configured to control a plurality of identical voice recognizers with voice recognition function to extract a plurality of wake-up logs generated after performing a wake-up operation on the environmental audio according to a preset trigger wake-up strategy, respectively. Wherein, the microphone and the voice recognizer are one-to-one paired.
[0028] A false wake-up audio extraction module configured to combine the wake-up logs into a data source for analysis and judgment based on the same time period in a time sequence according to a preset false wake-up strategy, and extract each segment of audio data that causes false wake-up from the environmental audio in the storage space according to the judgment result.
[0029] Further, each sound pickup hole is configured to be arranged at equal intervals between the audio source of the environmental audio.
[0030] Further, a plurality of sound pickup holes are arranged in a regular N-polygon, and the audio source is arranged on the central axis of a regular N-pyramid where the regular N-polygon is located.
[0031] In a fourth aspect, the present application provides a storage medium storing executable program codes, at least one processor reading the executable program codes to run a computer program corresponding to the executable program codes to execute at least one step of the false wake-up audio collection method according to any one of the first aspect.
[0032] The technical solutions provided in the embodiments of the present application have at least the following technical effects:
[0033] (1) Since the abnormal audio triggering false wake-up is directly extracted, the duration of false wake-up audio is shortened, the accuracy of the recognition result is improved, the audio occupied space is reduced, and the amount of training data is increased. Moreover, since the abnormal audio extracted accurately is used for false wake-up training, the data processing efficiency is improved, the throughput of data that can be processed by a server to which a plurality of processors with voice recognition function belong within a unit time is larger, so as to avoid causing data congestion and downtime, and the stability of the server is improved.
[0034] (2) Since the preset false wake-up strategy is used, the false wake-up strategy is directly used to analyze and judge the environmental audio, the environmental audio is quickly processed by a simple logic method, and the collection efficiency of false wake-up audio is improved.
[0035] (3) Since only the host, the sound collecting device and the voice recognition device are used to build the collection environment in the provided system, it is convenient for technical personnel to quickly simulate the audio collection scene, and the audio collection is not limited in a specific actual scene, and the data collection of different voice recognition devices is facilitated. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 A flow chart of a false wake-up audio collection method in the first embodiment of the present application;
[0037] Figure 2 An example diagram of abnormal audio delay extraction principle in the first embodiment of the present application;
[0038] Figure 3 A system simulation structure schematic block diagram of a false wake-up audio collection system in the second embodiment of the present application;
[0039] Figure 4 A system simulation structure block diagram of a three-voice recognition device scene in the second embodiment of the present application;
[0040] Figure 5 A side-by-side placement schematic diagram of a voice recognition device in the second embodiment of the present application;
[0041] Figure 6 A side-by-side placement schematic diagram of a voice recognition device in the second embodiment of the present application;
[0042] Figure 7Figure 1 is a structure diagram of an audio collection device for false wake-up in Embodiment Three of the present application.
[0043] Figure 8 Figure 2 is a structure diagram of the installation position of a sound source and a microphone in Embodiment Three of the present application. DETAILED DESCRIPTION
[0044] In order to better understand the above technical solutions, the above technical solutions will be described in detail below in combination with the drawings of the specification and specific embodiments.
[0045] Embodiment One
[0046] Referring to the drawings, the embodiments of the present application provide a false wake-up audio collection method, which comprises the following steps. Figure 1
[0047] Step S100: receiving and storing environmental audio.
[0048] In step S100, the sound source of the received and stored environmental audio is environmental audio generated in an actual scene or environmental audio of various places simulated by an audio playing device.
[0049] Step S200: receiving multiple wake-up logs. The wake-up logs are configured to be generated by multiple same devices with voice recognition function according to a preset trigger wake-up strategy, and the multiple wake-up logs are generated after the same environmental audio of the same sound source is executed in the same environment.
[0050] The device in this step S200 can be understood as any AI voice product on the market with voice recognition function, of course, not limited to the product, but also the main control chip in the product. In this embodiment, the audio data causing false wake-up is used for false wake-up training of the AI voice product or the related main control chip.
[0051] Step S300: based on the same time period in the time sequence, according to the preset false wake-up strategy, the multiple wake-up logs are combined as a data source for analysis and judgment, and according to the judgment result, each segment of audio data causing false wake-up is extracted from the environmental audio.
[0052] In step S300, the false wake-up strategy is configured as follows:
[0053] When the trigger wake-up results of all wake-up logs in the data source are consistent, the environmental audio in the current time period is determined as normal audio, and when the environmental audio of the next time period is received, the environmental audio of the current time period is discarded.
[0054] When the trigger wake-up result of any wake-up log in the data source is inconsistent with the results of other wake-up logs, the environmental audio in the current time period is determined as abnormal audio, and the time period to which the abnormal audio belongs is extracted, and the audio data is extracted from the environmental audio.
[0055] Therefore, from the step S300, it can be seen that after judging by multiple trigger wake-up results, the abnormal audio segment triggering false wake-up is directly extracted, thereby shortening the output duration of false wake-up audio, improving the accuracy of false wake-up rate training recognition results, and directly using the abnormal audio segment, thereby reducing the audio occupation space required for false wake-up rate training, and increasing the training data amount. In the embodiment, the audio data causing false wake-up is accurately extracted for false wake-up training, thereby improving the data processing efficiency, making the device with speech recognition function and the server performing training have larger data throughput in unit time, thereby avoiding data congestion and downtime, and improving the stability of server training. In the embodiment, the preset false wake-up strategy is used to analyze the wake-up results in the wake-up log, the analysis logic is simple, and multiple wake-up logs can be directly formed into a table format based on time series, thereby obtaining the judgment result of the stored environmental audio in the same time period in the time series based on the trigger wake-up results of the wake-up logs in the same time period, thereby realizing simple logic method for quickly analyzing and processing environmental audio, and improving the collection efficiency of false wake-up audio.
[0056] Further, in the embodiment, the false wake-up strategy in step S300 is configured to represent the trigger wake-up result of each device by 0 or 1 through the wake-up log. Among them, 1 represents that the device triggers wake-up successfully, and 0 represents that the device triggers wake-up fails, so that when the trigger wake-up results of the wake-up logs in the composed data source are all 1 or all 0, it means that all devices trigger wake-up successfully or trigger wake-up fails at the same time in the same time period, and the environmental audio in the current time period is determined as normal audio; otherwise, the environmental audio in the current time period is determined as abnormal audio.
[0057] Further, assuming three identical devices with voice recognition function, the data source of the trigger wake-up result group of the wake-up log generated by the three devices can be represented as 111, 110, 101, 000, and the like, wherein the data source 111 and 000 represent that the three devices trigger wake-up at the same time period and the trigger wake-up is successful or the trigger wake-up recognition is successful, and it is determined that the environmental audio at the time period is normal audio, otherwise, the data source determines that the environmental audio is abnormal audio according to the trigger wake-up result. In the embodiment, the time period of the abnormal audio or the normal audio is obtained through the determination of the data source, and then the stored environmental audio is cut to achieve the purpose of saving the abnormal audio and discarding the normal audio segment. In the embodiment, the normal audio is discarded through cyclic coverage. For example, refer to Table 1 for a kind of data source in table format.
[0058] Time series Wake log 1 Wake log 2 Wake log 3 Determination Time period 1 1 1 1 PASS Time period 2 1 0 0 FALL Time period 3 1 0 1 FALL Time period 4 1 1 0 FALL Time period 5 0 0 1 FALL Time period 6 0 1 0 FALL Time period 7 0 1 1 FALL Time period 8 0 0 0 PASS
[0059] Table 1
[0060] Further, for the determination operation of the wake-up log in the false wake-up strategy, one way is to pre-set a device local wake-up recognition log printer mechanism, output the wake-up log of each device, and compare the recognition results of each wake-up log. When the trigger wake-up results of the wake-up log of a certain time period in the data source are inconsistent, it is determined that the environmental audio recorded in the time period is abnormal audio, and the time period to which the abnormal audio belongs is saved. When the log results of the wake-up log of each device in a certain time period are consistent, it is determined that the environmental audio recorded in the time period is normal audio, and the normal audio is discarded without storage.
[0061] Further, in the embodiment, a preset comparison algorithm is used to obtain the wake-up log of multiple devices in real time, combine the wake-up recognition results of each wake-up log in the same form based on time sequence, and then use the comparison algorithm to compare and analyze to obtain the time period to which the abnormal audio belongs, extract the environmental audio of the time period, and then save it.
[0062] Furthermore, this embodiment can also pre-set a first processing area and a second processing area. The first processing area is used to process ambient audio, and the second processing area has a pre-set comparison algorithm. Based on the same timer and the same time series, the first processing area receives ambient audio, and the second processing area receives wake-up logs. The second processing area stores the recognition results of the wake-up logs in a data source form set based on the time series according to the false wake-up strategy. Then, it uses the comparison algorithm to compare the trigger wake-up results of the same time series. When different results are found, the comparison stops, and then the ambient audio received by the first processing area is processed and saved. When no different recognition results are found, the comparison algorithm is used to complete the comparison and then the analysis and processing of the next time series continues. When the first processing area receives new ambient audio, it directly overwrites the previous ambient audio with the latest ambient audio without saving the ambient audio of the previous period. Therefore, only a small cache space is needed to temporarily store the ambient audio, thereby cyclically overwriting the accurately recognized audio data and leaving only the audio data that causes false wake-ups and false recognitions, shortening the duration of the audio data that causes false wake-ups, which is beneficial to algorithm optimization.
[0063] In this embodiment, step S300 involves extracting audio data segments causing false wake-ups from the ambient audio by segmenting the stored ambient audio. Specifically, when the judgment result indicates the presence of abnormal audio in the ambient audio, the ambient audio within the time period to which the abnormal audio belongs is segmented with both preceding and following delays. Further explanation: when abnormal audio is detected, the ambient audio needs to be processed to extract the abnormal audio. In this embodiment, the ambient audio within the time period of the abnormal audio is segmented with both preceding and following delays to ensure the integrity of the stored abnormal audio.
[0064] For example, see reference. Figure 2 As shown, during an experiment, when the test started, the ambient audio source output was activated at 5 seconds, triggering the voice recognition device's wake-up. However, the wake-up log didn't return until 7 seconds, resulting in a 2-second delay. Within those 2 seconds, there's a 1.5-second timeframe from the sound being triggered to its end, plus 0.4 seconds for algorithm processing and 0.1 seconds for transmission. Therefore, a 2-second delay is needed, plus a 2-second redundancy, resulting in a 7-second lead time. Because the wake-up algorithm processing times of the three voice recognition devices are inconsistent, it's impossible to guarantee that their processing times are identical. Therefore, a 2-second algorithm delay and a 2-second redundancy are needed. Thus, in this embodiment, the ambient audio editing needs an extension of approximately 11 seconds to ensure audio integrity.
[0065] Example 2
[0066] refer to Figures 3-4As shown, the embodiment of the present application provides a false wake-up audio acquisition system, which adopts the false wake-up audio acquisition method of any one of embodiment one, and the system comprises: a host 110, an audio source 140, a sound collecting device 130, a plurality of same voice recognition devices 120 with voice recognition function; the host 110 is connected with the voice recognition device 120 through a communication serial port and connected with the sound collecting device 130 through an audio port. The audio source 140 is configured to output environmental audio. The sound collecting device 130 is configured to record environmental audio in real time and transmit the environmental audio to the host 110. The voice recognition device 120 is configured to synchronously collect and recognize the environmental audio of the same audio source 140, execute a wake-up operation based on a preset trigger wake-up strategy, generate a plurality of wake-up logs, and transmit them to the host. The host 110 is configured to combine the plurality of wake-up logs into a data source for analysis and judgment based on the same time period in the time sequence according to the preset false wake-up strategy, and extract each segment of audio data causing false wake-up from the environmental audio according to the judgment result.
[0067] Further explanation, the audio source 140 in the embodiment can be the environmental audio generated in the actual scene, or the environmental audio of various places simulated by the audio playing device. For example, the environmental audio formed by the speech of a person in the environment may trigger the preset wake-up strategy of the voice recognition device. The audio playing device is used to play various environmental audio to simulate various application places of the experimental object. Among them, the environmental audio may contain the trigger corpus in the voice recognition device, or may not contain it, or only contain other noise, such as noise data and superimposed noise that may exist in different scenes. The audio playing device can be a loudspeaker. The sound collecting device records the environmental audio output by the audio source in real time and transmits the environmental audio to the host.
[0068] The host 110 presets a false wake-up strategy, receives environmental audio and a plurality of wake-up logs based on the same time period in the time sequence, combines the wake-up logs into a data source for analysis and judgment, and extracts audio data causing false wake-up from the environmental audio according to the judgment result. Among them, the host can be various computer terminals, and has a controller, a communication serial port, an audio port, a memory and the like.
[0069] Further explanation, since the voice recognition device 120 in the embodiment has voice recognition function and preset trigger wake-up strategy, and the voice recognition device executes wake-up operation on the predetermined corpus. For example, the voice recognition device is a smart air conditioner, which executes opening or closing through the corpus. For example, when the corpus contains "turn on the air conditioner", the smart air conditioner collects the "turn on the air conditioner" corpus and immediately executes the wake-up start mode.
[0070] Further, the controller of the host 110 controls the sound collecting device to record the environmental audio output by the sound source in real time through the audio port, and controls the voice recognition device to collect the environmental audio, and then performs audio wake-up recognition according to the trigger wake-up strategy set by the voice recognition device, so as to take the wake-up results of the voice recognition devices as the data source for judging false wake-up. The host extracts the audio data causing false wake-up from the environmental audio according to the judgment on the data source, and saves the audio data into the memory.
[0071] To improve the accuracy of audio data collection, the sound collecting device and the sound pickup holes of the plurality of voice recognition devices in the embodiment are arranged at equal intervals between the sound source. Further, the sound pickup holes of the sound collecting device and the plurality of voice recognition devices are in the same direction, so as to ensure that the distance between the sound pickup holes and the sound source is consistent. Further, when the sound collecting device and the plurality of voice recognition devices are located in the same plane, the distances between the sound collecting device and the plurality of voice recognition devices are equal, and the sound source is located on the central axis of the regular N-polygon in which the sound collecting device and the plurality of voice recognition devices are located. When the sound collecting device and the plurality of voice recognition devices are not located in the same plane, the distances between the sound collecting device and the plurality of voice recognition devices can be equal and located on the same spherical surface, and the sound source is located at the center of the spherical surface. Of course, it is not limited to the above-described cases, and the implementation position can be set arbitrarily according to the requirements in the case of negligible error. In addition, when the voice recognition device has an audio storage function, the sound collecting device can be directly omitted, so as to reduce the implementation cost.
[0072] The voice recognition devices in the embodiment are of the same type, have the same hardware, and have the same electrical performance, and can adopt embedded firmware, and the wake-up algorithm mode adopted is also consistent. The sound collecting device and the voice recognition device in the embodiment have basically the same software and hardware requirements, and have the same placement requirements, which require real-time recording of the original environmental audio in a closed space, so as to obtain audio data for algorithm model training in the original environmental audio.
[0073] Reference Figures 5-6As shown, the distances between the voice recognition devices and the sound collection devices are equal, and the distances from the sound sources are also equal. For example, the voice recognition device is included, and the sound collection holes of the voice recognition device and the sound collection device are in the same direction. Taking two sound collection holes as an example: S1, S2, S3 and S4 are the distances between the two sound collection holes, and S1=S2=S3=S4 is required, that is, the distances between the sound collection holes of the four devices are the same, H1, H2 and H3 are the distances between the four devices, and the distances between each device are the same, so H1=H2=H3 is maintained. At the same time, N1 is the center point, N1=N2=N3 is maintained, and the equal distances are set. It can be seen that if the placement is not accurate, the sound collection hole directions will be inconsistent, so that the sound received by the devices will have a time difference, and the consistency of the simultaneous reception of the devices cannot be guaranteed. The test data has inaccuracy. The wake-up results output by the three voice recognition devices can be represented as 111, 110, 101, 000, etc. 1 represents wake-up recognition, and 0 represents no wake-up recognition. Among them, 111 and 000 represent that the three voice recognition devices are simultaneously woken up and recognized or not woken up and not recognized at the same time, and then it is judged to be a normal audio, and the remaining results are false wake-up and false recognition. Taking the response time of the first device as 0S, because the algorithm has different processing times after the sound source outputs, a time is needed to react, so it is judged to be a normal audio if the three devices respond within three seconds, and if one of the three devices does not respond within three seconds, it is judged to be a false wake-up and false recognition, and the consistency of the devices is ensured.
[0074] In this embodiment, after saving the abnormal audio, the environmental audio collected by each voice recognition device is also stored and named, for example: when the audio wake-up word is "Xiaomei Xiaomei", it can be named as "Xiaomei Xiaomei-1" to represent the environmental audio collected by the first voice recognition device, and the like.
[0075] Embodiment three
[0076] Reference Figures 7-8 As shown, the present application provides a false wake-up audio collection device, which adopts the false wake-up audio collection method of any one of embodiment one, and the device comprises:
[0077] The environmental audio acquisition module 210 is configured to control the multiple microphones 320 to synchronously collect the environmental audio of the same sound source 310 through the corresponding sound collection holes.
[0078] The wake-up log acquisition module 220 is configured to control multiple same voice recognition devices with voice recognition function to extract multiple wake-up logs generated after performing a wake-up operation on the environmental audio according to a preset trigger wake-up strategy. Among them, the microphone 320 and the voice recognition device are one-to-one matched. The microphone can be understood as a microphone.
[0079] The false wake-up audio extraction module 230 is configured to combine the wake-up logs into a data source for analysis and judgment based on the same time period in the time sequence according to a preset false wake-up strategy, and extract each segment of audio data causing false wake-up from the environmental audio in the memory according to the judgment result.
[0080] The pickup holes in the embodiment are arranged at equal intervals with the sound sources of the environmental audio. Further, the pickup holes are arranged in a regular N-polygon, and the sound sources are arranged on the central axis of a regular N-pyramid in which the regular N-polygon is located.
[0081] Embodiment four
[0082] The storage medium stores executable program codes, and at least one processor reads the executable program codes to run a computer program corresponding to the executable program codes to execute at least one step of the false wake-up audio acquisition method of any one of the steps in the first aspect.
[0083] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by programs instructing related hardware, and the programs can be stored in a computer readable storage medium, which can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0084] The above only describes the preferred embodiments of the present application, and the protection scope of the present application is not limited to the above-mentioned embodiments. Any technical solution falling within the concept of the present application shall be considered as falling within the protection scope of the present application. It should be noted that, for ordinary skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements shall also be considered as falling within the protection scope of the present application.
[0085] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above-mentioned embodiments are not described, however, as long as the combinations of the technical features do not exist contradictory, they shall be considered as falling within the scope of the present application.
[0086] It should be noted that the above-mentioned embodiments can be freely combined as needed. The above only describes the preferred embodiments of the present application, and it should be noted that, for ordinary skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements shall also be considered as falling within the protection scope of the present application.
[0087] The software programs of the present application can be executed by a processor to implement the steps or functions described above. Similarly, the software programs (including related data structures) of the present application can be stored in a computer readable recording medium, such as a RAM memory, a magnetic or optical drive or diskette, and the like. In addition, some of the steps or functions of the present application can be implemented in hardware, such as a circuit that cooperates with the processor to perform various functions or steps. The methods disclosed in the embodiments of the present specification can be applied to a processor or implemented by a processor. The processor can be an integrated circuit chip having a processing capability of signals. In the implementation process, the steps of the above method can be completed by integrated logic circuits or instructions in the form of software in the processor. The processor described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The disclosed methods, steps and logic block diagrams in the embodiments of the present specification can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the method disclosed in conjunction with the embodiments of the present specification can be directly embodied as a hardware code processor to execute and complete, or be executed and completed by a combination of hardware and software modules in the code processor. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above method.
[0088] The embodiments also provide a computer readable storage medium storing one or more programs, which when executed by an electronic system comprising a plurality of applications, cause the electronic system to perform the method described in Embodiment 1. This will not be repeated here.
[0089] Computer-readable media includes permanent and non-permanent, moveable and non- moveable media that can be implemented in any method or technology for storage of information, such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, without limitation, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disks (DVDs) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information for access by a computing device. According to the definitions herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0090] The systems, apparatuses, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices. The computer readable medium includes permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules, or other data. Examples of the storage medium of the computer include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, the computer readable medium does not include transitory media such as modulated data signals and carriers. It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0091] In addition, part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present application can be invoked or provided. The program instructions invoking the method of the present application can be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal bearing medium, and / or stored in the working memory of the computer device running according to the program instructions. Here, according to an embodiment of the present application, the device includes a memory for storing computer program instructions and a processor for executing program instructions, wherein when the computer program instructions are executed by the processor, the device triggers the operation of the method and / or technical solutions based on the aforementioned embodiments according to the present application.
[0092] While the preferred embodiments of the application have been described, additional variations and modifications can be made to these embodiments by those skilled in the art once they have the benefit of the present disclosure without departing from the spirit and scope of the application. Accordingly, it is intended that the appended claims include all such modifications and variations as fall within the scope of the present application.
[0093] It is apparent that those skilled in the art can make various changes and modifications to the application without departing from the spirit and scope of the application. Thus, it is intended that the application include all such modifications and alterations insofar as they come within the scope of the claims or the equivalents thereof.
Claims
1. A false wake-up audio acquisition method, characterized in that, The method comprises: S100: receiving and storing environmental audio; S200: receiving a plurality of wake-up logs; the wake-up logs are configured to generate a plurality of wake-up logs by using a plurality of same devices with voice recognition function to perform wake-up operations on the same environmental audio in the same environment according to a preset trigger wake-up strategy; S300: based on the same time period in the time sequence, combining the plurality of wake-up logs as a data source according to a preset false wake-up strategy, and extracting each segment of audio data causing false wake-up from the environmental audio according to the judgment result.
2. The false wake-up audio collection method of claim 1, wherein, The false wake-up strategy in the step S300 is configured to: when the trigger wake-up results of all the wake-up logs in the data source are consistent, the environmental audio in the current time period is determined as normal audio, and the environmental audio in the current time period is discarded when the environmental audio in the next time period is received; when the trigger wake-up result of any of the wake-up logs in the data source is inconsistent with the results of other wake-up logs, the environmental audio in the current time period is determined as abnormal audio, and the time period to which the abnormal audio belongs is extracted, and the audio data is extracted from the environmental audio.
3. The false wake-up audio capture method of claim 1 or 2, wherein, The false wake-up strategy in the step S300 is configured to represent the trigger wake-up result of each device by 0 or 1 through the wake-up logs; wherein 1 represents that the device triggers wake-up successfully, and 0 represents that the device triggers wake-up unsuccessfully, so that when the trigger wake-up results of the wake-up logs in the composed data source are all 1 or all 0, it means that all the devices trigger wake-up successfully or fail to trigger wake-up at the same time in the same time period, and the environmental audio in the current time period is determined as normal audio; otherwise, the environmental audio in the current time period is determined as abnormal audio.
4. The false wake-up audio capture method of claim 2, wherein, The method for extracting each segment of audio data causing false wake-up from the environmental audio in the step S300 is to cut the stored environmental audio; wherein when the judgment result indicates that abnormal audio appears in the environmental audio, the environmental audio in the time period to which the abnormal audio belongs is cut in front and back.
5. The false wake-up audio collection method of claim 1, wherein, In the step S100, the audio source of the received and stored environmental audio is the environmental audio generated in the actual scene, or the environmental audio of various places simulated by the audio playing device.
6. A false wake-up audio acquisition system, comprising: The system comprises: a host, an audio source, a sound collecting device, and a plurality of same voice recognition devices with voice recognition function; the host is connected with each voice recognition device through a communication serial port, and connected with the sound collecting device through an audio port; The audio source is configured to output environmental audio; The sound collecting device is configured to record the environmental audio in real time, and transmit the environmental audio to the host; The voice recognition device is configured to synchronously collect and recognize the environmental audio of the same audio source, perform a wake-up operation based on a preset trigger wake-up strategy, generate a plurality of wake-up logs, and transmit the wake-up logs to the host; The host is configured to combine multiple wake-up logs into a data source for analysis and judgment based on the same time period in the time sequence according to a preset false wake-up strategy, and extract each piece of audio data causing false wake-up from the environmental audio according to the judgment result.
7. A false wake-up audio acquisition device, comprising: The device comprises: An environmental audio acquisition module configured to control multiple microphones to synchronously acquire environmental audio of the same sound source in the same environment through corresponding sound pickup holes; A wake-up log acquisition module configured to control multiple identical voice recognizers with voice recognition function to extract multiple wake-up logs generated after performing wake-up operations on the environmental audio according to a preset trigger wake-up strategy; wherein the microphones and the voice recognizers are paired one by one; A false wake-up audio extraction module configured to combine wake-up logs into a data source for analysis and judgment based on the same time period in the time sequence according to a preset false wake-up strategy, and extract each piece of audio data causing false wake-up from the environmental audio in the memory according to the judgment result.
8. The false wake-up audio capture device of claim 7, wherein, Each sound pickup hole is configured to be arranged at equal intervals with the sound source of the environmental audio.
9. The false wake-up audio capture device of claim 8, wherein, Multiple sound pickup holes are arranged in a regular N-polygon, and the sound source is arranged on the central axis of a regular N-pyramid where the regular N-polygon is located.
10. A storage medium storing executable program code, at least one processor reading the executable program code to run a computer program corresponding to the executable program code to execute at least one step of the false wake-up audio acquisition method according to any one of claims 1-5.
Citation Information
Patent Citations
Method using mobile terminal to carry out voice awakening and device thereof
CN103561175A
Method and device for awakening applications
CN104660792A