Target corpus rejection method and device, storage medium and electronic device

By building forward and reverse corpus optimization rejection models, the problem of wrong start of smart home devices due to noise and accent is solved, and high-precision corpus recognition and control instructions are achieved accurately executed, improving user experience and system security.

CN120260563APending Publication Date: 2025-07-04HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510385205.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

User interaction with smart home devices is susceptible to noise and pronunciation accents, causing the device to be started incorrectly or executed incorrect control instructions, affecting user experience and security.

Method used

Forward and reverse corpus are constructed, target corpus is identified by optimizing the rejection model, and corpus with and without control intentions are distinguished, and corpus is filtered and optimized using phoneme sequences and corpus warning models to generate high-precision rejection model.

Benefits of technology

Improve the identification accuracy of smart home devices, avoid incorrect start and control instructions, and improve user experience and system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260563A_ABST
    Figure CN120260563A_ABST
Patent Text Reader

Abstract

The invention discloses a target corpus rejection method and device, a storage medium and an electronic device, and relates to the technical field of smart home, and the target corpus rejection method comprises the following steps: determining a forward corpus related to a control keyword of smart home equipment; wherein the forward corpus comprises a first corpus having a control intention for the smart home device; determining a reverse corpus similar to the control keyword through the first phoneme sequence of the control keyword; wherein the reverse corpus comprises a second corpus which does not have a control intention for the smart home equipment; a rejection model is optimized through the forward corpus and the reverse corpus, so that target corpus is rejected to be recognized through the optimized rejection model; wherein the target corpus is a corpus which is sent by a target object in a smart home scene and has no control intention for the smart home equipment, wherein the similarity between the target object and the control keyword meets a set condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of smart homes. Specifically, it relates to a method and device for rejecting recognition of target corpus, a storage medium, and an electronic device. Background Art

[0002] In the application of smart home voice assistants, users can perform voice interactions with smart speakers or other smart home devices bound under a household account. However, the interaction between users and smart home devices is easily affected by noise in the smart home scenario, the user's pronunciation accent, and the incompleteness of the voice assistant algorithm model, resulting in the smart home device being easily started incorrectly or executing incorrect control instructions. This not only brings great potential safety hazards but also seriously affects the user's interaction experience.

[0003] In related technologies, there is no effective solution to the problem that the interaction between users and smart home devices is easily affected by noise in the scenario, the user's pronunciation accent, etc., which in turn leads to the smart home device being started incorrectly or executing incorrect control instructions.

[0004] Therefore, it is necessary to improve related technologies to overcome the above-mentioned defects in related technologies. Summary of the Invention

[0005] Embodiments of this application provide a method and device for rejecting recognition of target corpus, a storage medium, and an electronic device, so as to at least solve the problem in related technologies that the interaction between users and smart home devices is easily affected by noise in the scenario, the user's pronunciation accent, etc., which in turn leads to the smart home device being started incorrectly or executing incorrect control instructions.

[0006] According to one aspect of the embodiments of this application, a method for rejecting recognition of target corpus is provided, including: determining a positive corpus related to control keywords of a smart home device; wherein, the positive corpus includes a first corpus with a control intention for the smart home device; and determining a reverse corpus similar to the control keywords through a first phoneme sequence of the control keywords; wherein, the reverse corpus includes a second corpus without a control intention for the smart home device; optimizing a rejection recognition model through the positive corpus and the reverse corpus, so as to reject and recognize a target corpus through the optimized rejection recognition model; wherein, the target corpus is a corpus that a target object emits in a smart home scenario, has a similarity with the control keywords that meets a set condition, and does not have a control intention for the smart home device.

[0007] In an exemplary embodiment, determining a positive corpus related to a control keyword of a smart home device includes: adding a first slot to the control keyword; wherein, a grammatical attribute corresponding to the first slot is different from a grammatical attribute corresponding to a second slot in the control keyword; replacing a first word of each slot in the control keyword after adding the slot with a second word; wherein the second word is a synonym or near-synonym of the first word; and / or changing a first position order of each slot in the control keyword after adding the slot to a second position order; determining the positive corpus through the control keyword after word replacement and / or the control keyword after order change.

[0008] In an exemplary embodiment, determining a negative corpus similar to the control keyword through a first phoneme sequence of the control keyword includes: obtaining the first phoneme sequence of the control keyword, and obtaining a candidate library for screening the second corpus; wherein, the candidate library includes: a third corpus that does not have a control intention for the smart home device, and a second phoneme sequence corresponding to the third corpus; determining a fourth corpus in the third corpus as the second corpus to obtain the negative corpus through the second corpus; wherein, a similarity between the second phoneme sequence of the fourth corpus and the first phoneme sequence is higher than a target value.

[0009] In an exemplary embodiment, optimizing a rejection recognition model through the positive corpus and the negative corpus to reject and recognize a target corpus through the optimized rejection recognition model includes: determining a corpus warning model through the positive corpus and the negative corpus; screening a fifth corpus obtained through the corpus warning model to obtain a sixth corpus; wherein, the fifth corpus is a corpus uttered by a target object in a smart home scenario; the sixth corpus is the fifth corpus with a warning mark or a non-warning mark, the warning mark is used to indicate that the sixth corpus does not have a control intention, and the non-warning mark is used to indicate that the sixth corpus has a control intention; optimizing the rejection recognition model through the sixth corpus to reject and recognize the target corpus through the optimized rejection recognition model.

[0010] In an exemplary embodiment, screening a fifth corpus obtained through the corpus warning model to obtain a sixth corpus includes: extracting a corpus feature of the fifth corpus according to a corpus type of the fifth corpus, wherein, the corpus feature includes at least one of the following: an audio feature, a text feature, a phoneme feature; determining a likelihood value between the fifth corpus and the corpus warning model through the corpus feature; determining a marking result for the fifth corpus through the likelihood value to obtain the sixth corpus, wherein, the marking result includes one of the following: a warning mark, a non-warning mark.

[0011] In an exemplary embodiment, the rejection recognition model is optimized by the forward corpus and the reverse corpus to reject the recognition of the target corpus by the optimized rejection recognition model, including: obtaining a pre-constructed corpus warning white model, where the corpus warning white model includes: a first channel and a second channel for extracting input data features, and a comparison network respectively connected to the first channel and the second channel; performing audio supplementation operations on the forward corpus and the reverse corpus; determining a corpus warning model through the supplemented forward corpus, the supplemented reverse corpus, and the corpus warning white model; optimizing the rejection recognition model through the corpus warning model to reject the recognition of the target corpus by the optimized rejection recognition model.

[0012] In an exemplary embodiment, determining a corpus warning model through the supplemented forward corpus, the supplemented reverse corpus, and the corpus warning white model includes: inputting the supplemented forward corpus and the supplemented reverse corpus into the first channel of the corpus warning white model; where the first channel includes: a plurality of filters, and a first network respectively connected to the plurality of filters; inputting the seed library where the control keyword is located into the second channel of the corpus warning white model; where the second channel includes: a second network; comparing the first output result of the first channel and the second output result of the second channel through the comparison network, and optimizing the corpus warning white model through the comparison result to obtain the corpus warning model.

[0013] In an exemplary embodiment, performing audio supplementation operations on the forward corpus and the reverse corpus includes: synthesizing a first audio corresponding to a seventh corpus through a speech synthesis tool, and adding noise to the first audio to obtain a second audio; supplementing the second audio to the corpus corresponding to the seventh corpus; where the seventh corpus includes one of the following: the first corpus, the second corpus; obtaining a third audio corresponding to the seventh corpus recorded by the target object in the smart home scenario; supplementing the third audio to the corpus corresponding to the seventh corpus; where the corpus corresponding to the seventh corpus includes one of the following: the forward corpus, the reverse corpus.

[0014] According to another aspect of the embodiments of the present application, there is also provided a rejection recognition device for target corpus, including: a determination module, configured to determine a positive corpus related to the control keywords of the smart home device; wherein, the positive corpus includes a first corpus with a control intention for the smart home device; and determine a reverse corpus similar to the control keywords through a first phoneme sequence of the control keywords; wherein, the reverse corpus includes a second corpus without a control intention for the smart home device; an optimization module, configured to optimize a rejection recognition model through the positive corpus and the reverse corpus, so as to reject and recognize the target corpus through the optimized rejection recognition model; wherein, the target corpus is a corpus that satisfies a set condition in terms of similarity to the control keywords and is uttered by a target object in a smart home scenario and does not have a control intention for the smart home device.

[0015] According to yet another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and wherein the computer program is configured to execute the above-mentioned rejection recognition method for target corpus when running.

[0016] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and wherein the above-mentioned processor executes the above-mentioned rejection recognition method for target corpus through the computer program.

[0017] According to yet another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program, and the computer program realizes the above-mentioned rejection recognition method for target corpus when executed by a processor.

[0018] Through this application, a positive corpus related to the control keywords of the smart home device is determined; wherein, the positive corpus includes a first corpus with a control intention for the smart home device; and a reverse corpus similar to the control keywords is determined through the first phoneme sequence of the control keywords; wherein, the reverse corpus includes a second corpus without a control intention for the smart home device; the rejection model is optimized through the positive corpus and the reverse corpus to reject and identify the target corpus through the optimized rejection model; wherein, the target corpus is a corpus that is uttered by a target object in a smart home scenario, has a similarity to the control keywords that meets a set condition, and does not have a control intention for the smart home device. That is to say, first, a positive corpus and a reverse corpus related to the control keywords are determined, and then the rejection model can be optimized through the positive corpus and the reverse corpus, so that the rejection model can more accurately reject and identify the corpus that is uttered by the target object in the smart home scenario, has a similarity to the control keywords that meets the set condition, and does not have a control intention for the smart home device. In addition, since the phoneme sequence of the control keywords is also introduced when determining the reverse corpus, the rejection model optimized through this reverse corpus can further avoid the interference of the target object's accent. Therefore, by adopting the above technical solution, the problem in the related art that in the interaction between the user (i.e., the target object) and the smart home device, it is easily affected by noises in the scenario and the user's pronunciation accent, etc., resulting in the smart home device being wrongly started or executing wrong control instructions is solved; the high-precision rejection and identification of the corpus uttered by the target object through the rejection model are realized, so that the smart home device will not be wrongly started and will not execute wrong control instructions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments that conform to the present application, and are used together with the specification to explain the principles of the present application.

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic diagram of the hardware environment of an optional method for rejecting and identifying a target corpus according to an embodiment of the present application;

[0022] Figure 2 It is a flowchart of an optional method for rejecting and identifying a target corpus according to an embodiment of the present application;

[0023] Figure 3It is a schematic structural diagram of a corpus early warning model according to an embodiment of the present application;

[0024] Figure 4 It is a structural block diagram of an optional rejection device for target corpus according to an embodiment of the present application;

[0025] Figure 5 It is another structural block diagram of an optional rejection device for target corpus according to an embodiment of the present application. Detailed implementation manners

[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] According to one aspect of the embodiments of the present application, a method for rejecting target corpus is provided. The method for rejecting target corpus is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, and IntelligenceHouse ecosystem. Optionally, in this embodiment, the above method for rejecting target corpus can be applied to, for example, Figure 1 the hardware environment composed of multiple terminal devices 102 and a server 104 as shown in Figure 1As shown in the figure, the server 104 is connected to multiple terminal devices 102 through a network. It can be used to provide services (such as application services, etc.) for the terminals or the clients installed on the terminals. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server 104.

[0029] The above network can include but is not limited to at least one of the following: wired network, wireless network. The above wired network can include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network can include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to being a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projection device, smart TV, smart drying rack, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification device, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc.

[0030] In this embodiment, a method for rejecting recognition of target corpus is provided, including but not limited to being applied to the above terminal devices. Figure 2 It is a flowchart of the method for rejecting recognition of target corpus according to an embodiment of the present application. The process includes the following steps:

[0031] Step S202: Determine a positive corpus related to the control keywords of the smart home device; wherein, the positive corpus includes a first corpus with a control intention for the smart home device; and determine a negative corpus similar to the control keywords through the first phoneme sequence of the control keywords; wherein, the negative corpus includes a second corpus without a control intention for the smart home device.

[0032] Step S204: Optimize the rejection recognition model through the positive corpus and the negative corpus, so as to reject and recognize the target corpus through the optimized rejection recognition model; wherein, the target corpus is the corpus that the target object emits in the smart home scenario, whose similarity with the control keywords meets the set conditions and does not have a control intention for the smart home device.

[0033] After optimizing the rejection recognition model with the forward corpus and the reverse corpus, the corpus collected by the smart home device in the smart home scenario (including noise and all the corpus uttered by the target object in the smart home scenario, where all the corpus includes the above-mentioned target corpus) is input into the rejection recognition model. Then, the rejection recognition model will reject the above-mentioned target corpus and noise, and allow further recognition of the device control corpus uttered by the target object in the smart home scenario, so as to control the smart home device.

[0034] Optionally, the set condition is, for example, that the phoneme similarity is higher than a preset threshold.

[0035] Through the above steps, a forward corpus related to the control keyword of the smart home device is determined; wherein, the forward corpus includes a first corpus with a control intention for the smart home device; and a reverse corpus similar to the control keyword is determined through the first phoneme sequence of the control keyword; wherein, the reverse corpus includes a second corpus without a control intention for the smart home device; the rejection recognition model is optimized with the forward corpus and the reverse corpus to reject the recognition of the target corpus through the optimized rejection recognition model; wherein, the target corpus is the corpus uttered by the target object in the smart home scenario, whose similarity with the control keyword meets the set condition and does not have a control intention for the smart home device. Therefore, by adopting the above technical solution, the problem in the related technology that the interaction between the user and the smart home device is easily affected by noise in the scenario, the user's pronunciation accent, etc., resulting in the smart home device being wrongly started or executing wrong control instructions is solved; the high-precision rejection recognition of the corpus uttered by the target object is realized through the rejection recognition model, so that the smart home device will not be wrongly started and will not execute wrong control instructions.

[0036] In an exemplary embodiment, determining a forward corpus related to the control keyword of the smart home device includes: performing a target operation on the control keyword to obtain the forward corpus, where the target operation includes at least one of the following: slot expansion operation, word order replacement operation, word replacement operation.

[0037] Optionally, an implementation manner of performing a target operation on the control keyword to obtain the positive corpus includes: adding a first slot to the control keyword; wherein, the grammatical attribute corresponding to the first slot is different from the grammatical attribute corresponding to the second slot in the control keyword; replacing the first word of each slot in the control keyword after adding the slot with a second word; wherein the second word is a synonym or near-synonym of the first word; and / or changing the first position order of each slot in the control keyword after adding the slot to a second position order; determining the positive corpus through the control keyword after word replacement and / or the control keyword after order change.

[0038] It should be noted that the control keyword generally consists of a verb + a noun. For example, "open the oven", "open the washing machine", etc. Adding a first slot to the control keyword, where the "slot" refers to the part in the control keyword that can be replaced or changed. For example, in the instruction "open the washing machine", a slot can be added to describe the location or state of the washing machine, such as "open the washing machine in the laundry room", and here "in the laundry room" is the addition of the first slot. By adding information slots such as location, state, and time, more specific and practical control instruction corpora can be generated. The first slot and the second slot (such as "washing machine") should have different grammatical attributes to ensure that the generated corpus is reasonable in grammatical structure and clear in semantics.

[0039] Furthermore, the first word of each slot in the control keyword after adding the slot can be replaced with a second word, where the second word is a synonym or near-synonym of the first word. For example, in the instruction "raise the temperature in the living room", "raise" can be replaced with "increase", or "living room" can be replaced with "sitting room", so that more variant control instructions can be generated to help the model learn different expressions that users may use, enhancing the robustness and coverage of recognition. Furthermore, the first position order in each slot of the control keyword can also be changed to the second position order. For example, changing the instruction "turn on the light in the bedroom" to "turn on the light in the bedroom", although the grammar may not be as natural as the original sentence, such a change can help the model learn the ability to correctly understand instructions in different word orders, especially when dealing with non-native users or users with diverse accents, improving the adaptability and flexibility of the model.

[0040] Through the above steps, a positive corpus containing a large number of device control instructions can be generated. This corpus not only covers the device control keywords themselves, but also includes instructions derived from the keywords that are closer to the actual usage scenarios. The generated positive corpus can be used to train various models in the voice assistant, such as the rejection recognition model, the wake-up model, and the device control model, etc., to help these models learn how to accurately understand and execute the user's device control instructions in a complex and changing environment.

[0041] In an exemplary embodiment, determining a reverse corpus similar to the control keyword through the first phoneme sequence of the control keyword includes: determining the reverse corpus through the first phoneme sequence and a similarity metric method.

[0042] Specifically, determining the reverse corpus through the first phoneme sequence and a similarity metric method includes: obtaining the first phoneme sequence of the control keyword, and obtaining a candidate library for screening the second corpus; wherein, the candidate library includes: a third corpus that does not have a control intention for the smart home device, and a second phoneme sequence corresponding to the third corpus; determining a fourth corpus in the third corpus as the second corpus, so as to obtain the reverse corpus through the second corpus; wherein, the similarity between the second phoneme sequence of the fourth corpus and the first phoneme sequence is higher than a target value.

[0043] First, convert the control keyword of the smart home device into a first phoneme sequence through a phoneme analysis tool. For example, "turn off the dehumidifier" can be converted to "guan b i chu sh i q i", which is the phoneme representation after removing the tones. Secondly, create a corpus that contains a large number of corpora without device control intentions, that is, a third corpus, and the corresponding phoneme sequences. This library can contain conversations in daily life, song lyrics, news broadcasts, etc. that do not involve device control. For example, sentences such as "I like watching TV dramas" and "The weather is nice today", and their phoneme representations. Finally, in the candidate library, by calculating the similarity between the second phoneme sequence of the third corpus and the first phoneme sequence of the control keyword, screen out those fourth corpora whose phoneme sequence similarity is higher than the target value, thereby determining the reverse corpus. This screening process can use various similarity calculation methods (i.e., similarity metric methods), such as Euclidean distance, cosine similarity, or edit distance, etc. The corpora in the reverse corpus are similar to the control keyword at the phoneme level, but do not have the intention of controlling the smart home device at the semantic level.

[0044] For example, if the third corpus is: I like to close the window. (Similarity: 0.74, easily confused device control statement: turn off the dehumidifier, device control phonemes: guan b i chu sh i q i, daily life phonemes: guan b i chuang hu), then if the target value is a value lower than 0.74, "I like to close the window" can be selected into the reverse corpus.

[0045] In an optional embodiment, the rejection recognition model is optimized by the forward corpus and the reverse corpus to reject the recognition of the target corpus by the optimized rejection recognition model, including: determining a corpus warning model by the forward corpus and the reverse corpus; screening the obtained fifth corpus by the corpus warning model to obtain a sixth corpus; wherein, the fifth corpus is the corpus issued by the target object in the smart home scenario; the sixth corpus is the fifth corpus with a warning mark or a non-warning mark, the warning mark is used to indicate that the sixth corpus does not have a control intention, and the non-warning mark is used to indicate that the sixth corpus has a control intention; optimizing the rejection recognition model by the sixth corpus to reject the recognition of the target corpus by the optimized rejection recognition model.

[0046] It can be understood that, based on the forward corpus and the reverse corpus, an initial corpus warning model is constructed in the embodiment of the present application. The goal of this warning model (i.e., the corpus warning model) is to distinguish between the corpus with a control intention and the corpus without a control intention. The warning model can be a binary classification model, such as logistic regression, support vector machine, decision tree or neural network model, which predicts the intention category of the newly input corpus by learning the features in the forward and reverse corpora.

[0047] In the actual usage scenario, the voice assistant in the smart home will receive the fifth corpus issued by the user. These corpora may include device control instructions or other daily conversation contents. The initially constructed corpus warning model is used to screen the fifth corpus, and it is marked with a warning mark or a non-warning mark to indicate whether the fifth corpus has a control intention, so as to obtain the sixth corpus. Furthermore, the sixth corpus constitutes a data set for further optimizing the rejection recognition model, because they contain diverse samples of real user inputs, including both correct device control instructions and non-control corpora that may be misrecognized as control instructions.

[0048] Next, the marked sixth corpus is used to optimize the rejection recognition model. This process may include retraining the model and adjusting parameters to improve the accuracy of the model in recognizing and rejecting corpora. By introducing diverse corpora in the real scenario, the model can learn more details, such as specific pronunciation methods, different contexts and environmental noises, etc., so as to improve its recognition performance in complex environments.

[0049] The optimized rejection recognition model can more accurately distinguish between the corpus with a control intention and the corpus without a control intention, avoid misoperations, and improve the user experience and security of the smart home voice assistant. For example, if the model might previously misrecognize a corpus with similar phonemes to "turn on the washing machine" such as "open the western restaurant" as a control instruction, the optimized model can effectively distinguish between the two, only execute the real device control instructions, and "reject" those corpora without a control intention.

[0050] The optimized rejection recognition model will be used in the actual smart home scenario to make decisions on recognizing and rejecting the corpus issued by users. The target corpus is any voice command or conversation content issued by users in the smart home. According to the recognition rules obtained through training, the model classifies the target corpus. For the corpus marked as not having a control intention, the voice assistant will perform the "rejection recognition" action and not initiate any device control to ensure the security and stability of the system.

[0051] In summary, in this embodiment, a corpus early warning model based on a positive and a negative corpus is constructed and optimized to enhance the recognition and rejection capabilities of the voice assistant for user corpus in the smart home scenario. This method can effectively process various possible voice inputs, avoid misoperations, improve the user experience, and ensure the security and reliability of the smart home system.

[0052] In another alternative embodiment, the rejection recognition model is optimized through the positive corpus and the negative corpus to reject and recognize the target corpus through the optimized rejection recognition model, including: obtaining a pre-constructed corpus early warning blank model, where the corpus early warning blank model includes: a first channel and a second channel for extracting input data features, and a comparison network respectively connected to the first channel and the second channel; performing audio supplementation operations on the positive corpus and the negative corpus; determining a corpus early warning model through the supplemented positive corpus, the supplemented negative corpus, and the corpus early warning blank model; optimizing the rejection recognition model through the corpus early warning model to reject and recognize the target corpus through the optimized rejection recognition model.

[0053] Furthermore, determining a corpus early warning model through the supplemented positive corpus, the supplemented negative corpus, and the corpus early warning blank model includes: inputting the supplemented positive corpus and the supplemented negative corpus into the first channel of the pre-said corpus early warning blank model; where the first channel includes: a plurality of filters, and a first network respectively connected to the plurality of filters; inputting the seed library where the control keywords are located into the second channel of the corpus early warning blank model; where the second channel includes: a second network; comparing a first output result of the first channel and a second output result of the second channel through the comparison network, and optimizing the corpus early warning blank model through the comparison result to obtain the corpus early warning model.

[0054] To enhance the recognition ability of the model, first, audio supplementation operations are performed on the forward corpus and the reverse corpus. The supplemented forward corpus and reverse corpus are input into the first channel of a pre-constructed corpus early warning white model for processing. This channel includes multiple filters and a first network connected thereto. The multiple filters are used to further expand the forward corpus and the reverse corpus, and further supplement the noise related to the smart home scenario, such as the working noise of a washing machine; the role of the first network is to extract phoneme features, text features, audio features (such as phonemes, spectrograms, cepstral coefficients), etc. As Figure 3 shown, the first network can be a Mixture Of Expert NLP network, and these features will be used as the input for subsequent model training. At the same time, the seed library containing the control keywords is input into the second channel of the white model. This channel contains a second network. As Figure 3 shown, the second network can be an Original Model NLP, which is used to process and understand the characteristics of the control keywords themselves. Further, the role of the comparison network is to evaluate the similarity between the outputs of the two channels and further distinguish the intention types of the input corpus. According to the comparison result, the model will optimize itself and adjust the network parameters to improve the recognition rate of the corpus with control intention and the corpus without control intention. This process may involve multiple iterations until the model reaches satisfactory performance on the test set, thereby finally obtaining an optimized corpus early warning model.

[0055] Among them, performing the audio supplementation operation on the forward corpus and the reverse corpus includes: synthesizing the first audio corresponding to the seventh corpus through a speech synthesis tool, and adding noise to the first audio to obtain the second audio; supplementing the second audio into the corpus corresponding to the seventh corpus; where the seventh corpus includes one of the following: the first corpus, the second corpus; obtaining the third audio corresponding to the seventh corpus recorded by the target object in the smart home scenario; supplementing the third audio into the corpus corresponding to the seventh corpus; where the corpus corresponding to the seventh corpus includes one of the following: the forward corpus, the reverse corpus.

[0056] In the early stage of building the model, it is also necessary to expand the forward corpus and the reverse corpus to ensure the diversity and representativeness of the model training data. This includes using a speech synthesis tool to generate audio corresponding to the corpus and adding noise to these synthesized audios to simulate the actual environment; and it can also include recording the corpus audio emitted by the target object in the real home scenario. The target object can be users of different ages and different accents to ensure that the model can adapt to a wide range of speech inputs.

[0057] In an exemplary embodiment, the obtained fifth corpus is screened by the corpus early warning model to obtain a sixth corpus, including: extracting the corpus features of the fifth corpus according to the corpus type of the fifth corpus, where the corpus features include at least one of the following: audio features, text features, phoneme features; determining the likelihood value between the fifth corpus and the corpus early warning model through the corpus features; determining the marking result of the fifth corpus through the likelihood value to obtain the sixth corpus, where the marking result includes one of the following: early warning mark, non-early warning mark.

[0058] In the usage stage of the corpus early warning model, it is first necessary to extract features from the obtained fifth corpus, and these features include audio features, text features, and phoneme features. Audio features can be spectrograms, Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding (LPC), or other statistical features extracted from audio signals; text features are derived from the text transcription of the corpus and can be bag-of-words models, TF-IDF (Term Frequency-Inverse Document Frequency) weights, or word embedding vectors; phoneme features refer to the phoneme sequences extracted from the transcribed text, which are the key to speech recognition and semantic understanding.

[0059] Next, the extracted corpus features are input into the corpus early warning model to determine the likelihood value between the fifth corpus and the model. The likelihood value output by the corpus early warning model actually reflects the matching degree of the corpus with the positive (having control intention) and negative (not having control intention) corpus features learned by the model. Specifically: If the fifth corpus highly matches the features in the positive corpus, the model will output a relatively high likelihood value, indicating that the corpus is very likely to have a control intention. This likelihood value can be understood as the probability that the corpus belongs to the positive category. If the fifth corpus more matches the features in the negative corpus, the likelihood value output by the model will be lower, indicating that the corpus may not have a control intention. The likelihood value can be regarded as the probability that the corpus belongs to the negative category. In practical applications, a threshold (such as the third value) can be set to determine whether the likelihood value output by the model is sufficient to indicate that the corpus to be identified has a control intention. If the likelihood value is higher than the third value, the corpus will be marked as having a control intention and given a non-early warning mark; if the likelihood value is less than or equal to the third value, the corpus will be marked as not having a control intention and given an early warning mark.

[0060] Optionally, in another implementation scenario, if the corpus warning model is trained only with the reverse corpus, if the likelihood value obtained by inputting the fifth corpus is higher than the fourth value, it indicates that the fifth corpus may not have a control intention, and a warning mark is assigned to the fifth corpus.

[0061] In summary, the likelihood value output by the corpus warning model trained with the positive corpus and the reverse corpus is a decision-making index for judging whether the corpus to be identified has the intention to control smart home devices. Combining the likelihood value with the decision boundary of model training and the preset threshold (the third value) can effectively achieve the classification of the corpus, ensure that the intelligent home voice assistant can accurately identify and execute the user's control commands in actual applications, and at the same time "reject the recognition" of those corpora without control intention, thereby improving the security and user experience of the system.

[0062] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. To better understand the above method for rejecting the target corpus, the following will describe the above process in conjunction with embodiments, but it is not used to limit the technical solutions of the embodiments of the present application. Specifically:

[0063] The embodiments of the present application propose to generate a detection sample set (equivalent to the above-mentioned reverse corpus and positive corpus) through the principle of phoneme similarity of the corpus for the rejection model and other system-related functions in the voice interaction system. The samples in this set include basic corpora and the corresponding audio of the corpora. Specifically, it includes: defining keywords for controlling smart home devices and obtaining the phoneme representations of these keywords (i.e., the above-mentioned first phoneme sequence); generating commonly used corpora in life and searching for words with similar pronunciations to the keywords for controlling smart home devices (i.e., the second corpus); generating meaningful sentences and saving them to files; generating voices through tool synthesis and adding noise or manual recording (i.e., the voice supplement operation);

[0064] Input the generated voice into the voice assistant system for testing. The generated audio can be combined with the algorithm of the dialogue system through multi-dimensional filtering processing (equivalent to the above-mentioned constructed corpus warning model). First, a large number of audio (equivalent to the supplemented positive and reverse corpora) are automatically generated and compared with the marked results of the seed audio to automatically detect incorrect audio samples (equivalent to giving a warning mark); analyze and correct the detected incorrect samples. Optionally, the construction and use of the corpus warning model may include: constructing a simple convolutional neural network (CNN) or recurrent neural network (RNN), training the model with the prepared data, evaluating the performance of the model, classifying new voice inputs, etc.

[0065] That is to say, after verifying the supplemented positive and negative corpora, they can be used as a training set to train a corpus warning model for optimizing the rejection model, thereby completing the closed loop of automated testing. The supplemented positive and negative corpora can also be used as the enhanced learning training sets for the wake-up model and the intent parsing model in the voice assistant. In summary, in the embodiments of the present application, by using the principle of phoneme similarity, the discovery and application of similarity confusion can be carried out from multiple dimensions such as text, voice, and other audio signals, so as to realize error detection in the entire voice interaction link and model enhanced learning in the link.

[0066] Specifically, define the control keywords of the smart home device (constituting the seed bank), and obtain the phoneme representations of these keywords (i.e., the first phoneme sequence), removing the tones. For example:

[0067] Turn on the oven -> Open the oven; Turn on the washing machine -> Open the washing machine; Set the refrigerator -> Set the refrigerator; Increase the temperature of the oven -> Increase the temperature of the oven; Start the dehumidifier -> Start the dehumidifier; Pause the floor cleaning robot -> Pause the floor cleaning robot; Restart the light -> Restart the light; Query the TV -> Query the TV; Restore the camera -> Restore the camera; Query the microwave oven -> Query the microwave oven; Start the washing machine -> Start the washing machine; Query the air conditioner -> Query the air conditioner; Restart the floor cleaning robot -> Restart the floor cleaning robot; Unlock the washing machine -> Unlock the washing machine; Open the air purifier -> Open the air purifier; Turn off the microwave oven -> Turn off the microwave oven; Restore the floor cleaning robot -> Restore the floor cleaning robot; Unlock the oven -> Unlock the oven; Unlock the camera -> Unlock the camera; Pause the air conditioner -> Pause the air conditioner; Unlock the smart socket -> Unlock the smart socket; Increase the temperature of the dehumidifier -> Increase the temperature of the dehumidifier; Restore the water heater -> Restore the water heater; Turn off the refrigerator -> Turn off the refrigerator; Decrease the power of the smart socket -> Decrease the power of the smart socket; Decrease the power of the coffee machine -> Decrease the power of the coffee machine; Decrease the brightness of the light -> Decrease the brightness of the light; Lock the floor cleaning robot -> Lock the floor cleaning robot; Stop the water heater -> Stop the water heater.

[0068] Construct a confusion-prone corpus (i.e., reverse corpus) based on control keywords, including: creating a candidate library containing common words or other audio in life and their phoneme representations, calculating the similarity with a specific scenario corpus dictionary through the similarity of phoneme sequences, such as using the Euclidean distance algorithm, and screening through a similarity threshold. Those higher than the specified threshold (i.e., the target value) can enter the confusion-prone corpus. For example:

[0069] I like to close the window (Similarity: 0.74, Confusing device control statement: Turn off the dehumidifier, Device control phoneme: guan bi chu shi qi, Daily life phoneme: guan bi chuang hu);

[0070] Lock the screen (Similarity: 0.78, Confusing device control statement: Lock the air conditioner, Device control phoneme: suo ding kong diao, Daily life phoneme: suo ding ping mian);

[0071] Light (Similarity: 0.76, Confusing device control statement: Turn down the light, Device control phoneme: diao di deng, Daily life phoneme: dian deng);

[0072] Resume playback (Similarity: 0.73, Confusing device control statement: Resume the sound system, Device control phoneme: hui fu yin xiang, Daily life phoneme: hui fu bo fang);

[0073] Check the weather (Similarity: 0.73, Confusing device control statement: Check the dehumidifier, Device control phoneme: cha xun chu shi qi, Daily life phoneme: cha xun tian qi);

[0074] Stop running (Similarity: 0.73, Confusing device control statement: Stop the oven, Device control phoneme: ting zhi kao xiang, Daily life phoneme: ting zhi pao bu);

[0075] Open the schoolbag (Similarity: 0.71, Confusing device control statement: Turn on the water heater, Device control phoneme: da kai re shui qi, Daily life phoneme: da kai shu bao);

[0076] Pause playback (Similarity: 0.71, Confusing device control statement: Pause the air conditioner, Device control phoneme: zan ting kong diao, Daily life phoneme: zan ting bo fang);

[0077] Turn down the volume (Similarity: 0.71, Confusing device control statement: Turn down the oven, Device control phoneme: tiao di kao xiang, Daily life phoneme: tiao di yin liang);

[0078] Close the window (Similarity: 0.71, Confusing device control statements: Turn off the fan, Device control phonemes: guanbi feng shan, Daily life phonemes: guan bi chuang hu);

[0079] Generate an extended library (i.e., the positive corpus) that includes more generalized device control corpora based on the seed library. For example: Turn on the oven -> Turn on the oven in the kitchen; Turn on the washing machine -> Turn on the washing machine in the laundry room; Set the refrigerator -> Set the refrigerator in the living room; Turn up the oven -> Turn up the oven in the kitchen; Start the dehumidifier -> Start the dehumidifier in the bedroom; Pause the floor cleaning robot -> Pause the floor cleaning robot in the living room; Restart the light -> Restart the light in the bedroom; Query the TV -> Query the TV in the living room; Resume the camera -> Resume the camera at the door; Query the microwave oven -> Query the microwave oven in the kitchen; Start the washing machine -> Start the washing machine in the laundry room; Query the air conditioner -> Query the air conditioner in the bedroom; Restart the floor cleaning robot -> Restart the floor cleaning robot in the living room; Unlock the washing machine -> Unlock the washing machine in the laundry room; Turn on the air purifier -> Turn on the air purifier in the living room; Turn off the microwave oven -> Turn off the microwave oven in the kitchen; Resume the floor cleaning robot -> Resume the floor cleaning robot in the living room; Unlock the oven -> Unlock the oven in the kitchen; Unlock the camera -> Unlock the camera at the door; Pause the air conditioner -> Pause the air conditioner in the bedroom; Unlock the smart socket -> Unlock the smart socket in the living room; Turn up the dehumidifier -> Turn up the dehumidifier in the bedroom; Resume the water heater -> Resume the water heater in the bathroom; Turn off the refrigerator -> Turn off the refrigerator in the living room; Turn down the smart socket -> Turn down the smart socket in the living room; Turn down the coffee machine -> Turn down the coffee machine in the kitchen; Turn down the light -> Turn down the light in the bedroom; Lock the floor cleaning robot -> Lock the floor cleaning robot in the living room; Stop the water heater -> Stop the water heater in the bathroom.

[0080] Through the above steps, the confusing corpus, the seed library, and the extended library are obtained. Optionally, other large-scale corpora that select similar pronunciations (i.e., obtain the candidate library) are exemplified by databases of audio media such as song lyrics, film and television works, news, etc. In addition to the confusing pronunciation corpora of device control, the above algorithm is also applicable to trigger scenarios based on specific sounds such as wake-up words and specific words. Examples of wake-up words: Some segments of the opening music in the News Broadcast may cause mis-wake-up of wake-up words similar to "Xiaoyou", thus resulting in poor user experience.

[0081] Optionally, there are two ways to generate the corpus voice (i.e., audio supplement operations): 1) Tool synthesis and adding noise: Use a speech synthesis tool to generate speech and add noise to simulate the actual usage scenario. 2) Manual recording in a real home scenario: Record speech in a real home scenario to obtain more realistic test data.

[0082] Based on the above corpus and the generated speech, the end-to-end speech test is performed as follows: 1) Generate a confusing corpus. 2) Generate speech through tool synthesis and adding noise or manual recording. 3) Input the generated speech into the speech assistant system for testing. 4) Compare with the labeled results of the seed audio for consistency to detect incorrect audio samples. 5) Analyze and correct the detected incorrect samples.

[0083] During the process of performing the end-to-end speech test, when inputting the generated speech into the speech assistant system for testing, it is necessary to construct a corpus warning model. As Figure 3 shown, the corpus warning model can select the MOE NLP model. The specific training process of the corpus warning model includes:

[0084] 1) Training set preparation: Prepare the wav files corresponding to the confusing corpus, the seed library, and the extended library corpus as training data. 2) Feature extraction: Extract features (such as MFCC Mel Frequency Cepstral Coefficients) from the audio files. 3) Model construction: Construct a simple Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN) (which can be regarded as another case of the corpus warning white model). 4) Model training: Train the model using the prepared data. 5) Model evaluation: Evaluate the model performance. 6) Model inference: Classify new speech inputs and label non-smart home control speech as "rejected recognition" (equivalent to the above warning label). Finally, the obtained corpus warning model can be used for the above end-to-end test and can also be used to screen training data for the rejected recognition model, etc.

[0085] In summary, the embodiment of the present application effectively improves the accuracy and user experience of the smart home speech assistant through the construction and modeling method of the confusing corpus test set based on the phoneme similarity principle. By generating a large number of confusing corpora and conducting automated tests, errors can be quickly discovered and corrected to ensure the stability and reliability of the system. The establishment of the confusing corpus can be used as the enhanced learning iteration data of multiple sub-module models such as wake-up, rejected recognition, and intent parsing in the speech interaction assistant for model iteration optimization, thereby closed-loop optimizing the interaction performance and user experience of the speech assistant.

[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0087] In this embodiment, a rejection recognition device for target corpus is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0088] Figure 4 It is a structural block diagram of an optional rejection recognition device for target corpus according to an embodiment of the present application. The device includes:

[0089] A determination module 42, configured to determine a positive corpus related to a control keyword of a smart home device; wherein, the positive corpus includes a first corpus with a control intention for the smart home device; and determine a reverse corpus similar to the control keyword through a first phoneme sequence of the control keyword; wherein, the reverse corpus includes a second corpus without a control intention for the smart home device;

[0090] An optimization module 44, configured to optimize a rejection recognition model through the positive corpus and the reverse corpus, so as to reject and recognize a target corpus through the optimized rejection recognition model; wherein, the target corpus is a corpus emitted by a target object in a smart home scenario, whose similarity to the control keyword meets a set condition and does not have a control intention for the smart home device.

[0091] Through the above device, a positive corpus related to a control keyword of a smart home device is determined; wherein, the positive corpus includes a first corpus with a control intention for the smart home device; and a reverse corpus similar to the control keyword is determined through a first phoneme sequence of the control keyword; wherein, the reverse corpus includes a second corpus without a control intention for the smart home device; the rejection recognition model is optimized through the positive corpus and the reverse corpus, so as to reject and recognize a target corpus through the optimized rejection recognition model; wherein, the target corpus is a corpus emitted by a target object in a smart home scenario, whose similarity to the control keyword is higher than meeting a set condition and does not have a control intention for the smart home device. Therefore, by adopting the above technical solution, the problem in the related art that the interaction between the user and the smart home device is easily affected by noises in the scenario, the user's pronunciation accent, etc., and thus the smart home device is wrongly started or executes wrong control instructions is solved; high-precision rejection recognition of the corpus emitted by the target object is achieved through the rejection recognition model, so that the smart home device will not be wrongly started and will not execute wrong control instructions.

[0092] In an exemplary embodiment, the determination module 42 is further configured to: add a first slot to the control keyword; wherein, the grammatical attribute corresponding to the first slot is different from the grammatical attribute corresponding to a second slot in the control keyword; replace the first word of each slot in the control keyword after adding the slot with a second word; wherein the second word is a synonym or near-synonym of the first word; and / or change the first position order of each slot in the control keyword after adding the slot to a second position order; determine the positive corpus through the control keyword after word replacement and / or the control keyword after order change.

[0093] In an exemplary embodiment, the determination module 42 is further configured to: obtain the first phoneme sequence of the control keyword, and obtain a candidate library for screening the second corpus; wherein, the candidate library includes: a third corpus that does not have a control intention for the smart home device, and a second phoneme sequence corresponding to the third corpus; determine a fourth corpus in the third corpus as the second corpus, so as to obtain the reverse corpus through the second corpus; wherein, the similarity between the second phoneme sequence of the fourth corpus and the first phoneme sequence is higher than a target value.

[0094] In an exemplary embodiment, the optimization module 44 is further configured to: determine a corpus early warning model through the positive corpus and the reverse corpus; screen a fifth corpus obtained through the corpus early warning model to obtain a sixth corpus; wherein, the fifth corpus is a corpus uttered by a target object in a smart home scenario; the sixth corpus is the fifth corpus with a warning mark or a non-warning mark, the warning mark is used to indicate that the sixth corpus does not have a control intention, and the non-warning mark is used to indicate that the sixth corpus has a control intention; optimize the rejection recognition model through the sixth corpus, so as to reject and recognize a target corpus through the optimized rejection recognition model.

[0095] In an exemplary embodiment, the optimization module 44 is further configured to: obtain a pre-constructed corpus early warning white model, wherein, the corpus early warning white model includes: a first channel and a second channel for extracting input data features, and a comparison network respectively connected to the first channel and the second channel; perform an audio supplement operation on the positive corpus and the reverse corpus; determine a corpus early warning model through the supplemented positive corpus, the supplemented reverse corpus and the corpus early warning white model; optimize the rejection recognition model through the corpus early warning model, so as to reject and recognize a target corpus through the optimized rejection recognition model.

[0096] In an exemplary embodiment, the optimization module 44 is further configured to: input the supplemented positive corpus and the supplemented negative corpus into the first channel of the corpus early warning white model; wherein, the first channel includes: a plurality of filters, and a first network respectively connected to the plurality of filters; input the seed library where the control keyword is located into the second channel of the corpus early warning white model; wherein, the second channel includes: a second network; compare the first output result of the first channel and the second output result of the second channel through the comparison network, and optimize the corpus early warning white model according to the comparison result to obtain the corpus early warning model.

[0097] In an exemplary embodiment, as Figure 5 shown, the device further includes an audio supplement module 46, configured to: synthesize a first audio corresponding to the seventh corpus through a speech synthesis tool, and add noise to the first audio to obtain a second audio; supplement the second audio to the corpus corresponding to the seventh corpus; wherein, the seventh corpus includes one of the following: the first corpus, the second corpus; obtain a third audio corresponding to the seventh corpus recorded by the target object in the smart home scenario; supplement the third audio to the corpus corresponding to the seventh corpus; wherein, the corpus corresponding to the seventh corpus includes one of the following: the positive corpus, the negative corpus.

[0098] In an exemplary embodiment, the optimization module 44 is further configured to: extract the corpus feature of the fifth corpus according to the corpus type of the fifth corpus, wherein the corpus feature includes at least one of the following: audio feature, text feature, phoneme feature; determine the likelihood value between the fifth corpus and the corpus early warning model through the corpus feature; determine the marking result of the fifth corpus through the likelihood value to obtain the sixth corpus, wherein the marking result includes one of the following: early warning mark, non-early warning mark.

[0099] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0100] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc and other various media that can store computer programs.

[0101] For the specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary embodiments, and details thereof will not be repeated herein.

[0102] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.

[0103] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. The transmission device is connected to the processor, and the input / output device is connected to the processor.

[0104] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0105] For the specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary embodiments, and details thereof will not be repeated herein.

[0106] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps of them can be made into a single integrated circuit module to implement. Thus, the present application is not limited to any specific combination of hardware and software.

[0107] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present application.

Claims

1. A method for rejecting recognition of a target corpus, characterized in that, Including: Determine a positive corpus related to the control keywords of the smart home device; wherein, the positive corpus includes a first corpus with a control intention for the smart home device; and Determine a reverse corpus similar to the control keyword through the first phoneme sequence of the control keyword; wherein, the reverse corpus includes a second corpus without a control intention for the smart home device; Optimize the rejection recognition model through the positive corpus and the reverse corpus, so as to reject and recognize the target corpus through the optimized rejection recognition model; wherein, the target corpus is a corpus issued by a target object in a smart home scenario, whose similarity to the control keyword meets a set condition and does not have a control intention for the smart home device.

2. The rejection method for the target corpus according to claim 1, characterized in that Determine a positive corpus related to the control keywords of the smart home device, including: Add a first slot to the control keyword; wherein, the grammatical attribute corresponding to the first slot is different from the grammatical attribute corresponding to the second slot in the control keyword; Replace the first word of each slot in the control keyword after adding the slot with a second word; wherein, the second word is a synonym or near-synonym of the first word; and / or Change the first position order of each slot in the control keyword after adding the slot to a second position order; Determine the positive corpus through the control keyword after word replacement and / or the control keyword after order change.

3. The rejection method of the target corpus according to claim 1, characterized in that Determine a reverse corpus similar to the control keyword through the first phoneme sequence of the control keyword, including: Obtain the first phoneme sequence of the control keyword, and obtain a candidate library for screening the second corpus; wherein, the candidate library includes: a third corpus without a control intention for the smart home device, and a second phoneme sequence corresponding to the third corpus; Determine the fourth corpus in the third corpus as the second corpus, so as to obtain the reverse corpus through the second corpus; wherein, the similarity between the second phoneme sequence of the fourth corpus and the first phoneme sequence is higher than a target value.

4. The method for rejecting recognition of the target corpus according to claim 1, wherein, Optimize the rejection recognition model through the positive corpus and the reverse corpus, so as to reject and recognize the target corpus through the optimized rejection recognition model, including: Determine a corpus warning model through the positive corpus and the reverse corpus; Screen the obtained fifth corpus through the corpus warning model to obtain a sixth corpus; wherein, the fifth corpus is a corpus issued by a target object in a smart home scenario; the sixth corpus is the fifth corpus with a warning mark or a non-warning mark, the warning mark is used to indicate that the sixth corpus does not have a control intention, and the non-warning mark is used to indicate that the sixth corpus has a control intention; Optimize the rejection recognition model through the sixth corpus, so as to reject and recognize the target corpus through the optimized rejection recognition model.

5. The rejection method for the target corpus according to claim 4, characterized in that Screen the obtained fifth corpus through the corpus warning model to obtain a sixth corpus, including: Extract the corpus feature of the fifth corpus according to the corpus type of the fifth corpus; wherein, the corpus feature includes at least one of the following: audio feature, text feature, phoneme feature; Determine the likelihood value between the fifth corpus and the corpus warning model based on the corpus features; Determine the marking result for the fifth corpus based on the likelihood value to obtain the sixth corpus; wherein, the marking result includes one of the following: warning mark, non-warning mark.

6. The rejection method for the target corpus according to claim 1, characterized in that, Optimize the rejection recognition model through the positive corpus and the negative corpus to reject the target corpus through the optimized rejection recognition model, including: Obtain a pre-constructed corpus warning white model, wherein the corpus warning white model includes: a first channel and a second channel for extracting input data features, and a comparison network respectively connected to the first channel and the second channel; Perform audio supplementation operations on the positive corpus and the negative corpus; Determine a corpus warning model through the supplemented positive corpus, the supplemented negative corpus, and the corpus warning white model; Optimize the rejection recognition model through the corpus warning model to reject the target corpus through the optimized rejection recognition model.

7. The rejection method for the target corpus according to claim 6, characterized in that, Determine a corpus warning model through the supplemented positive corpus, the supplemented negative corpus, and the corpus warning white model, including: Input the supplemented positive corpus and the supplemented negative corpus into the first channel of the corpus warning white model; wherein, the first channel includes: a plurality of filters, and a first network respectively connected to the plurality of filters; Input the seed library where the control keyword is located into the second channel of the corpus warning white model; Wherein, the second channel includes: a second network; Compare the first output result of the first channel and the second output result of the second channel through the comparison network, and optimize the corpus warning white model through the comparison result to obtain the corpus warning model.

8. The method for rejecting recognition of the target corpus according to claim 6, characterized in that, Perform audio supplementation operations on the positive corpus and the negative corpus, including: Synthesize a first audio corresponding to the seventh corpus through a speech synthesis tool, and add noise to the first audio to obtain a second audio; supplement the second audio to the corpus corresponding to the seventh corpus; wherein, the seventh corpus includes one of the following: the first corpus, the second corpus; Obtain a third audio corresponding to the seventh corpus recorded by the target object in the smart home scenario; supplement the third audio to the corpus corresponding to the seventh corpus; wherein, the corpus corresponding to the seventh corpus includes one of the following: the positive corpus, the negative corpus.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 7.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.

Citation Information

Cited By

  • Speech recognition method and device and terminal equipment

    CN121838738A

  • A speech recognition method, device and terminal equipment

    CN121838738B