Semantic analysis method and device based on intelligent earphone, intelligent earphone and medium

By verifying voiceprint information in smart headphones, preprocessing voice signals and using semantic analytical models, the problem that smart headphones cannot understand user voice intentions is solved, and a better user experience and feature-rich smart headphones are achieved.

CN120220675APending Publication Date: 2025-06-27SHENZHEN AIRSMART TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343071.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing smart headphones cannot effectively understand the complex intentions of user voice information, resulting in poor user experience.

Method used

By obtaining the user's wake-up voice, verifying the voiceprint information, allowing voice interaction, preprocessing the voice information to reduce noise interference, and input the preprocessed voice information into the semantic analysis model to obtain semantic analysis results.

Benefits of technology

It realizes the accurate understanding of the user's voice information intentions by smart headphones, enhances the user experience, and provides richer functions, such as music playback, multi-wheel dialogue and headphone control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220675A_ABST
    Figure CN120220675A_ABST
Patent Text Reader

Abstract

The invention relates to a semantic analysis method and device based on an intelligent earphone, the intelligent earphone and a storage medium, and the method comprises the steps: obtaining a wake-up voice sent by a user, verifying the voiceprint information of the wake-up voice, and allowing the user to carry out voice interaction with the intelligent earphone if the voiceprint information passes verification, and obtaining voice information of voice interaction between a user and the intelligent earphone, performing preprocessing operation on the voice information, inputting the preprocessed voice information into the semantic analysis model, and obtaining a semantic analysis result corresponding to the voice information. The intention of the voice information of the user can be accurately understood through the intelligent earphone, so that the intelligent earphone can execute the instruction corresponding to the voice information according to the intention of the user, the method is suitable for multiple scenes such as music playing, multi-round dialogue and earphone control, and the functions of the intelligent earphone are enriched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent earphones, and in particular, to a semantic parsing method, device, intelligent earphone and storage medium based on intelligent earphones. Background Art

[0002] With the popularization of intelligent wearable devices, intelligent earphones, as a convenient audio playback device, are favored by users. In different environments and scenarios, users' usage requirements for intelligent earphones are different. Currently, most intelligent earphones on the market only provide simple audio playback functions, or select corresponding playback content according to the connected mobile phone device, or users can achieve basic control of intelligent earphones through voice, but cannot understand more complex intentions of users' voice information, resulting in a poor user experience.

[0003] Therefore, how to understand the complex intentions of users' voice information through intelligent earphones has become a technical problem that needs to be solved urgently by those skilled in the art. It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] In view of the above, this application provides a semantic parsing method, device, intelligent earphone and storage medium based on intelligent earphones, aiming to solve the above technical problems.

[0005] In a first aspect, this application provides a semantic parsing method based on intelligent earphones. The method is applied to intelligent earphones and includes:

[0006] Obtain the wake-up voice issued by the user, and verify the voiceprint information of the wake-up voice;

[0007] If the voiceprint information is verified successfully, allow the user to perform voice interaction with the intelligent earphone;

[0008] Obtain the voice information of the user's voice interaction with the intelligent earphone, and perform a preprocessing operation on the voice information;

[0009] Input the preprocessed voice information into a semantic parsing model to obtain a semantic parsing result corresponding to the voice information.

[0010] In a second aspect, this application provides a semantic parsing device based on intelligent earphones. The device includes:

[0011] Verification module: used to obtain the wake-up voice issued by the user, and verify the voiceprint information of the wake-up voice;

[0012] Permission module: configured to, if the voiceprint information is verified successfully, permit the user to perform voice interaction with the intelligent earphone;

[0013] Preprocessing module: configured to obtain the voice information of the user performing voice interaction with the intelligent earphone and perform preprocessing operations on the voice information;

[0014] Parsing module: configured to input the preprocessed voice information into a semantic parsing model to obtain a semantic parsing result corresponding to the voice information.

[0015] In a third aspect, the present application provides an intelligent earphone, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0016] The memory is used to store a computer program;

[0017] The processor is configured to, when executing the program stored on the memory, implement the semantic parsing method based on an intelligent earphone according to any one of the embodiments in the first aspect.

[0018] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the semantic parsing method based on an intelligent earphone according to any one of the embodiments in the first aspect.

[0019] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:

[0020] The present application obtains the wake-up voice issued by the user and verifies the voiceprint information of the wake-up voice. Only after the voiceprint information is verified successfully is the user permitted to perform voice interaction with the intelligent earphone, which can prevent unauthorized users from accessing the earphone functions and protect the user's privacy and data security. Through voice feedback and corresponding interaction designs, the convenience and comfort of the user using the intelligent earphone are enhanced. By obtaining the voice information of the user performing voice interaction with the intelligent earphone and performing preprocessing operations on the voice information, the interference of environmental noise can be reduced and the clarity of the voice information can be enhanced. By inputting the preprocessed voice information into a semantic parsing model to obtain a semantic parsing result corresponding to the voice information, the intelligent earphone can accurately understand the intention of the user's voice information, enabling the intelligent earphone to execute the instruction corresponding to the voice information according to the user's intention, applicable to multiple scenarios such as music playback, multi-round conversations, and earphone control, enriching the functions of the intelligent earphone. Description of the Drawings

[0021] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a schematic flowchart of an embodiment of the semantic parsing method based on an intelligent earphone of the present application;

[0024] Figure 2 It is a schematic module diagram of an embodiment of the semantic parsing device based on an intelligent earphone of the present application;

[0025] Figure 3 It is a schematic diagram of an embodiment of the intelligent earphone of the present application;

[0026] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0027] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0028] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0029] The present application provides a semantic parsing method based on an intelligent earphone. Referring to Figure 1 As shown, it is a schematic flowchart of the method of an embodiment of the semantic parsing method based on an intelligent earphone of the present application. This method can be executed by an intelligent earphone, and the intelligent earphone can be implemented by software and / or hardware. The semantic parsing method based on an intelligent earphone includes:

[0030] Step S10: Obtain the wake-up voice issued by the user, and verify the voiceprint information of the wake-up voice;

[0031] Step S20: If the voiceprint information is verified, allow the user to perform voice interaction with the intelligent earphone;

[0032] Step S30: Obtain the voice information of the user's voice interaction with the intelligent earphone, and perform a preprocessing operation on the voice information;

[0033] Step S40: Input the preprocessed voice information into a semantic parsing model to obtain a semantic parsing result corresponding to the voice information.

[0034] The intelligent earphone of the present application can be connected to an AI large model. Users can control the related operations of the earphone through voice. The cabin of the intelligent earphone has a 4G card insertion function and can be connected to an exclusive audio content APP. Therefore, when users use the intelligent earphone, they can obtain the audio content they want to listen to at any time and anywhere without relying on devices such as mobile phones. The intelligent earphone can use 4G and Bluetooth communication systems. When interacting with users by voice, it preferentially occupies the 4G uplink bandwidth, and the audio streaming media uses the Bluetooth channel.

[0035] The intelligent earphone listens to the surrounding sounds in real time through a highly sensitive microphone and obtains the user's wake-up voice. The wake-up voice is usually set to a specific word. Users can use a preset voice wake-up word (for example: "Kitty Kitty") to wake up the intelligent earphone. The digital signal processing algorithm (DSP) is used to extract the voiceprint features of the wake-up voice in real time. The voiceprint features can be Mel frequency cepstral coefficients and linear predictive coding features, which can characterize the user's voice characteristics.

[0036] After obtaining the wake-up voice issued by the user, verify the voiceprint information of the wake-up voice. The purpose of the verification is to ensure that only registered users can activate the voice interaction function of the earphone to prevent unauthorized access, thereby protecting the user's information security and privacy information.

[0037] If the voiceprint information verification is passed, the user is allowed to have a voice interaction with the intelligent earphone to ensure that the intelligent voice interaction of the earphone is not misused. After successfully verifying the user's identity, the earphone will issue a voice prompt (for example, "Hello, please tell me your needs"), and at the same time turn on the voice recognition mode to prepare to receive further voice information instructions. Conversely, if the voiceprint information verification fails, the user is refused to have a voice interaction with the intelligent earphone. Specifically, verifying the voiceprint information of the wake-up voice includes:

[0038] Extract the target voiceprint features of the wake-up voice;

[0039] Match the target voiceprint features with a pre-constructed voiceprint feature library;

[0040] If there is a voiceprint feature in the voiceprint feature library that matches the target voiceprint feature, the voiceprint information verification is passed.

[0041] The extracted target voiceprint features can be Mel Frequency Cepstral Coefficients (MFCCs) and Linear Predictive Coding (LPC) features. The pre - constructed voiceprint feature library stores the voiceprint information of users who are allowed to perform voice interactions with the smart headset. By using a distance metric algorithm, the similarity between the target voiceprint features and each voiceprint feature in the voiceprint feature library is calculated. If the similarity is greater than a preset threshold (e.g., 95%), the verification of the voiceprint information for waking up the voice is passed.

[0042] After the verification of the voiceprint information is passed, the user can perform further voice interactions with the smart headset, and the smart headset obtains the voice information of the user's voice interaction. Since the voice information obtained by the smart headset may be affected by noise, it is necessary to perform pre - processing operations on the voice information. The smart headset can use noise suppression technology to remove background music and human voice interference in the environment, and then extract clear voice information, which can improve the accuracy of the subsequent voice information parsing results. Specifically, the pre - processing operation performed on the voice information includes:

[0043] Perform a noise removal operation on the voice information to obtain denoised voice information;

[0044] Perform a voice enhancement operation on the denoised voice information to obtain pre - processed voice information.

[0045] The noise removal operation can be to perform a Fast Fourier Transform (FFT) on the voice information, convert the voice information from a time - domain signal to a frequency - domain signal, extract the noise components from a pre - established noise model, subtract the estimated noise components from the mixed spectrum and retain the frequency - domain information, and then perform an inverse Fast Fourier Transform on the denoised frequency - domain signal to restore it to a time - domain signal, thereby obtaining the denoised voice information.

[0046] After obtaining the denoised voice information, although the voice quality of the voice information has been improved, due to possible problems such as low volume or insufficient clarity of the voice signal, it is still necessary to perform a voice enhancement operation on the denoised voice information, and use the voice after the voice enhancement operation as the pre - processed voice information. The voice enhancement operation can improve the subsequent semantic parsing accuracy by increasing the sound pressure level, improving the clarity and brightness of the voice signal. Among them, the voice enhancement operation can use Linear Predictive Coding (LPC) and filtering processing algorithms. By using a linear prediction filter to enhance the timbre characteristics of the voice and using a high - pass filter for high - frequency enhancement, the clarity of the voice can be improved.

[0047] Input the preprocessed voice information into the semantic parsing model to obtain the semantic parsing result corresponding to the voice information. The semantic parsing model can accurately understand the user's intention, enabling the smart headset to respond correctly according to the user's intention. The semantic parsing model can be trained based on an LSTM model or a Transformer deep learning model, and can capture the context information and semantic relationships of the user's voice information.

[0048] In this application, by obtaining the wake-up voice emitted by the user and verifying the voiceprint information of the wake-up voice, the user is only allowed to perform voice interaction with the smart headset after the voiceprint information verification passes, which can prevent unauthorized users from accessing the headset functions and protect the user's privacy and data security. Through voice feedback and corresponding interaction designs, the convenience and comfort of using the smart headset by the user are enhanced. Obtain the voice information of the user's voice interaction with the smart headset, and perform preprocessing operations on the voice information, which can reduce the interference of environmental noise and enhance the clarity of the voice information. Input the preprocessed voice information into the semantic parsing model to obtain the semantic parsing result corresponding to the voice information. The semantic parsing model can accurately understand the intention of the voice information, enabling the smart headset to execute the instructions corresponding to the voice information according to the user's intention, and is applicable to multiple scenarios such as music playback, multi-turn conversations, and headset control, enriching the functions of the smart headset.

[0049] In one embodiment, after obtaining the semantic parsing result corresponding to the voice information, the method further includes:

[0050] Control the smart headset to execute the relevant operations according to the semantic parsing result.

[0051] Control the smart headset to execute relevant operations according to the semantic parsing result to respond to the instructions corresponding to the user's voice information. For example, for the semantic parsing result of the instruction "inquire about the weather in Shenzhen today", the smart headset will voice feedback the weather in Shenzhen to the user. Thus, it adapts to the needs of different users and provides personalized services.

[0052] In one embodiment, the semantic parsing model is deployed in the cloud, and the smart headset can input the preprocessed voice information into the semantic parsing model in the cloud. Among them, inputting the preprocessed voice information into the semantic parsing model to obtain the semantic parsing result corresponding to the voice information specifically includes:

[0053] Input the preprocessed voice information into the voice feature extraction layer to obtain the initial features corresponding to the voice information;

[0054] Use the multi-scale acoustic feature encoding layer to extract the acoustic features of multiple scales corresponding to the initial features;

[0055] After fusing the acoustic features of the multiple scales, input them into the context-aware semantic encoding layer to obtain the target features corresponding to the speech information;

[0056] Input the target features into the dialogue state management layer to obtain the state vector of the speech information;

[0057] Input the state vector into the intent recognition layer to obtain the semantic parsing result of the speech information.

[0058] The speech feature extraction layer can extract features such as MFCC features, sound pressure level, and fundamental frequency of the preprocessed speech information, and these features can be fused into a feature matrix as the initial features corresponding to the speech information.

[0059] The multi-scale acoustic feature encoding layer has convolution kernels of different sizes. By using convolution kernels of different sizes to perform convolution operations on the initial features, multiple scales of acoustic features can be generated. The size of the convolution kernel can be set to different time and frequency resolutions to adapt to the extraction of various acoustic features. The multi-scale acoustic feature encoding layer can capture different time and frequency features in the speech signal, thereby obtaining a rich representation of acoustic features.

[0060] Since context information plays an important role in accurately understanding the user's intention, the context-aware semantic encoding layer can also combine multiple scales of acoustic features with the context to obtain more semantic target features. Concatenate or weighted average the features extracted from the multi-scale acoustic feature encoding layer to fuse these multi-dimensional features, and process the fused features through a long short-term memory network to extract the dependencies between contexts, thereby generating target features containing context information.

[0061] The dialogue state management layer is used to integrate and maintain the state information of the conversation, including the context of the current dialogue, the user's historical instructions, and the context. Input the target features into the dialogue state management layer, and maintain the state record of the current dialogue through a state update mechanism (such as a state transition diagram). The generated state vector will reflect various information in the current dialogue, including key information such as the user's intention and emotional state. Pass the obtained state vector to the intent recognition layer, use multiple hidden layers of a fully connected neural network to solve the final intent classification, and the output result includes multiple intent categories and their confidence levels. Select the intent with the highest confidence as the semantic parsing result of the speech information.

[0062] Further, the step of inputting the preprocessed speech information into the speech feature extraction layer to obtain the initial features corresponding to the speech information includes:

[0063] Extract various acoustic features of the preprocessed speech information;

[0064] Calculate the first-order difference features and second-order difference features of each of the described acoustic features;

[0065] Concatenate the described multiple acoustic features, the first-order difference features, and the second-order difference features to obtain the initial features corresponding to the speech information.

[0066] It is possible to extract the Mel Frequency Cepstral Coefficients (MFCCs), Linear Prediction Cepstral Coefficients (LPCCs), spectrogram features, and speech energy of the preprocessed speech information. The first-order difference features are obtained by calculating the difference between the current frame features and the previous frame features, and the second-order difference features are calculated as the change in the first-order difference features, i.e., the difference between the current first-order difference feature and the previous first-order difference feature. Concatenating multiple acoustic features, first-order, and second-order difference features to obtain the initial features corresponding to the speech information enables the initial features to more comprehensively express the attributes of the speech information. For example, assuming the MFCC features have 12 dimensions, the first-order difference features and the second-order difference features each have 12 dimensions, then the final feature vector will be 36 dimensions (12 + 12 + 12).

[0067] Furthermore, using the multi-scale acoustic feature encoding layer, extract the acoustic features of multiple scales corresponding to the initial features, including:

[0068] Input the initial features into multiple convolutional layer networks for encoding to obtain the first feature, the second feature, and the third feature, where each feature is obtained by encoding the initial features with different convolutional kernels;

[0069] Use the dual-parameterized convolutional layer network to perform a convolutional operation on the first feature to obtain the first-scale intermediate feature;

[0070] Use the temporal pyramid pooling layer network to perform a pooling operation on the second feature to obtain the second-scale intermediate feature;

[0071] Use the self-attention mechanism to perform a decoding operation on the third feature to obtain the third-scale intermediate feature.

[0072] Each convolutional layer network uses convolutional kernels of different sizes (e.g., 3x3, 5x5, and 7x7) to extract features of different scales. Input the initial features into these convolutional layers, and after convolutional operations, obtain the first feature, the second feature, and the third feature respectively. The first feature can characterize the features of short-term frequency changes, the second feature can characterize the features of mid-term pitch changes, and the third feature can characterize the features of long-term timbre changes.

[0073] The dual-parameterized convolutional layer network can enhance the model's ability to express features by introducing an additional parameterization mechanism. The first feature is input into the dual-parameterized convolutional layer for convolution operations. The dual-parameterized convolutional layer uses two independent convolutional kernels to extract different features respectively, and generates the first-scale intermediate feature by combining the convolution results. The first-scale intermediate feature contains richer local information and reflects more complex acoustic features.

[0074] The second feature is input into the time pyramid pooling layer. Multiple pooling windows are set to perform pooling operations at different time scales respectively, and the outputs of each pooling window are concatenated to generate the second-scale intermediate feature. The second-scale intermediate feature retains the diversity of time information and enhances the understanding of speech signals.

[0075] The self-attention mechanism can dynamically adjust the weights of features according to the relationships between input features, thereby capturing long-range dependencies. When processing speech signals, the self-attention mechanism can effectively extract context information, enabling the model to understand more complex semantics. The third feature is input into the self-attention mechanism layer, the correlation between each feature is calculated and weights are assigned to each feature, and the third-scale intermediate feature is generated through weighted averaging. The third-scale intermediate feature combines context information and reflects a deeper level of semantic understanding.

[0076] Refer to Figure 2 As shown, it is a schematic diagram of the functional modules of the semantic parsing device 100 based on smart headphones according to the present application.

[0077] The semantic parsing device 100 based on smart headphones described in the present application can be installed in smart headphones. According to the implemented functions, the semantic parsing device 100 based on smart headphones can include a verification module 110, a permission module 120, a preprocessing module 130, and a parsing module 140. The modules described in the present application can also be referred to as units, which refer to a series of computer program segments that can be executed by a smart headphone processor and can complete fixed functions, and are stored in the memory of the smart headphone.

[0078] In this embodiment, the functions of each module / unit are as follows:

[0079] Verification module 110: used to obtain the wake-up speech issued by the user and verify the voiceprint information of the wake-up speech;

[0080] Permission module 120: used to, if the voiceprint information is verified to be passed, allow the user to perform voice interaction with the smart headphone;

[0081] Preprocessing module 130: used to obtain the speech information of the user's voice interaction with the smart headphone and perform preprocessing operations on the speech information;

[0082] Parsing module 140: configured to input the preprocessed speech information into a semantic parsing model to obtain a semantic parsing result corresponding to the speech information.

[0083] In one embodiment, the semantic parsing device 100 based on the intelligent earphone further includes a control module, and the control module is configured to:

[0084] Control the intelligent earphone to perform the relevant operations according to the semantic parsing result.

[0085] In one embodiment, the verification of the voiceprint information of the wake-up voice includes:

[0086] Extract the target voiceprint feature of the wake-up voice;

[0087] Match the target voiceprint feature with a pre-constructed voiceprint feature library;

[0088] If there is a voiceprint feature in the voiceprint feature library that matches the target voiceprint feature, the voiceprint information is verified to pass.

[0089] In one embodiment, the preprocessing operation performed on the speech information includes:

[0090] Perform a noise removal operation on the speech information to obtain denoised speech information;

[0091] Perform a speech enhancement operation on the denoised speech information to obtain preprocessed speech information.

[0092] In one embodiment, the inputting the preprocessed speech information into a semantic parsing model to obtain a semantic parsing result corresponding to the speech information includes:

[0093] Input the preprocessed speech information into a speech feature extraction layer to obtain initial features corresponding to the speech information;

[0094] Use a multi-scale acoustic feature encoding layer to extract acoustic features of multiple scales corresponding to the initial features;

[0095] After fusing the acoustic features of the multiple scales, input them into a context-aware semantic encoding layer to obtain target features corresponding to the speech information;

[0096] Input the target features into a dialogue state management layer to obtain a state vector of the speech information;

[0097] Input the state vector into an intent recognition layer to obtain a semantic parsing result of the speech information.

[0098] In one embodiment, inputting the preprocessed voice information into a voice feature extraction layer to obtain initial features corresponding to the voice information includes:

[0099] Extracting multiple acoustic features of the preprocessed voice information;

[0100] Calculating the first-order difference features and second-order difference features of each of the acoustic features;

[0101] Concatenating the multiple acoustic features, the first-order difference features, and the second-order difference features to obtain initial features corresponding to the voice information.

[0102] In one embodiment, using a multi-scale acoustic feature encoding layer to extract acoustic features of multiple scales corresponding to the initial features includes:

[0103] Inputting the initial features into multiple convolutional layer networks for encoding respectively to obtain a first feature, a second feature, and a third feature, where each feature is obtained by encoding the initial features with different convolutional kernels;

[0104] Performing a convolution operation on the first feature using a dual-parameterized convolutional layer network to obtain a first-scale intermediate feature;

[0105] Performing a pooling operation on the second feature using a temporal pyramid pooling layer network to obtain a second-scale intermediate feature;

[0106] Performing a decoding operation on the third feature using a self-attention mechanism to obtain a third-scale intermediate feature.

[0107] Refer to Figure 3 shown, which is a schematic diagram of a preferred embodiment of the intelligent earphone of the present application.

[0108] The intelligent earphone includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114;

[0109] The memory 113 is used to store computer programs, for example, a semantic parsing program based on the intelligent earphone; among them, the processor 111 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 111 is generally used to control the overall operation of the intelligent earphone, for example, to execute control and processing related to data interaction or communication, etc. In this embodiment, the processor 111 is used to run the program code stored in the memory 113 or process data, for example, to run the program code of the semantic parsing program based on the intelligent earphone, etc.

[0110] The communication interface 112 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and the communication interface 112 can also be used to establish a communication connection between the intelligent earphone and other intelligent earphones.

[0111] The memory 113 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 113 may be an internal storage unit of the intelligent earphone, such as the hard disk or memory of the intelligent earphone. In other embodiments, the memory 113 may also be an external storage device of the intelligent earphone, such as a plug-in hard disk equipped with the intelligent earphone, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory 113 may also include both the internal storage unit of the intelligent earphone and its external storage device. In this embodiment, the memory 11 is generally used to store the operating system installed in the intelligent earphone and various computer programs, such as the program code of the semantic parsing program based on the intelligent earphone. In addition, the memory 113 can also be used to temporarily store various data that have been output or will be output.

[0112] Figure 3 Only the intelligent earphone having components 111-114 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0113] In an embodiment of the present application, when the processor 111 is used to execute the program stored on the memory 113, it implements the semantic parsing method based on the intelligent earphone provided by any one of the foregoing method embodiments, including:

[0114] Obtain the wake-up voice issued by the user, and verify the voiceprint information of the wake-up voice;

[0115] If the voiceprint information is verified to be passed, allow the user to perform voice interaction with the intelligent earphone;

[0116] Obtain the voice information of the user's voice interaction with the intelligent earphone, and perform a preprocessing operation on the voice information;

[0117] Input the preprocessed voice information into the semantic parsing model to obtain the semantic parsing result corresponding to the voice information.

[0118] For a detailed introduction to the above steps, please refer to the above Figure 1 Description of the flowchart of the embodiment of the semantic parsing method based on intelligent headphones.

[0119] In addition, an embodiment of the present application also proposes a computer-readable storage medium, which can be non-volatile or volatile. The computer-readable storage medium includes a storage data area and a storage program area. The storage program area stores a semantic parsing program based on intelligent headphones. When the semantic parsing program based on intelligent headphones is executed by a processor, the following operations are implemented:

[0120] Obtain the wake-up voice issued by the user and verify the voiceprint information of the wake-up voice;

[0121] If the voiceprint information is verified, allow the user to perform voice interaction with the intelligent headphones;

[0122] Obtain the voice information of the user's voice interaction with the intelligent headphones and perform a preprocessing operation on the voice information;

[0123] Input the preprocessed voice information into the semantic parsing model to obtain the semantic parsing result corresponding to the voice information.

[0124] The specific implementation manner of the computer-readable storage medium of the present application is substantially the same as the specific implementation manner of the above-mentioned semantic parsing method based on intelligent headphones, and will not be elaborated here.

[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0127] It should be noted that the descriptions involving "first", "second", etc. in this application are only for descriptive purposes, and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0128] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that alternative or additional steps may be used.

[0129] The above are only specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A semantic parsing method based on smart headphones, characterized in that: The method is applied to a smart headset, and the method comprises: Acquire a wake-up voice issued by a user, and verify the voiceprint information of the wake-up voice; If the voiceprint information is verified, the user is allowed to perform voice interaction with the smart headset; Acquire voice information of voice interaction between the user and the smart headset, and perform preprocessing operations on the voice information; The preprocessed speech information is input into the semantic analysis model to obtain the semantic analysis result corresponding to the speech information.

2. The semantic parsing method based on smart earphones according to claim 1, characterized in that: After obtaining the semantic analysis result corresponding to the voice information, the method further includes: The smart headset is controlled to perform the related operations according to the semantic parsing result.

3. The semantic parsing method based on smart earphones according to claim 2, characterized in that: The verifying the voiceprint information of the wake-up voice includes: Extracting target voiceprint features of the wake-up speech; Matching the target voiceprint feature with a pre-built voiceprint feature library; If the voiceprint feature database contains a voiceprint feature that matches the target voiceprint feature, the voiceprint information verification is successful.

4. The semantic parsing method based on smart earphones according to claim 1, characterized in that: The performing of a preprocessing operation on the voice information comprises: Performing a noise removal operation on the voice information to obtain denoised voice information; A speech enhancement operation is performed on the denoised speech information to obtain preprocessed speech information.

5. The semantic parsing method based on smart earphones according to claim 1, characterized in that: The step of inputting the preprocessed speech information into a semantic parsing model to obtain a semantic parsing result corresponding to the speech information includes: Inputting the preprocessed speech information into the speech feature extraction layer to obtain initial features corresponding to the speech information; Using a multi-scale acoustic feature encoding layer, extracting acoustic features of multiple scales corresponding to the initial features; After fusing the acoustic features of the multiple scales, the features are input into a context-aware semantic coding layer to obtain target features corresponding to the speech information; Inputting the target feature into the dialogue state management layer to obtain the state vector of the speech information; The state vector is input into the intention recognition layer to obtain the semantic analysis result of the speech information.

6. The semantic parsing method based on smart earphones according to claim 5, characterized in that: The step of inputting the preprocessed speech information into the speech feature extraction layer to obtain initial features corresponding to the speech information includes: Extracting various acoustic features of preprocessed speech information; Calculating the first-order difference feature and the second-order difference feature of each of the acoustic characteristics; The multiple acoustic features, the first-order difference features and the second-order difference features are concatenated to obtain initial features corresponding to the speech information.

7. The semantic parsing method based on smart earphones according to claim 5, characterized in that: The method of using a multi-scale acoustic feature encoding layer to extract acoustic features of multiple scales corresponding to the initial features includes: Inputting the initial features into multiple convolutional layer networks for encoding respectively to obtain a first feature, a second feature and a third feature, wherein each feature is obtained by encoding the initial features by using a different convolution kernel; Using a dual parameterized convolutional layer network, a convolution operation is performed on the first feature to obtain a first-scale intermediate feature; Using a temporal pyramid pooling layer network, a pooling operation is performed on the second feature to obtain an intermediate feature of a second scale; The self-attention mechanism is used to decode the third feature to obtain an intermediate feature of the third scale.

8. A semantic parsing device based on smart headphones, characterized in that: The device comprises: Verification module: used to obtain the wake-up voice issued by the user and verify the voiceprint information of the wake-up voice; Permission module: used to allow the user to perform voice interaction with the smart headset if the voiceprint information verification is passed; Preprocessing module: used to obtain voice information of the user's voice interaction with the smart headset, and perform preprocessing operations on the voice information; Parsing module: used to input the preprocessed voice information into the semantic parsing model to obtain the semantic parsing result corresponding to the voice information.

9. A smart headset, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the semantic parsing method based on smart headphones as described in any one of claims 1 to 7 when executing the program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the semantic parsing method based on a smart headset as described in any one of claims 1 to 7 is implemented.