Semantic information determination method, storage medium and electronic device

By slicing and encoding audio data and matching semantic information using an audio vector library, the latency and accuracy issues in speech recognition technology are solved, resulting in more efficient speech recognition and a better user experience.

CN121053973APending Publication Date: 2025-12-02HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD +3
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202410682604.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing speech recognition technologies suffer from significant latency and low accuracy, resulting in a poor user experience.

Method used

By slicing and encoding real-time received audio data based on time windows, and using an audio vector library for matching, semantic information can be directly identified, skipping the ASR and NLU processing flow.

Benefits of technology

It improves speech recognition efficiency, reduces overall processing latency, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053973A_ABST
    Figure CN121053973A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic information determination method, a storage medium and an electronic device, and relates to the technical field of smart home, the semantic information determination method comprises the following steps: slicing first audio data received in real time based on a first time window to obtain second audio data; performing voice coding on the second audio data to obtain a first audio vector of the second audio data; whether a second audio vector exists in a target audio vector library or not is determined, the second audio vector is a voice vector meeting a preset first matching condition with the first audio vector, multiple tuple information is stored in the target audio vector library, and the tuple information comprises audio vectors and semantic information corresponding to the audio vectors; and when it is determined that the second audio vector exists in the target audio vector library, determining first semantic information corresponding to the second audio vector, and identifying the first semantic information as semantic information corresponding to the first audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home technology, and more specifically, to a method for determining semantic information, a storage medium, and an electronic device. Background Technology

[0002] Currently, in the home appliance control field, third-party Automatic Speech Recognition (ASR) services are mainly used to recognize user speech. The recognized results are then used for subsequent processing, such as Natural Language Understanding (NLU) and Dialogue Management (DM). However, the sequential processing logic results in a long overall time consumption, severely impacting the user experience.

[0003] In existing technologies, the common process of voice interaction systems is as follows:

[0004] Audio data → Automatic speech recognition → Natural language understanding → Natural language generation (NLG). Among these, ASR and NLU are processed serially, resulting in excessively long latency for the entire link, which is the sum of the latency of each module.

[0005] Regarding the issues of significant latency and low accuracy in speech recognition technology, no effective solutions have yet been proposed. Summary of the Invention

[0006] This application provides a method for determining semantic information, a storage medium, and an electronic device to at least solve the problems of large time delay and low efficiency in related technologies, such as speech recognition technology.

[0007] According to one embodiment of this application, a method for determining semantic information is provided, comprising: slicing first audio data received in real time based on a first time window to obtain second audio data; performing speech encoding on the second audio data to obtain a first audio vector; determining whether a second audio vector exists in a target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, the target audio vector library storing multiple tuple information, the tuple information including: an audio vector and semantic information corresponding to the audio vector; and, if it is determined that the second audio vector exists in the target audio vector library, determining the first semantic information corresponding to the second audio vector, and identifying the first semantic information as the semantic information corresponding to the first audio data.

[0008] In an exemplary embodiment, after determining whether a second audio vector exists in the target audio vector library, the method further includes: if the second audio vector does not exist in the target audio vector library, slicing the third audio data received in real time based on a second time window to obtain fourth audio data, wherein the third audio data includes at least the first audio data, and the fourth audio data includes at least the second audio data; performing speech encoding on the fourth audio data to obtain a third audio vector of the fourth audio data; determining whether a fourth audio vector exists in the target audio vector library, wherein the fourth audio vector is a speech vector that satisfies a preset second matching condition with the third audio vector; if the fourth audio vector exists, determining the second semantic information corresponding to the fourth audio vector, and identifying the second semantic information as the semantic information corresponding to the third audio data.

[0009] In an exemplary embodiment, before slicing the third audio data received in real time based on the second time window to obtain the fourth audio data, the method further includes: determining the audio duration of the second audio data and determining the first text data of the second audio data; determining the fifth audio vector with the highest similarity to the second audio vector in the audio vector library and determining the second text data corresponding to the fifth audio vector; determining the integrity of the second audio data based on the first text data and the second text data; and determining the second time window based on the integrity and the audio duration.

[0010] In an exemplary embodiment, before slicing the received first audio data based on a preset time window to obtain the second audio data, the method further includes: acquiring sample audio data and sample text data corresponding to the sample audio data; parsing the sample text data to obtain domain information, intent information, and slot information of the sample text data; determining the semantic information of the sample audio data based on the domain information, the intent information, and the slot information; establishing first tuple information based on the audio vector of the sample audio data and the semantic information of the sample audio data, and storing the first tuple information in the target audio vector library.

[0011] In one exemplary embodiment, parsing the sample text data to obtain domain information, intent information, and slot information of the sample text data includes: performing word segmentation on the sample text data to obtain word groups corresponding to the sample text data; determining the word vectors of the word groups and determining the meaning of the word groups based on the word vectors; and determining the domain information, slot information, and intent information of the sample text data based on the meaning of the word groups.

[0012] In one exemplary embodiment, determining whether a second audio vector exists in a target audio vector library includes: determining the part-of-speech vector and syntactic vector of the second audio data; determining the domain information of the second audio data based on the part-of-speech vector and the syntactic vector; and determining a target audio vector library among multiple audio vector libraries based on the domain information of the second audio data, wherein the domain information corresponding to the target audio vector library is consistent with the domain information of the second audio data.

[0013] In an exemplary embodiment, after determining whether a second audio vector exists in the target audio vector library, the method further includes: if the second audio vector does not exist in the target audio vector library, determining the audio duration of the second audio vector; if the audio duration is greater than a preset audio duration, inputting the second audio data into a speech recognition model so that the speech recognition model outputs first text data corresponding to the second audio data; and determining semantic information corresponding to the first audio data based on the first text data.

[0014] In one exemplary embodiment, slicing real-time received first audio data based on a first time window to obtain second audio data includes: loading the first audio data into an audio processing library; determining the start and end time points of the second audio data based on the first time window; and controlling the audio processing library to slice the first audio data according to the start and end time points to obtain the second audio data.

[0015] According to another embodiment of this application, a semantic information determination device is also provided, comprising: a slicing module, configured to slice real-time received first audio data based on a first time window to obtain second audio data; an encoding module, configured to perform speech encoding on the second audio data to obtain a first audio vector; a first determination module, configured to determine whether a second audio vector exists in a target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, the target audio vector library storing multiple tuple information, the tuple information including: an audio vector and semantic information corresponding to the audio vector; and a second determination module, configured to, when it is determined that the second audio vector exists in the target audio vector library, determine the first semantic information corresponding to the second audio vector, and identify the first semantic information as the semantic information corresponding to the first audio data.

[0016] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the method for determining the semantic information described above when running.

[0017] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the semantic information determination method described above through the computer program.

[0018] In this embodiment, the first audio data received in real time is sliced ​​based on a first time window to obtain second audio data; the second audio data is speech encoded to obtain a first audio vector; it is determined whether a second audio vector exists in a target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, and the target audio vector library stores multiple tuple information, the tuple information including: audio vector and semantic information corresponding to the audio vector; if it is determined that the second audio vector exists in the target audio vector library, the first semantic information corresponding to the second audio vector is determined, and the first semantic information is identified as the semantic information corresponding to the first audio data. That is, this embodiment slices the audio data received in real time based on a preset time window, determines the semantic information of the audio data based on the encoded audio vector in the audio vector library; judges the semantic information of the audio information based on streaming, and the corresponding semantic information is stored in the cached audio vector library. The high-confidence hit result can directly eliminate the processing flow of ASR and NLU, solving the problems of large latency and low efficiency in related technologies, thereby improving the speech recognition efficiency, reducing the overall processing latency, and improving the user experience. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the hardware environment for a method of determining semantic information according to an embodiment of this application;

[0022] Figure 2 This is a flowchart of a method for determining semantic information according to an embodiment of this application;

[0023] Figure 3 This is a schematic diagram of speech recognition of audio data in existing technology;

[0024] Figure 4 This is a schematic diagram illustrating the construction of an audio vector library according to an optional embodiment of this application;

[0025] Figure 5 This is a flowchart of a method for determining semantic information according to an optional embodiment of this application;

[0026] Figure 6 This is a structural block diagram of a semantic information determination device according to an embodiment of this application. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] According to one aspect of the embodiments of this application, a method for determining semantic information is provided. This method for determining semantic information is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned method for determining semantic information can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0030] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0031] This embodiment provides a method for determining semantic information, applied to the aforementioned terminal device. Figure 2 This is a flowchart of a method for determining semantic information according to an embodiment of this application, the process including the following steps:

[0032] Step S202: Slice the first audio data received in real time based on the first time window to obtain the second audio data;

[0033] Optionally, slicing the first audio data received in real time based on a first time window to obtain second audio data includes: loading the first audio data into an audio processing library; determining the start time point and end time point of the second audio data based on the first time window; and controlling the audio processing library to slice the first audio data according to the start time point and the end time point to obtain the second audio data.

[0034] In this embodiment, the first audio data is loaded into memory, for example, by loading an audio file using an audio processing library such as librosa or pydub; the start and end time points of the audio data to be sliced ​​(i.e. the start and end time points of the second audio data) are determined according to the first time window; the audio processing library slices the first audio data using the start and end time points of the slice to obtain the second audio data; and the second audio data is saved as a new file.

[0035] Step S204: Perform speech encoding on the second audio data to obtain the first audio vector;

[0036] Alternatively, the second audio data can be speech encoded in the following manner:

[0037] 1) The second audio data is encoded for speech based on the Mel frequency cepstral coefficient encoding method.

[0038] 2) The second audio data is speech encoded based on a linear prediction model.

[0039] 3) The second audio data is encoded for speech based on the short-time Fourier transform method.

[0040] 4) The second audio data is speech encoded using deep learning models such as convolutional neural networks (CNN) and recurrent neural networks (RNN).

[0041] Optionally, before encoding the second audio data, it needs to be preprocessed, including noise reduction and normalization. Then, according to the selected encoding method, the audio signal is converted into an audio vector.

[0042] Step S206: Determine whether a second audio vector exists in the target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, and the target audio vector library stores multiple tuple information, the tuple information including: audio vector and semantic information corresponding to the audio vector;

[0043] Optionally, the tuple information may also include: audio vector, semantic information corresponding to the audio vector, and text information corresponding to the audio vector.

[0044] It should be noted that the second audio vector mentioned above can also be understood as an audio vector whose similarity to the first audio vector is greater than a preset threshold similarity.

[0045] Optionally, the similarity between the first audio vector and each audio vector in the target audio vector library is determined as follows:

[0046] 1) If the lengths of the two audio vectors are equal, calculate the Euclidean distance or Manhattan distance between the first audio vector and each audio vector in the target audio vector library.

[0047] 2) Calculate the distance between two audio vectors based on the dynamic time warping algorithm, find a path with minimum cost, and determine the similarity between the first audio vector and each audio vector in the target audio vector library based on the path with minimum cost.

[0048] 3) Calculate the dot product of the first audio vector and each audio vector in the target audio vector library, and normalize it to a cosine value between -1 and 1; determine the similarity between the first audio vector and each audio vector in the target audio vector library based on the cosine value.

[0049] 4) Use machine learning algorithms or deep learning models to determine the similarity between the first audio vector and each audio vector in the target audio vector library to learn the representation of the audio data, and use the learned features for similarity calculation.

[0050] Given the similarity between the first audio vector and each audio vector in the target audio vector library, the similarity between the first audio vector and each audio vector in the target audio vector library determines whether the first audio vector and each audio vector in the target audio vector library satisfy a preset first matching condition.

[0051] Specifically: if the similarity is greater than the first similarity, determine that the first audio vector and each audio vector in the target audio vector library meet the preset first matching condition; if the similarity is less than or equal to the first similarity, determine that the first audio vector and each audio vector in the target audio vector library do not meet the preset first matching condition.

[0052] Step S208: If it is determined that the second audio vector exists in the target audio vector library, the first semantic information corresponding to the second audio vector is determined, and the first semantic information is identified as the semantic information corresponding to the first audio data.

[0053] Through the above steps, the first audio data received in real time is sliced ​​based on a first time window to obtain second audio data; the second audio data is speech encoded to obtain a first audio vector; it is determined whether a second audio vector exists in a target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, and the target audio vector library stores multiple tuple information, the tuple information including: audio vector and semantic information corresponding to the audio vector; if it is determined that the second audio vector exists in the target audio vector library, the first semantic information corresponding to the second audio vector is determined, and the first semantic information is identified as the semantic information corresponding to the first audio data. This solves the problems of large latency and low accuracy in related technologies, thereby improving speech recognition efficiency, reducing overall processing latency, and enhancing user experience.

[0054] Optionally, after determining whether a second audio vector exists in the target audio vector library, the following steps also need to be performed:

[0055] Step S11: If it is determined that the second audio vector does not exist in the target audio vector library, the third audio data received in real time based on the second time window is sliced ​​to obtain the fourth audio data, wherein the third audio data includes at least the first audio data, and the fourth audio data includes at least the second audio data.

[0056] Step S12: Perform speech encoding on the fourth audio data to obtain the third audio vector of the fourth audio data;

[0057] Step S13: Determine whether a fourth audio vector exists in the target audio vector library, wherein the fourth audio vector is a speech vector that satisfies a preset second matching condition with the third audio vector;

[0058] It should be noted that the aforementioned fourth audio vector can also be understood as an audio vector whose similarity to the third audio vector is greater than that to the second audio vector.

[0059] Step S14: If the fourth audio vector exists, determine the second semantic information corresponding to the fourth audio vector, and identify the second semantic information as the semantic information corresponding to the third audio data.

[0060] In other words, this embodiment employs a streaming detection method. If the semantic information of the first audio data is not identified, it may be because the received first audio data is incomplete. Therefore, it is necessary to continue receiving audio data. The received third audio data is then sliced ​​again to obtain fourth audio data. It is then determined whether a fourth audio vector with a similarity greater than the second similarity can be found in the target audio vector library. If a fourth audio vector exists, the second semantic information corresponding to the fourth audio vector is used as the semantic information corresponding to the third audio data. In other words, this embodiment determines the semantic information of the third audio data based solely on the fourth audio data, solving the problems of large latency and low accuracy in related technologies, thereby improving speech recognition efficiency, reducing overall processing latency, and enhancing user experience.

[0061] It should be noted that, before the fourth audio vector is determined, the fifth audio data received in real time based on the third time window is sliced ​​to obtain the sixth audio data. The fifth audio data includes at least the first audio data and the third audio data, and the sixth audio data includes at least the second audio data and the fourth audio data. The sixth audio data is then speech encoded to obtain the sixth audio vector. It is determined whether a seventh audio vector exists in the target audio vector library, wherein the seventh audio vector is a speech vector that satisfies a preset third matching condition with the sixth audio vector. If the seventh audio vector exists, the third semantic information corresponding to the seventh audio vector is determined, and the third semantic information is identified as the semantic information corresponding to the sixth audio data.

[0062] Optionally, the values ​​of the first similarity and the second similarity can be the same or different; if the values ​​of the first similarity and the second similarity are different, the first text data of the second audio data is determined; the fifth audio vector with the highest similarity to the second audio vector is determined in the audio vector library, and the second text data corresponding to the fifth audio vector is determined; the integrity of the second audio data is determined based on the first text data and the second text data; the first similarity is determined based on the integrity.

[0063] For example, if the completeness of the first audio data is 40%, the similarity can be set to 70%; if the completeness of the first audio data is 80%, the similarity can be set to 95%.

[0064] Optionally, the duration of the second time window can be the same as or different from that of the first time window. If the durations of the second time window and the first time window are different, the second time window can be determined based on the following scheme: determining the audio duration of the second audio data and determining the first text data of the second audio data; determining the fifth audio vector with the highest similarity to the second audio vector in the audio vector library and determining the second text data corresponding to the fifth audio vector; determining the integrity of the second audio data based on the first text data and the second text data; and determining the second time window based on the integrity and the audio duration.

[0065] In this embodiment of the application, in order to obtain complete audio data more quickly and accurately, the time when the target object outputs complete audio data is predicted based on the target object's speech rate and the completeness of the currently received audio data, and then a second time window is determined based on the time when the target object outputs complete audio data.

[0066] For example, if the duration of the first audio data is 2 seconds and the completeness of the first audio data is 40%, then the second time window should be 3 seconds.

[0067] It should be noted that slicing the third audio data received in real time based on the second time window includes: loading the third audio data into the audio processing library; determining the start and end time points of the third audio data based on the first and second time windows; and controlling the audio processing library to slice the third audio data according to the start and end time points to obtain the fourth audio data.

[0068] Optionally, if it is determined that the second audio vector does not exist in the target audio vector library, the following procedure also needs to be performed: determine the audio duration of the second audio vector; if the audio duration is greater than a preset audio duration, input the second audio data into the speech recognition model so that the speech recognition model outputs the first text data corresponding to the second audio data; determine the semantic information corresponding to the first audio data based on the first text data.

[0069] It should be noted that when the duration of the second audio vector is greater than the preset audio duration, it can be understood that the target object has output sufficient audio data. However, since the number of audio vectors in the audio vector library is limited, it is impossible to recognize the audio data output by the target object based on the audio vector library. Therefore, in order to ensure that the speech recognition effect can be achieved, the first text data corresponding to the second audio data is output based on the speech recognition model; and then the semantic information corresponding to the first audio data is determined based on the first text data.

[0070] Optionally, before performing step S202, it is also necessary to establish a target speech vector library. The method for establishing the target speech vector library is as follows: obtain sample audio data and sample text data corresponding to the sample audio data; parse the sample text data to obtain the domain information, intent information, and slot information of the sample text data; determine the semantic information of the sample audio data based on the domain information, the intent information, and the slot information; establish first tuple information based on the audio vector of the sample audio data and the semantic information of the sample audio data, and store the first tuple information in the target audio vector library.

[0071] In this embodiment, the speech information corresponding to the sample audio data is determined based on the domain information, intent information and slot information of the sample text data, and tuple information is established according to the sample audio data and the speech information corresponding to the sample audio data, and the tuple information is saved to the target audio vector library.

[0072] Specifically, parsing the sample text data to obtain the domain information, intent information, and slot information of the sample text data includes: performing word segmentation on the sample text data to obtain the word groups corresponding to the sample text data; determining the word vectors of the word groups and determining the meaning of the word groups based on the word vectors; and determining the domain information, slot information, and intent information of the sample text data based on the meaning of the word groups.

[0073] Optionally, domain information, slot information, and intent information can also be determined through the following methods: Preprocessing the sample audio data, including but not limited to: noise reduction, segmentation, and silence detection; converting the preprocessed audio data into sample text data; inputting the sample text data into an intent recognition system, wherein the intent recognition system includes, but is not limited to: a pre-trained machine learning model, such as a Support Vector Machine (SVM), Recurrent Neural Network (RNN), or Transformer model; after recognizing the user's intent, extracting entities or key information related to the intent from the sample text data; wherein specific information in the entity or key information text, such as time, location, or name; mapping the extracted entities or key information to predefined slots. For example, if the user's intent is to book a hotel, then the slots include, but are not limited to, "date," "hotel name," and "room type"; for example, if the user's intent is to control the air conditioner, then the slots include, but are not limited to, "temperature" and "fan speed."

[0074] Optionally, determining whether a second audio vector exists in the target audio vector library includes: determining the part-of-speech vector and syntactic vector of the second audio data; determining the domain information of the second audio data based on the part-of-speech vector and the syntactic vector; and determining the target audio vector library from multiple audio vector libraries based on the domain information of the second audio data, wherein the domain information corresponding to the target audio vector library is consistent with the domain information of the second audio data.

[0075] In other words, due to the sheer number of audio vectors, storing all types of audio vectors in a single audio vector library would lead to excessive resource consumption and low recognition accuracy. Therefore, it is necessary to store different types of audio vectors in different audio vector libraries. Then, upon receiving audio data, the target audio vector library can be determined based on the domain information of the audio data, thereby achieving the effect of quickly recognizing the semantic information corresponding to the audio vectors.

[0076] To better understand the process of determining the semantic information described above, the implementation flow of the method for determining the semantic information will be further described below in conjunction with optional embodiments, but this is not intended to limit the technical solution of the embodiments of this application.

[0077] In existing technologies, besides the latency issue, when performing pipelined processing based on semi-streaming NLP techniques, ASR inevitably encounters semantically incomplete audio truncation, leading to unreliable intermediate ASR results. For example, as shown in the attached... Figure 3 As shown, when the user speaks slowly, Voice Activity Detection (VAD) will incorrectly determine that the black line is the end of the voice command, and then the ASR service will return an incomplete command recognition result.

[0078] In other words, the existing technology still suffers from the problem of low accuracy in audio data recognition.

[0079] To address the aforementioned issues of latency and low accuracy, this embodiment provides a method for determining semantic information. Firstly, an audio vector library needs to be established. Figure 4 This is a schematic diagram illustrating the construction of an audio vector library according to an optional embodiment of this application, such as... Figure 4 As shown, the specific steps are as follows:

[0080] Step 1: Prepare sample data;

[0081] The sample data includes audio data and text corresponding to manually annotated audio data.

[0082] Step 2: Semantic information parsing;

[0083] In the NLU module, the BERT model is used to parse text data to obtain the corresponding domain information, slot information, and intent information. Then, based on the domain information, slot information, and intent information, the semantic information corresponding to the audio data is determined.

[0084] Step 3: Audio data encoding;

[0085] (1) Train the speech encoder based on (audio, text) data so that the audio vectors corresponding to audio data with the same semantic information have high similarity, and vice versa.

[0086] (2) The audio data is encoded according to the trained speech encoder to obtain the audio vector.

[0087] Step 4: Audio vector library storage;

[0088] The audio vectors are combined with the corresponding text and semantic information to form tuple information, which is then stored in the audio vector library.

[0089] After establishing the audio vector library, speech recognition is performed based on the aforementioned audio vector library. Figure 5 This is a flowchart of a method for determining semantic information according to an optional embodiment of this application, such as... Figure 5 As shown, the specific steps are as follows:

[0090] Step 1: Stream the audio data;

[0091] Specifically: the audio data is sliced ​​according to fixed time windows (each slice starts at the same time);

[0092] Step 2: Encode the audio data;

[0093] That is, the streaming slices are encoded sequentially using a speech encoder to obtain the corresponding audio vectors.

[0094] Step 3: Perform a search based on the audio vector library;

[0095] Specifically: The audio vector is compared with all data in the audio vector library using a cosine similarity algorithm to obtain the data with the highest similarity; the semantic integrity and validity of the current audio data are judged based on the confidence level (i.e., similarity).

[0096] Step 4: Obtain semantic information;

[0097] When the confidence level obtained by streaming slice retrieval is higher than the preset threshold (0.95), the semantic information corresponding to the retrieved data is used as the NLU result and returned directly.

[0098] If the value is below a preset threshold, the corresponding semantic information will be discarded directly, and subsequent streaming audio processing will continue.

[0099] In this embodiment, streaming audio data can be sliced ​​according to a fixed time window, encoded to obtain audio vectors, and then the audio vector library is searched to determine the semantic completeness and semantic validity of the user's voice commands. Based on the above method, the user's valid voice commands can be identified as quickly as possible, and when the retrieval confidence is higher than a preset threshold, the ASR and NLU modules can be skipped to directly obtain the semantic information of the command, greatly reducing the system's response latency and improving the user experience.

[0100] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0101] This application also provides a semantic information determination device in its embodiments. Figure 6 This is a structural block diagram of a semantic information determining device according to an embodiment of this application; as shown below. Figure 6 As shown, it includes:

[0102] The slicing module 62 is used to slice the first audio data received in real time based on a first time window to obtain the second audio data;

[0103] Encoding module 64 is used to perform speech encoding on the second audio data to obtain a first audio vector;

[0104] The first determining module 66 is used to determine whether a second audio vector exists in the target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, and the target audio vector library stores multiple tuple information, the tuple information including: audio vector and semantic information corresponding to the audio vector;

[0105] The second determining module 68 is used to determine the first semantic information corresponding to the second audio vector when it is determined that the second audio vector exists in the target audio vector library, and to identify the first semantic information as the semantic information corresponding to the first audio data.

[0106] Using the aforementioned device, the first audio data received in real time is sliced ​​based on a first time window to obtain second audio data; the second audio data is then encoded into speech to obtain a first audio vector; it is determined whether a second audio vector exists in a target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, and the target audio vector library stores multiple tuple information, the tuple information including: an audio vector and semantic information corresponding to the audio vector; if it is determined that the second audio vector exists in the target audio vector library, the first semantic information corresponding to the second audio vector is determined, and the first semantic information is identified as the semantic information corresponding to the first audio data. This solves the problems of large latency and low accuracy in related technologies, thereby improving speech recognition efficiency, reducing overall processing latency, and enhancing user experience.

[0107] In an exemplary embodiment, the slicing module 62 is configured to slice the third audio data received in real time based on the second time window to obtain the fourth audio data when it is determined that the second audio vector does not exist in the target audio vector library. The third audio data includes at least the first audio data, and the fourth audio data includes at least the second audio data.

[0108] Encoding module 64 is used to perform speech encoding on the fourth audio data to obtain the third audio vector of the fourth audio data;

[0109] The first determining module 66 is used to determine whether a fourth audio vector exists in the target audio vector library, wherein the fourth audio vector is a speech vector that satisfies a preset second matching condition with the third audio vector;

[0110] The second determining module 68 is used to determine the second semantic information corresponding to the fourth audio vector when the fourth audio vector exists, and to identify the second semantic information as the semantic information corresponding to the third audio data.

[0111] In one exemplary embodiment, the apparatus further includes: a third determining module, configured to determine the audio duration of the second audio data and determine first text data of the second audio data; determine a fifth audio vector in the audio vector library that has the highest similarity to the second audio vector, and determine second text data corresponding to the fifth audio vector; determine the integrity of the second audio data based on the first text data and the second text data; and determine a second time window based on the integrity and the audio duration.

[0112] In one exemplary embodiment, the above apparatus further includes: a building module, configured to acquire sample audio data and sample text data corresponding to the sample audio data; parse the sample text data to obtain domain information, intent information, and slot information of the sample text data; determine semantic information of the sample audio data based on the domain information, the intent information, and the slot information; build first tuple information based on the audio vector of the sample audio data and the semantic information of the sample audio data, and store the first tuple information in the target audio vector library.

[0113] In an exemplary embodiment, a module is established to perform word segmentation on the sample text data to obtain word groups corresponding to the sample text data; determine the word vectors of the word groups and determine the meaning of the word groups based on the word vectors; and determine the domain information, slot information, and intent information of the sample text data based on the meaning of the word groups.

[0114] In an exemplary embodiment, a first determining module is configured to determine part-of-speech vectors and syntactic vectors of the second audio data; determine domain information of the second audio data based on the part-of-speech vectors and the syntactic vectors; and determine a target audio vector library among multiple audio vector libraries based on the domain information of the second audio data, wherein the domain information corresponding to the target audio vector library is consistent with the domain information of the second audio data.

[0115] In an exemplary embodiment, the above apparatus further includes: a recognition module, configured to determine the audio duration of the second audio vector when it is determined that the second audio vector does not exist in the target audio vector library; input the second audio data into a speech recognition model when the audio duration is greater than a preset audio duration, so that the speech recognition model outputs first text data corresponding to the second audio data; and determine semantic information corresponding to the first audio data based on the first text data.

[0116] In an exemplary embodiment, the slicing module is further configured to load the first audio data into an audio processing library; determine the start time point and end time point of the second audio data based on the first time window; and control the audio processing library to slice the first audio data according to the start time point and the end time point to obtain the second audio data.

[0117] Embodiments of this application also provide a storage medium including a stored program, wherein the program executes any of the methods described above when it is run.

[0118] Optionally, in this embodiment, the storage medium may be configured to store program code for performing the following steps:

[0119] S1, the first audio data received in real time is sliced ​​based on the first time window to obtain the second audio data;

[0120] S2, perform speech encoding on the second audio data to obtain the first audio vector;

[0121] S3, determine whether a second audio vector exists in the target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, and the target audio vector library stores multiple tuple information, the tuple information including: audio vector and semantic information corresponding to the audio vector;

[0122] S4, if it is determined that the second audio vector exists in the target audio vector library, determine the first semantic information corresponding to the second audio vector, and identify the first semantic information as the semantic information corresponding to the first audio data.

[0123] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0124] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0125] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0126] S1, the first audio data received in real time is sliced ​​based on the first time window to obtain the second audio data;

[0127] S2, perform speech encoding on the second audio data to obtain the first audio vector;

[0128] S3, determine whether a second audio vector exists in the target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, and the target audio vector library stores multiple tuple information, the tuple information including: audio vector and semantic information corresponding to the audio vector;

[0129] S4, if it is determined that the second audio vector exists in the target audio vector library, determine the first semantic information corresponding to the second audio vector, and identify the first semantic information as the semantic information corresponding to the first audio data.

[0130] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0131] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0132] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0133] The embodiments described herein also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.

[0134] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0135] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0136] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for determining semantic information, characterized in that, include: The first audio data received in real time is sliced ​​based on the first time window to obtain the second audio data; The second audio data is speech encoded to obtain the first audio vector; Determine whether a second audio vector exists in the target audio vector library, wherein the second audio vector is a speech vector that satisfies a preset first matching condition with the first audio vector, and the target audio vector library stores multiple tuple information, the tuple information including: audio vector and semantic information corresponding to the audio vector; If it is determined that the second audio vector exists in the target audio vector library, the first semantic information corresponding to the second audio vector is determined, and the first semantic information is identified as the semantic information corresponding to the first audio data.

2. The method for determining semantic information according to claim 1, characterized in that, After determining whether a second audio vector exists in the target audio vector library, the method further includes: If it is determined that the second audio vector does not exist in the target audio vector library, the third audio data received in real time based on the second time window is sliced ​​to obtain the fourth audio data, wherein the third audio data includes at least the first audio data, and the fourth audio data includes at least the second audio data; The fourth audio data is speech encoded to obtain the third audio vector; Determine whether a fourth audio vector exists in the target audio vector library, wherein the fourth audio vector is a speech vector that satisfies a preset second matching condition with the third audio vector; In the presence of the fourth audio vector, the second semantic information corresponding to the fourth audio vector is determined, and the second semantic information is identified as the semantic information corresponding to the third audio data.

3. The method for determining semantic information according to claim 2, characterized in that, Before slicing the third audio data received in real time based on the second time window to obtain the fourth audio data, the method further includes: Determine the audio duration of the second audio data, and determine the first text data of the second audio data; In the audio vector library, determine the fifth audio vector that has the highest similarity to the second audio vector, and determine the second text data corresponding to the fifth audio vector; The integrity of the second audio data is determined based on the first text data and the second text data; The second time window is determined based on the integrity and the audio duration.

4. The method for determining semantic information according to claim 1, characterized in that, After determining whether a second audio vector exists in the target audio vector library, the method further includes: If it is determined that the second audio vector does not exist in the target audio vector library, the audio duration of the second audio vector is determined; If the audio duration is longer than the preset audio duration, the second audio data is input into the speech recognition model so that the speech recognition model outputs the first text data corresponding to the second audio data; The semantic information corresponding to the first audio data is determined based on the first text data.

5. The method for determining semantic information according to claim 1, characterized in that, Before slicing the received first audio data based on a preset time window to obtain the second audio data, the method further includes: acquiring sample audio data and sample text data corresponding to the sample audio data; The sample text data is parsed to obtain the domain information, intent information, and slot information of the sample text data; The semantic information of the sample audio data is determined based on the domain information, the intent information, and the slot information; A first tuple is established based on the audio vector and semantic information of the sample audio data, and the first tuple is stored in the target audio vector library.

6. The method for determining semantic information according to claim 5, characterized in that, The sample text data is parsed to obtain the domain information, intent information, and slot information of the sample text data, including: The sample text data is segmented into words to obtain the word groups corresponding to the sample text data; Determine the word vector of the phrase, and determine the meaning of the phrase based on the word vector; Based on the meaning of the phrases, the domain information, slot information, and intent information of the sample text data are determined.

7. The method for determining semantic information according to claim 1, characterized in that, Determining whether a second audio vector exists in the target audio vector library includes: Determine the part-of-speech vector and syntactic vector of the second audio data; The domain information of the second audio data is determined based on the part-of-speech vector and the syntactic vector; A target audio vector library is determined from multiple audio vector libraries based on the domain information of the second audio data, wherein the domain information corresponding to the target audio vector library is consistent with the domain information of the second audio data.

8. The method for determining semantic information according to claim 1, characterized in that, The first audio data received in real time is sliced ​​based on the first time window to obtain the second audio data, including: Load the first audio data into the audio processing library; The start and end times of the second audio data are determined based on the first time window; the audio processing library is controlled to slice the first audio data according to the start and end times to obtain the second audio data.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method described in any one of claims 1 to 8.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 8 through the computer program.

Citation Information

Patent Citations

  • Voice information identifying method and device

    CN106205610A

  • Voice recognition method based on key words

    CN109545190A

  • Audio signal processing method and device, multimedia information processing method and device and electronic equipment

    CN111866542A

  • Voice retrieval method and device, equipment and storage medium

    CN112445934A

  • Voice-based disease early warning method and device, equipment and storage medium

    CN113935330A