Japanese speech model training method, interaction method, storage medium, and device
Patent Information
- Application Number
- CN202211321530.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-10-26
AI Technical Summary
[0003]相关技术中,语音识别和语义理解是割裂的,对于日语这种存在多种表记文本方式的语种来说,语义理解模型很难对不同表记方式的文本抽取意图和槽位
[0021] An electronic device according to an embodiment of the present invention includes a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the aforementioned training method for the Japanese voice interaction model and the Japanese voice interaction method. Thus, through the training method for the Japanese voice interaction model and the Japanese voice interaction method, semantic information can be correctly extracted from text in various writing styles, improving the accuracy of semantic recognition and making the interactive information more in line with people's daily reading and writing habits.
Smart Images

Figure CN115662399B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech technology, specifically to a training method, interaction method, storage medium, and device for a Japanese speech model. Background Technology
[0002] As electronic products become increasingly intelligent, intelligent interactive systems are being used more and more, with intelligent voice interaction systems being a crucial aspect of this intelligence. With the increasing demand for exported products, multilingual voice interaction technology is poised to become one of the most important technologies for showcasing the intelligence of exported products.
[0003] In related technologies, speech recognition and semantic understanding are often separated. For languages like Japanese, which have multiple ways of writing text, semantic understanding models struggle to extract intent and slots from text written in different formats. Furthermore, from a purely speech recognition perspective, it's difficult to determine which format best suits people's daily reading and writing habits. Therefore, the performance of these technologies is often poor. Summary of the Invention
[0004] This invention aims to at least partially address one of the technical problems in related technologies. Therefore, the first objective of this invention is to propose a training method for a Japanese speech model that can correctly extract semantic information from texts written in various formats, improve semantic recognition accuracy, and make it more consistent with people's daily reading and writing habits.
[0005] The second objective of this invention is to propose a Japanese voice interaction method.
[0006] The third objective of this invention is to provide an electronic device.
[0007] The fourth objective of this invention is to provide a computer-readable storage medium.
[0008] To achieve the above objectives, a first aspect of the present invention proposes a method for training a Japanese speech model, comprising: acquiring a first training set and training an initial speech recognition model using multiple Japanese speech information from the first training set to obtain a target speech recognition model, wherein the speech recognition model is used to recognize the text corresponding to the Japanese speech information; acquiring a second training set and training an initial semantic recognition model using multiple sets of notation information from the second training set to obtain a target semantic recognition model, wherein the notation information includes a text phoneme sequence and a notation method composed of at least one of Chinese character notation and kana character notation, and the semantic recognition model is used to recognize the semantic meaning of characters or words in the text; and concatenating the target speech recognition model with the target semantic recognition model to obtain a Japanese speech interaction model.
[0009] According to an embodiment of the present invention, a method for training a Japanese speech model firstly acquires a first training set and trains an initial speech recognition model using multiple Japanese speech information from the first training set to obtain a target speech recognition model. The speech recognition model is used to recognize the text corresponding to the Japanese speech information. Next, a second training set is acquired, and an initial semantic recognition model is trained using multiple sets of notation information from the second training set to obtain a target semantic recognition model. The notation information includes a text phoneme sequence and a notation method composed of at least one of kanji (Chinese characters) or kana (Japanese characters). The semantic recognition model is used to recognize the semantic meaning of characters or words in the text. Finally, the target speech recognition model and the target semantic recognition model are concatenated to obtain a Japanese speech interaction model. Therefore, this method for training a Japanese speech model can correctly extract semantic information from texts with various notation methods, improve the accuracy of semantic recognition, and make the interactive information more consistent with people's daily reading and writing habits.
[0010] Furthermore, the training method for the Japanese speech model proposed in the above embodiments of the present invention may also have the following additional technical features:
[0011] In one embodiment of the present invention, the step of training an initial speech recognition model using multiple Japanese speech information from the first training set to obtain a target speech recognition model includes: annotating the multiple Japanese speech information with transcription information to obtain multiple sets of transcription information; and training the initial speech recognition model based on the transcription information of the multiple Japanese speech information to obtain the target speech recognition model.
[0012] In one embodiment of the present invention, training the initial speech recognition model based on the transcription information of multiple Japanese speech information to obtain the target speech recognition model includes: for each Japanese speech information, performing speech recognition on the Japanese speech information to obtain a speech state sequence, and obtaining a set of predicted transcription information based on the speech state sequence; and performing supervised training on the initial speech recognition model using the predicted transcription information based on the transcription information of multiple Japanese speech information to obtain the target speech recognition model.
[0013] In one embodiment of the present invention, the step of training an initial semantic recognition model using multiple sets of transcription information in the second training set to obtain a target semantic recognition model includes: annotating the multiple sets of transcription information with speech information to obtain multiple semantic information; and training the initial semantic recognition model based on the semantic information of the multiple sets of transcription information to obtain the target semantic recognition model.
[0014] In one embodiment of the present invention, training the initial semantic recognition model based on the semantic information of multiple sets of the notation information to obtain the target semantic recognition model includes: for each set of notation information, extracting features from the notation information to obtain Chinese character feature vectors, kana feature vectors, and phoneme feature vectors; fusing the Chinese character feature vectors and the kana feature vectors to obtain text feature vectors, and concatenating the text feature vectors and the phoneme feature vectors to obtain concatenated features; decoding the concatenated features to obtain predicted semantic information; and training the initial semantic recognition model using the predicted semantic information based on the semantic information of multiple sets of notation information to obtain the target semantic recognition model.
[0015] In one embodiment of the present invention, the acquisition method of the first training set and the second training set includes: acquiring Japanese speech information, sending transcription information and semantic information based on the Japanese speech information; receiving feedback information on the transcription information and the semantic information, and determining whether the transcription information and the semantic information meet expectations based on the feedback information; if so, saving the Japanese speech information and the corresponding transcription information to form the first training set, and saving the transcription information and the semantic information to form the second training set.
[0016] To achieve the above objectives, a second aspect of the present invention proposes a Japanese voice interaction method, comprising: acquiring interactive Japanese voice; recognizing the interactive Japanese voice using a Japanese voice interaction model to obtain corresponding semantic information, wherein the Japanese voice interaction model is trained according to the above-described Japanese voice interaction model training method; and obtaining interactive information based on the semantic information.
[0017] According to an embodiment of the Japanese voice interaction method of the present invention, firstly, interactive Japanese voice is acquired; then, the interactive Japanese voice is recognized using a Japanese voice interaction model to obtain corresponding semantic information, wherein the Japanese voice interaction model is trained according to the aforementioned method; and finally, interactive information is obtained based on the semantic information. Therefore, this Japanese voice interaction method can correctly extract semantic information from texts written in various formats, improve the accuracy of semantic recognition, and make the interactive information more consistent with people's daily reading and writing habits.
[0018] In addition, the Japanese voice interaction method proposed in the above embodiments of the present invention may also have the following additional technical features:
[0019] In one embodiment of the present invention, the semantic information includes intent and slot information, wherein obtaining interaction information based on the semantic information includes: performing semantic format conversion on the intent and the slot information to obtain the interaction information.
[0020] To achieve the above objectives, a third aspect of the present invention provides an electronic device comprising: a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the above-described training method for the Japanese voice interaction model and the above-described Japanese voice interaction method.
[0021] An electronic device according to an embodiment of the present invention includes a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the aforementioned training method for the Japanese voice interaction model and the Japanese voice interaction method. Thus, through the training method for the Japanese voice interaction model and the Japanese voice interaction method, semantic information can be correctly extracted from text in various writing styles, improving the accuracy of semantic recognition and making the interactive information more in line with people's daily reading and writing habits.
[0022] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium storing a program file that can be executed to implement the above-described training method for the Japanese speech interaction model and the Japanese speech interaction method.
[0023] According to an embodiment of the present invention, when a program file thereon is executed, it implements the above-described training method for the Japanese speech interaction model and the Japanese speech interaction method. Therefore, through the training method for the Japanese speech interaction model and the Japanese speech interaction method, semantic information can be correctly extracted from texts written in various formats, improving the accuracy of semantic recognition and making the interactive information more consistent with people's daily reading and writing habits.
[0024] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0025] Figure 1 This is a flowchart of a training method for a Japanese speech model according to an embodiment of the present invention;
[0026] Figure 2 This is a schematic diagram of a speech state sequence of Japanese speech information as an example of the present invention;
[0027] Figure 3 This is a schematic diagram of the representation information of Japanese voice information as an example of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of a semantic recognition model as an example of the present invention;
[0029] Figure 5 This is a flowchart of a Japanese voice interaction method according to an embodiment of the present invention;
[0030] Figure 6 is a schematic flow diagram of an example Japanese voice interaction method according to the present invention. Detailed Description of Embodiments
[0031] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals throughout denote the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary, and are intended to explain the present invention, and should not be construed as a limitation on the present invention.
[0032] In a voice interaction model, an automatic speech recognition model and a semantic recognition model are generally included, but the training of the speech recognition model and the semantic recognition model are performed separately. For Japanese, which has a plurality of text notation forms, on one hand, if the speech recognition model provides a single text to the semantic recognition model, a large amount of information will be lost. For example, the sentence "open bluetooth" has multiple notations in Japanese, such as "bluetoothを開く", "bluetoothをひらく", "ブルートゥースを開く", "ブルートゥースをひらく", etc. If the speech recognition model recognizes that the text notation is "bluetoothを開く", the input information provided to the semantic recognition model loses the features of other notations of each text; on the other hand, if the semantic recognition model is also trained according to the notation of "bluetoothを開く", it can extract intent and slot information from the text "bluetoothを開く", but if the text is in the notation of "ブルートゥースをひらく", it is difficult for the semantic recognition model to extract the corresponding intent and slot information. Therefore, from the perspective of speech recognition, it is difficult to determine which notation is more in line with people's daily reading and writing habits.
[0033] Accordingly, the present invention provides a training method for a Japanese speech model, a Japanese speech interaction method, a storage medium, and an apparatus.
[0034] The training method for a Japanese speech model, the interaction method, the storage medium and the apparatus according to embodiments of the present invention are described below with reference to the accompanying drawings.
[0035] Figure 1 is a flow chart of a training method for a Japanese speech model according to an embodiment of the present invention.
[0036] As Figure 1 shown, the training method for a Japanese speech model according to an embodiment of the present invention includes the following steps:
[0037] S101, Obtain the first training set, and use multiple Japanese speech information in the first training set to train the initial speech recognition model to obtain the target speech recognition model, wherein the speech recognition model is used to recognize the text corresponding to the Japanese speech information.
[0038] In some embodiments of the present invention, the initial speech recognition model is trained using multiple Japanese speech information from a first training set to obtain a target speech recognition model, which may include:
[0039] S201, mark multiple Japanese phonetic information with annotation information to obtain multiple sets of annotation information.
[0040] Specifically, the aforementioned transcription information includes a text phoneme sequence and a transcription method consisting of at least one of Chinese character text transcription and kana text transcription, which can be described in detail with reference to step S202.
[0041] S202, the initial speech recognition model is trained based on the transcription information of multiple Japanese speech information to obtain the target speech recognition model.
[0042] Specifically, it may include:
[0043] A1. For each Japanese speech information, perform speech recognition on the Japanese speech information to obtain a speech state sequence, and obtain a set of predicted representation information based on the speech state sequence.
[0044] A2. Based on the transcription information of multiple Japanese speech information, the initial speech recognition model is trained in a supervised manner using the predicted transcription information to obtain the target speech recognition model.
[0045] Specifically, the speech recognition model consists of two parts: an acoustic model and a language model. The acoustic model outputs a sequence of speech states for each Japanese speech item, while the language model encodes the speech state sequence output by the acoustic model to obtain a sequence of text phonemes. This text phoneme sequence is then decoded to obtain a set of predicted notation information. (Refer to...) Figure 2 The process from frame to speech state sequence is a sequence numbering operation for text phonemes. Parts of the speech state sequence are encoded to obtain the corresponding text phoneme sequence "s IH ks". Then, the text phoneme sequence is passed through the decoder to obtain the corresponding recognized text "six".
[0046] However, in this embodiment of the invention, for each Japanese speech information, speech recognition is performed to obtain a speech state sequence. The speech state sequence is first encoded to obtain a text phoneme sequence, and then the text phoneme sequence is decoded. The decoder here not only outputs a single text sequence, but also the corresponding text phoneme sequence and a notation method composed of at least one of kanji text notation and kana text notation; that is, the predicted notation information for each Japanese speech information. For example, inputting "turn on Bluetooth" into the speech recognition model for training yields a set of predicted notation information, which can be referred to... Figure 3 .
[0047] S102, Obtain the second training set, and use multiple sets of notation information in the second training set to train the initial semantic recognition model to obtain the target semantic recognition model, wherein the semantic recognition model is used to recognize the semantic meaning of characters or words in the text.
[0048] In some embodiments of the present invention, the initial semantic recognition model is trained using multiple sets of representation information in the second training set to obtain the target semantic recognition model, which may include:
[0049] S301, annotate multiple sets of recorded information with speech information to obtain multiple semantic information.
[0050] S302, the initial semantic recognition model is trained based on the semantic information of multiple sets of table information to obtain the target semantic recognition model.
[0051] Specifically, it may include:
[0052] B1. For each set of information, perform feature extraction on the set of information to obtain Chinese character feature vectors, kana feature vectors, and phoneme feature vectors.
[0053] B2. Perform feature fusion on the Chinese character feature vector and the kana feature vector to obtain the text feature vector, and then concatenate the text feature vector and the phoneme feature vector to obtain the concatenated feature.
[0054] B3. Decode the spliced features to obtain the predicted semantic information.
[0055] B4. Based on the semantic information of multiple sets of table information, the initial semantic recognition model is trained using the predicted semantic information to obtain the target semantic recognition model.
[0056] Specifically, the semantic recognition model includes an intent model and a NER model. The intent model extracts the intent of the text corresponding to Japanese speech information, while the NER model extracts the entity slots of the text corresponding to Japanese speech information. For example, in the phrase "open bluetooth," the intent is "open," and the entity slot is "bluetooth." The structure of the NER model can be found in [reference needed]. Figure 4 First, feature vectors from various notation methods—namely, Chinese character feature vectors and kana feature vectors—are fused to obtain text feature vectors. Then, these text feature vectors and phoneme feature vectors are concatenated to obtain concatenated features. These concatenated features are then input into a multi-layer transformer and hidden layers for decoding, ultimately yielding the entity slot information "bluetooth" in the text. Similarly, the intent model is not elaborated upon here. Finally, after decoding the concatenated features, predicted semantic information is obtained, which includes the text's intent and entity slot information.
[0057] In some embodiments of the present invention, the methods for obtaining the first training set and the second training set include:
[0058] S401: Acquire Japanese voice information and send transcription and semantic information based on the Japanese voice information.
[0059] S402, receive feedback information regarding the representation information and semantic information, and determine whether the representation information and semantic information meet expectations based on the feedback information.
[0060] S403, if so, then save the Japanese phonetic information and the corresponding notation information to form the first training set, and save the notation information and semantic information to form the second training set.
[0061] S103, the target speech recognition model and the target semantic recognition model are concatenated to obtain the Japanese speech interaction model.
[0062] In this embodiment of the invention, by jointly training and optimizing the speech recognition model and the semantic recognition model, the text features of the kanji and kana text representations in Japanese are fused and concatenated with the text phoneme features. This allows for effective extraction of speech information during semantic recognition, regardless of whether the text is represented by kanji, kana, or a combination of both. Furthermore, based on the results of semantic recognition, semantic format conversion can be performed to optimize the presentation of interactive information, thereby improving the overall interactive effect and intelligence of the Japanese speech interaction model.
[0063] In summary, the training method for the Japanese speech interaction model in this embodiment of the invention first obtains a first training set and trains an initial speech recognition model using multiple Japanese speech information from the first training set to obtain a target speech recognition model. This speech recognition model is used to identify the text corresponding to the Japanese speech information. Next, a second training set is obtained, and an initial semantic recognition model is trained using multiple sets of notation information from the second training set to obtain a target semantic recognition model. This notation information includes a text phoneme sequence and a notation method composed of at least one of kanji (Chinese characters) or kana (Japanese characters). The semantic recognition model is used to identify the semantic meaning of characters or words in the text. Finally, the target speech recognition model and the target semantic recognition model are concatenated to obtain the Japanese speech interaction model. Therefore, this training method for the Japanese speech interaction model, by concatenating the target speech recognition model and the target semantic recognition model, integrates text features from multiple notation methods, thereby correctly extracting semantic information, improving the accuracy of semantic recognition, and making the interactive information more consistent with people's daily reading and writing habits.
[0064] Furthermore, this invention proposes a Japanese voice interaction method.
[0065] Figure 5 This is a flowchart of a Japanese voice interaction method according to an embodiment of the present invention.
[0066] like Figure 5 As shown, the Japanese voice interaction method of this invention includes the following steps:
[0067] S501, obtain interactive Japanese voice.
[0068] S502, using a Japanese speech interaction model to recognize interactive Japanese speech and obtain corresponding semantic information, wherein the Japanese speech interaction model is trained according to the above-mentioned Japanese speech interaction model training method.
[0069] S503, obtain interactive information based on semantic information.
[0070] In some embodiments of the present invention, semantic information includes intent and slot information, wherein obtaining interaction information based on semantic information may include: performing semantic format conversion on intent and slot information to obtain interaction information.
[0071] Specifically, after recognizing the interactive Japanese speech using a Japanese speech interaction model, the corresponding intent and slot information are obtained, and then semantic format conversion is performed on it, converting it into a pre-specified semantic protocol format. This semantic protocol format can be set according to the user's daily reading and writing habits in different interaction scenarios.
[0072] To better understand the Japanese voice interaction method of this invention, please refer to... Figure 6To elaborate on it.
[0073] like Figure 6 As shown, to obtain interactive Japanese speech, firstly, speech recognition is performed on the interactive Japanese speech information. First, the acoustic model in the speech recognition model outputs the speech state sequence of the interactive Japanese speech information. Then, the language model in the speech recognition model encodes the speech state sequence output by the acoustic model to obtain a text phoneme sequence. The text phoneme sequence is then decoded to obtain the representation information of the interactive Japanese speech information, namely the text phoneme sequence and the representation method composed of at least one of the kanji text representation and the kana text representation.
[0074] Furthermore, the transcription information of the interactive Japanese speech information is input into the semantic recognition model. First, the transcription information is used to extract features to obtain kanji feature vectors, kana feature vectors and phoneme feature vectors. Then, the kanji feature vectors and kana feature vectors are fused to obtain text feature vectors. Finally, the text feature vectors and phoneme feature vectors are concatenated to obtain concatenated features.
[0075] Furthermore, the splicing features are decoded to obtain the semantic information of the interactive Japanese speech information. Finally, the semantic format is converted to obtain the interactive information, thereby meeting the user's daily reading and writing habits.
[0076] In summary, the Japanese voice interaction method of this invention first acquires interactive Japanese voice, then uses a Japanese voice interaction model to recognize the interactive Japanese voice and obtain corresponding semantic information. The Japanese voice interaction model, based on the training method described above, obtains interactive information from the semantic information. Therefore, this Japanese voice interaction method, by concatenating the target speech recognition model and the target semantic recognition model, integrates text features from multiple writing methods, thereby correctly extracting semantic information, improving the accuracy of semantic recognition, and making the interactive information more in line with people's daily reading and writing habits.
[0077] Furthermore, the present invention proposes an electronic device.
[0078] In this embodiment of the invention, the electronic device includes a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute a training method for a Japanese voice interaction model and a Japanese voice interaction method.
[0079] When the program instructions on the electronic device of this invention are invoked by the processor, the aforementioned training method for the Japanese speech interaction model and the Japanese speech interaction method are executed. Thus, the training method for the Japanese speech interaction model and the Japanese speech interaction method, by utilizing the concatenation of the target speech recognition model and the target semantic recognition model, fuse text features from multiple writing methods, thereby correctly extracting semantic information, improving the accuracy of semantic recognition, and making the interactive information more in line with people's daily reading and writing habits.
[0080] Furthermore, the present invention proposes a computer-readable storage medium.
[0081] In this embodiment of the invention, a computer-readable storage medium stores a program file that can be executed to implement a training method for a Japanese voice interaction model and a Japanese voice interaction method.
[0082] The computer-readable storage medium of this invention, when the program instructions thereon are executed, implements the above-described training method for the Japanese speech interaction model and the Japanese speech interaction method. Thus, the training method for the Japanese speech interaction model and the Japanese speech interaction method, by utilizing the concatenation of a target speech recognition model and a target semantic recognition model, fuse text features from multiple writing methods, thereby correctly extracting semantic information, improving the accuracy of semantic recognition, and making the interactive information more in line with people's daily reading and writing habits.
[0083] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0084] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0085] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0086] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0087] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0088] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0089] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
[0090] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A training method for a Japanese voice interaction model, characterized in that, include: A first training set is obtained, and an initial speech recognition model is trained using multiple Japanese speech information in the first training set to obtain a target speech recognition model. The speech recognition model is used to identify the transcription information corresponding to the Japanese speech information. The transcription information includes a text phoneme sequence and a transcription method composed of at least one of Chinese character text transcription and kana text transcription. A second training set is obtained, and the initial semantic recognition model is trained using multiple sets of transcription information in the second training set to obtain the target semantic recognition model. The transcription information includes a text phoneme sequence and a transcription method composed of at least one of Chinese character text transcription and kana text transcription. The semantic recognition model is used to identify the semantic meaning of the characters or words in the text. The step of training the initial semantic recognition model using multiple sets of representation information from the second training set to obtain the target semantic recognition model includes: Semantic information annotation is performed on multiple sets of the aforementioned representation information to obtain multiple semantic information; The initial semantic recognition model is trained based on the semantic information of multiple sets of the aforementioned representation information to obtain the target semantic recognition model; The target speech recognition model and the target semantic recognition model are concatenated to obtain a Japanese speech interaction model.
2. The training method for the Japanese voice interaction model according to claim 1, characterized in that, The step of training the initial speech recognition model using multiple Japanese speech information from the first training set to obtain the target speech recognition model includes: Multiple sets of transcription information are obtained by annotating the various Japanese phonetic information; The initial speech recognition model is trained based on the transcription information of multiple Japanese speech information to obtain the target speech recognition model.
3. The training method for the Japanese voice interaction model according to claim 2, characterized in that, The process of training the initial speech recognition model based on the transcription information of multiple Japanese speech information to obtain the target speech recognition model includes: For each Japanese speech information, speech recognition is performed on the Japanese speech information to obtain a speech state sequence, and a set of predicted notation information is obtained based on the speech state sequence; Based on the transcription information of multiple Japanese speech information, the initial speech recognition model is trained in a supervised manner using the predicted transcription information to obtain the target speech recognition model.
4. The training method for the Japanese voice interaction model according to claim 1, characterized in that, The process of training the initial semantic recognition model based on semantic information from multiple sets of the aforementioned representation information to obtain the target semantic recognition model includes: For each set of the aforementioned information, feature extraction is performed on the set of information to obtain Chinese character feature vectors, kana feature vectors, and phoneme feature vectors; The feature vectors of the Chinese characters and the feature vectors of the kana are fused to obtain the text feature vector, and the text feature vector and the phoneme feature vector are concatenated to obtain the concatenated feature. Decoding the spliced features yields predicted semantic information; Based on the semantic information of multiple sets of the aforementioned representation information, the initial semantic recognition model is trained using the predicted semantic information to obtain the target semantic recognition model.
5. The training method for the Japanese voice interaction model according to claim 1, characterized in that, The methods for obtaining the first training set and the second training set include: Acquire Japanese voice information, and send transcription information and semantic information based on the Japanese voice information; Receive feedback information regarding the representation information and the semantic information, and determine whether the representation information and the semantic information meet expectations based on the feedback information; If so, the Japanese phonetic information and the corresponding notation information are saved to form the first training set, and the notation information and the semantic information are saved to form the second training set.
6. A Japanese voice interaction method, characterized in that, include: Get interactive Japanese voice recordings; The interactive Japanese speech is identified using a Japanese speech interaction model to obtain corresponding semantic information, wherein the Japanese speech interaction model is the training method of the Japanese speech interaction model according to any one of claims 1-5; Interaction information is obtained based on the semantic information.
7. The Japanese voice interaction method according to claim 6, characterized in that, The semantic information includes intent and slot information, wherein obtaining interaction information based on the semantic information includes: The semantic format of the intent and the slot information is converted to obtain the interaction information.
8. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the training method of the Japanese voice interaction model as claimed in any one of claims 1-5 and the Japanese voice interaction method as claimed in claim 6 or 7.
9. A computer-readable storage medium, characterized in that, The system stores a program file that can be executed to implement the training method for the Japanese voice interaction model as described in any one of claims 1-5 and the Japanese voice interaction method as described in claim 6 or 7.
Citation Information
Patent Citations
Terminal device, speech recognition method and speech recognition program
JP2012063526A
Processing method, device and electronic apparatus
US20200211533A1