A speech adaptation method and system based on spoken tone markers
By recording and adjusting the user's spoken tone, the dialect adaptation problem in voice interaction technology was solved, achieving a voice interaction experience consistent with the user's spoken language habits, expanding the product's applicability, especially for voice interaction products for the elderly.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2026-03-10
AI Technical Summary
Existing voice interaction technologies struggle to support the diverse dialects and habits of elderly people in different regions, resulting in low voice interaction recognition efficiency, especially since niche dialects are difficult to fully adapt to.
By recording the user's spoken tone and using preset tone markers to adjust the interactive voice output, a voice interaction experience consistent with the user's spoken language habits is formed. This includes recognizing user voice information, judging differences in tone sequences, adjusting tone sequences to match the user's accent, and providing feedback and optimizing the output content.
This greatly expands the application scope of voice interaction products, especially suitable for voice interaction products for the elderly, and improves voice interaction recognition efficiency and user experience.
Smart Images

Figure CN115641838B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of product development for the elderly and voice data processing technology, and in particular to a voice self-adapting method and system based on oral tone marking. BACKGROUND
[0002] With the continuous progress of intelligent technology, voice interaction technology is gradually applied to various life scenes, and the elderly are increasingly using voice interaction technology, including smart speakers, smart voice customer service, smart voice diagnosis, voice interaction training games, companion robots, etc. On the one hand, intelligent language interaction technology can assist the elderly in their daily lives, replacing input methods such as key presses and typing, and also helping the elderly maintain mental activity and improve cognitive and communication skills through strong logical voice interaction with the elderly. On the other hand, due to the prevalence of dialects, slow speech, low volume, and unclear pronunciation among the elderly, especially the different dialects in different regions of China, which have become a technical problem for the development of voice interaction technology.
[0003] One of the major technical problems that existing voice interaction technology must face is how to support different dialect habits of the elderly in various regions and improve the efficiency of voice interaction recognition. Although existing technologies have adopted the method of pre-setting multiple different major local dialects for users to choose from, which has to some extent expanded the application population and range of products, it still requires the elderly to adjust their language habits to adapt to the product, and it is difficult to completely develop and adapt to niche dialects. SUMMARY
[0004] To solve the problems of the prior art, the present application proposes a voice self-adapting method and system based on oral tone marking, which records and freely marks the user's oral tone on the basis of existing major dialect selection, and adjusts the interactive voice output accordingly, ultimately forming a voice interaction experience consistent with the user's oral habits, which can greatly expand the application range of voice interaction products, especially for the development and use of voice interaction products for the elderly.
[0005] To achieve the above purpose, the technical solution adopted by the present application includes:
[0006] A voice self-adapting method based on oral tone marking, characterized in that it comprises:
[0007] S1, according to the basic dialect type selection, load the preset tone marking;
[0008] S2, obtain the user's voice, identify the user's voice information and the corresponding first tone sequence;
[0009] S3, process the user's voice information using the preset tone marking to obtain the second tone sequence;
[0010] S4, judging whether the difference degree between the first pitch sequence and the second pitch sequence is greater than a preset threshold value, and when judging that the difference degree between the first pitch sequence and the second pitch sequence is not greater than the preset threshold value, marking the preset pitch mark as an output pitch mark;
[0011] S5, when judging that the difference degree between the first pitch sequence and the second pitch sequence is greater than the preset threshold value, adjusting the second pitch sequence to the first pitch sequence using a preset adjustment coefficient until the difference degree between the third pitch sequence obtained after adjustment and the first pitch sequence is not greater than the preset threshold value;
[0012] S6, modifying the preset pitch mark based on the third pitch sequence to obtain an output pitch mark;
[0013] S7, adjusting the voice output content using the output pitch mark.
[0014] Further, the method further comprises:
[0015] S8, feeding back the voice output content to the user, obtaining the voice output content repeated by the user, and identifying a corresponding fourth pitch sequence;
[0016] S9, replacing the first pitch sequence with the fourth pitch sequence, and repeating steps S4 to S7.
[0017] Further, the difference degree calculation method comprises:
[0018] Subtracting the second pitch sequence from the first pitch sequence to obtain a difference value;
[0019] Dividing the difference value by the first pitch sequence to obtain the difference degree.
[0020] Further, the preset adjustment coefficient is greater than 0 and less than or equal to 1.
[0021] Further, the adjusting the second pitch sequence to the first pitch sequence using the preset adjustment coefficient comprises:
[0022] Multiplying the preset adjustment coefficient by the difference degree to obtain an adjustment target coefficient;
[0023] Adjusting the second pitch sequence using the adjustment target coefficient.
[0024] The application also relates to a voice self-adapting system based on spoken language pitch marks, which is characterized by comprising:
[0025] A voice recognition module for recognizing user voice information and a corresponding first pitch sequence;
[0026] A pitch mark processing module for processing the user voice information using a preset pitch mark to obtain a second pitch sequence;
[0027] a tone sequence judgment module, configured to judge whether a difference between the first tone sequence and the second tone sequence is greater than a preset threshold value;
[0028] a tone sequence adjustment module, configured to adjust the second tone sequence to the first tone sequence using a preset adjustment coefficient to obtain a third tone sequence;
[0029] a tone mark management module, configured to modify the preset tone mark based on the third tone sequence to obtain an output tone mark;
[0030] a voice output module, configured to adjust voice output content using the output tone mark.
[0031] The application also relates to a computer readable storage medium, characterized in that the storage medium stores a computer program, and the computer program is executed by a processor to implement the above method.
[0032] The application also relates to an electronic device, characterized in that the electronic device comprises a processor and a memory.
[0033] The memory is configured to store a preset tone mark, a preset threshold value and a preset adjustment coefficient.
[0034] The processor is configured to execute the above method by calling the preset tone mark, the preset threshold value and the preset adjustment coefficient.
[0035] The application also relates to a computer program product, comprising a computer program and / or instructions, characterized in that the computer program and / or instructions are executed by a processor to implement the steps of the above method.
[0036] The application has the following beneficial effects:
[0037] The speech self-adaptive method and system based on spoken tone marking can greatly expand the application range of speech interaction products, and are especially suitable for the development and use of speech interaction products for the elderly. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 FIG. 1 is a flowchart of the speech self-adaptive method based on spoken tone marking.
[0039] Figure 2 FIG. 2 is a structural diagram of the speech self-adaptive system based on spoken tone marking. DETAILED DESCRIPTION
[0040] In order to more clearly understand the content of the present application, detailed description will be made in combination with the drawings and examples.
[0041] The first aspect of the present application relates to a speech adaptation method based on oral tone marking as shown in the following step flow: Figure 1 The speech adaptation method comprises the following steps:
[0042] S1, loading preset tone marking according to the selected basic dialect type.
[0043] The basic dialect type can be preferably a widely used dialect type other than the standard Chinese, such as Cantonese, Fujianese, Northeastern Chinese, etc. The user can select the basic dialect type close to the user's oral habits as the oral tone marking basis according to the user's oral habits.
[0044] The preset tone marking covers the regular special oral tone corresponding to the basic dialect type, such as initial stress, final stress, medial tone, and special tone marking for special words, etc.
[0045] S2, obtaining user speech, identifying user speech information and corresponding first tone sequence.
[0046] The user speech information is preferably processed by using a suitable semantic recognition method. For special dialects, there may be inaccurate user speech information recognition, and a user confirmation information step can be added as appropriate to ensure that the subsequent tone marking can accurately match the user speech content.
[0047] The corresponding first tone sequence is preferably a set of annotations directly matched with the user speech information by simulating the input of the user speech, so as to accurately and completely reflect the user's oral tone habits.
[0048] S3, processing the user speech information using the preset tone marking to obtain a second tone sequence.
[0049] On the basis of correctly obtaining the user speech information, the user speech information is completely independently processed by using the preset tone marking to form the second tone sequence expected to be output under the current setting. The generation of the second tone sequence makes the user speech content be processed only according to the system preset value without referring to the user's oral tone habits.
[0050] S4, judging whether the difference degree between the first tone sequence and the second tone sequence is greater than a preset threshold value, and when the difference degree between the first tone sequence and the second tone sequence is not greater than the preset threshold value, registering the preset tone marking as the output tone marking.
[0051] Specifically, the calculation method of the difference degree comprises: subtracting the second tone sequence from the first tone sequence to obtain a difference value; and dividing the difference value by the first tone sequence to obtain the difference degree.
[0052] The difference degree represents the difference between the actual accent and tone of the user and the tone generated by the system preset. The difference degree can be positive or negative, and the positive and negative of the difference degree only corresponds to the artificial setting of the direction of the tone difference, and there is no other additional meaning.
[0053] S5, when it is judged that the difference degree between the first tone sequence and the second tone sequence is greater than the preset threshold, the second tone sequence is adjusted to the first tone sequence using a preset adjustment coefficient until the difference degree between the third tone sequence obtained after adjustment and the first tone sequence is not greater than the preset threshold. Preferably, the preset adjustment coefficient is a positive decimal greater than 0 and less than or equal to 1, which is used to control the adjustment amplitude of the preset tone to the user tone.
[0054] Correspondingly, adjusting the second tone sequence to the first tone sequence using the preset adjustment coefficient comprises: multiplying the preset adjustment coefficient by the difference degree to obtain an adjustment target coefficient; and adjusting the second tone sequence using the adjustment target coefficient.
[0055] That is, in the adjustment process, it is not always pursued to adjust the system output to be completely consistent with the user's accent habit, but to deviate from the user's habit as much as possible within a reasonable range based on the system preset. Due to individual differences of users, using the system preset accent tone can achieve more extensive compatibility, for example, facilitating others to understand the content of the user's voice conversation, without misunderstanding caused by extremely special accent.
[0056] S6, modifying the preset tone mark based on the third tone sequence to obtain an output tone mark.
[0057] In actual application, the generation of the output tone mark is gradual, and the output tone mark conforming to the user's communication habit can be gradually improved and formed through the increase of the number of user usage conversations and the accumulation of related accent tone samples.
[0058] S7, adjusting the voice output content using the output tone mark.
[0059] S8, feeding back the voice output content to the user, obtaining the voice output content repeated by the user, and identifying a corresponding fourth tone sequence.
[0060] For some special dialects, there can be a large difference between the user's accent and the system preset value, or the local change of the user's accent in the repeating state, so the actual accent habit of the user can be determined in the form of repeated content to avoid excessive revision of the accent tone.
[0061] S9, replacing the first tone sequence with the fourth tone sequence, and repeating steps S4 to S7.
[0062] Another aspect of the present application also relates to a voice adaptive system based on spoken language tone mark, the structure of which is as followsFigure 2 As shown, comprising:
[0063] A voice recognition module is configured to recognize user voice information and a corresponding first tone sequence.
[0064] A tone mark processing module is configured to process the user voice information using a preset tone mark to obtain a second tone sequence.
[0065] A tone sequence judgment module is configured to judge whether a difference between the first tone sequence and the second tone sequence is greater than a preset threshold.
[0066] A tone sequence adjustment module is configured to adjust the second tone sequence to the first tone sequence using a preset adjustment coefficient to obtain a third tone sequence.
[0067] A tone mark management module is configured to modify the preset tone mark based on the third tone sequence to obtain an output tone mark.
[0068] A voice output module is configured to adjust voice output content using the output tone mark.
[0069] By using the system, the above-mentioned operation processing method can be executed and the corresponding technical effects can be achieved.
[0070] The embodiments of the present application also provide a computer readable storage medium capable of realizing all steps of the method in the above-mentioned embodiments, and the computer readable storage medium stores a computer program, which is executed by a processor to realize all steps of the method in the above-mentioned embodiments.
[0071] The embodiments of the present application also provide an electronic device for executing the above-mentioned method, which is an implementation device of the method, and the electronic device at least has a processor and a memory, and in particular, the memory stores data and related computer programs required for executing the method, such as a preset tone mark, a preset threshold and a preset adjustment coefficient, and all steps of the method are realized by calling the data and programs in the memory by the processor, and the corresponding technical effects are obtained.
[0072] Preferably, the electronic device can include a bus architecture, the bus can include any number of interconnected buses and bridges, the bus will include various circuits linked together by one or more processors and memories. The bus can also link various other circuits such as peripheral devices, voltage regulators and power management circuits, which are well known in the art, and therefore, will not be further described herein. The bus interface provides an interface between the bus and the receiver and transmitter. The receiver and transmitter can be the same element, i.e., a transceiver, which provides a unit for communicating with various other systems over a transmission medium. The processor is responsible for managing the bus and general processing, while the memory can be used to store data used by the processor in performing operations.
[0073] Additionally, the electronic device can further include a communication module, an input unit, an audio processor, a display, a power supply, etc. The processor (or controller, operating control) employed therein can include a microprocessor or other processor device and / or logic device, which receives input and controls the operation of various components of the electronic device; the memory can be one or more of a cache, a flash memory, a hard drive, a removable media, a volatile memory, a non-volatile memory, or other suitable device, which stores the above-mentioned related data information, and in addition, can store programs for executing related information, and the processor can execute the programs stored in the memory to achieve information storage or processing, etc.; the input unit is used to provide input to the processor, which can be a key or a touch input device, for example; the power supply is used to provide power to the electronic device; the display is used to display display objects such as images and text, which can be an LCD display, for example. The communication module is a transmitter / receiver that transmits and receives signals via an antenna. The communication module (transmitter / receiver) is coupled to the processor to provide input signals and receive output signals, which can be the same as in the case of a conventional mobile communication terminal. Based on different communication technologies, multiple communication modules can be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module (transmitter / receiver) is also coupled to the speaker and the microphone via the audio processor to provide audio output via the speaker and receive audio input from the microphone, thereby achieving the usual telecommunication functions. The audio processor can include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor is also coupled to the central processor, so that it can be able to record on the local machine through the microphone, and it can be able to play the sound stored on the local machine through the speaker.
[0074] Those skilled in the art will appreciate that embodiments of the present application can be readily used as software, hardware, or a combination of software and hardware. In a software embodiment, the methods can be tangibly embodied in a machine-readable storage medium having stored thereon instructions that can be used to program a computer to perform any of the methods. The software implementation can be initialized by loading and executing a set of instructions arranged to perform one of the methods into the computer's memory. Alternatively, hard-wired circuitry can be used in place of, or in combination with, software instructions. Thus, the
[0075] The present application is described in reference to the drawings using a flowchart and / or a block diagram of the method, apparatus (system) and computer program product according to embodiments of the application. It will be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in one or more of the flowchart or block diagram block or blocks. Figure 1 a system to perform the functions specified in one or more of the flowchart or block diagram block or blocks.
[0076] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in one or more of the flowchart or block diagram block or blocks. Figure 1 a system to perform the functions specified in one or more of the flowchart or block diagram block or blocks.
[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in one or more of the flowchart or block diagram block or blocks. Figure 1 a system to perform the functions specified in one or more of the flowchart or block diagram block or blocks.
[0078] The above merely provides the preferred but not limiting embodiments of the present application, and any modification or substitution within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A speech adaptation method based on spoken prosody tagging, characterized in that, The method comprises the following steps: S1, loading a preset tone mark according to a base dialect type selection; S2, obtaining user speech, identifying user speech information and a corresponding first tone sequence; S3, processing the user speech information using the preset tone mark to obtain a second tone sequence; S4, determining whether the difference between the first tone sequence and the second tone sequence is greater than a preset threshold value, and when the difference between the first tone sequence and the second tone sequence is not greater than the preset threshold value, registering the preset tone mark as an output tone mark; S5, when the difference between the first tone sequence and the second tone sequence is greater than the preset threshold value, adjusting the second tone sequence to the first tone sequence using a preset adjustment coefficient until the difference between the third tone sequence obtained after the adjustment and the first tone sequence is not greater than the preset threshold value; S6, modifying the preset tone mark based on the third tone sequence to obtain the output tone mark; S7, adjusting the voice output content using the output tone mark.
2. The method of claim 1, wherein, The method further comprises the following steps: S8, feeding back the voice output content to the user, obtaining the user's repeated voice output content, and identifying a corresponding fourth tone sequence; S9, replacing the first tone sequence with the fourth tone sequence, and repeating steps S4 to S7.
3. The method of claim 1, wherein, The difference degree calculation method comprises the following steps: Subtracting the second tone sequence from the first tone sequence to obtain a difference value; Dividing the difference value by the first tone sequence to obtain the difference degree.
4. The method of claim 3, wherein, The preset adjustment coefficient is greater than 0 and less than or equal to 1.
5. The method of claim 4, wherein, The method of adjusting the second tone sequence to the first tone sequence using the preset adjustment coefficient comprises the following steps: Multiplying the preset adjustment coefficient by the difference degree to obtain an adjustment target coefficient; Adjusting the second tone sequence using the adjustment target coefficient.
6. A speech adaptation system based on spoken prosody tagging, characterized by, The method comprises the following steps: A speech recognition module for identifying user speech information and a corresponding first tone sequence; A tone mark processing module for processing user speech information using a preset tone mark to obtain a second tone sequence; A tone sequence judgment module for determining whether the difference between the first tone sequence and the second tone sequence is greater than a preset threshold value; A tone sequence adjustment module for adjusting the second tone sequence to the first tone sequence using a preset adjustment coefficient to obtain a third tone sequence; A tone mark management module for modifying the preset tone mark based on the third tone sequence to obtain an output tone mark; A voice output module for adjusting the voice output content using the output tone mark.
7. A computer readable storage medium characterized by The storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1 to 5.
8. An electronic device, comprising: The device comprises a processor and a memory; The memory is used to store a preset tone mark, a preset threshold value and a preset adjustment coefficient; The processor is used to execute the method of any one of claims 1 to 5 by calling the preset tone mark, the preset threshold value and the preset adjustment coefficient.
9. A computer program product comprising computer programs and / or instructions, characterized in that, The computer program and / or instructions are executed by the processor to implement the steps of the method of any one of claims 1 to 5.
Citation Information
Patent Citations
AI voice rate adjusting method and device and electronic equipment
CN110619888A
Offline speech recognition matching device and method for finite sentence library
CN113744722A