Voice processing method and device and electronic equipment
By collecting and adjusting the weight of the speech recognition model of children's audio information and identifying and adjusting the rhythmic characteristics of children's audio information, the problem of inaccurate feedback in children's voice interaction is solved, and the children's voice interaction experience is improved.
Patent Information
- Application Number
- CN202510767530.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-08-12
AI Technical Summary
The feedback from existing smart terminal devices in children's voice interaction is not accurate enough and fails to properly guide the children's group, resulting in poor interaction effects.
By collecting children's audio information, adjusting the weight of the speech recognition model, identifying the predicted rhythm characteristics and current rhythm characteristics of children's audio information, and adjusting them when mismatch, the target rhythm characteristics are obtained for speech recognition and identifying the semantic information of children's audio information.
It realizes accurate recognition and feedback of children's audio information, and improves children's voice interaction experience.
Smart Images

Figure CN120472890A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing, and in particular to a speech processing method, device and electronic equipment. Background Art
[0002] In the wave of intelligentization, smart terminal devices have become a key area. With growing demand for intelligent interactive experiences, children's voice interaction experiences are attracting much attention. Simultaneously, voice interaction technology has seen significant progress thanks to advances in natural language processing, with improved recognition accuracy and comprehension capabilities. It is now widely used in various smart devices, laying the foundation for natural voice conversations across these devices.
[0003] In the existing technology, the designs of smart terminal devices used for human-computer interaction do not take the needs of children into consideration. The feedback on children's questions is not accurate enough, and the feedback on some questions is not properly guided for this special group of children, resulting in poor voice interaction effects for children. Summary of the Invention
[0004] The purpose of the present invention is to provide a voice processing method, device and electronic equipment that can perform voice interaction for children and improve the voice interaction effect.
[0005] In order to achieve the above objectives, the technical solutions adopted in the embodiments of the present application are as follows:
[0006] In a first aspect, an embodiment of the present application provides a speech processing method, the method comprising:
[0007] Collect children's audio information within a preset frequency band;
[0008] Determining a first weight of a trained speech recognition model based on the child audio information;
[0009] Adjusting the weight of the trained speech recognition model based on the first weight to obtain a target speech recognition model;
[0010] Determining, based on the target speech recognition model, predicted prosodic features corresponding to the child's audio information and current prosodic features corresponding to the child's audio information;
[0011] When the predicted prosodic feature does not match the current prosodic feature, adjusting the current prosodic feature based on the predicted prosodic feature to obtain a target prosodic feature;
[0012] Speech recognition is performed based on the target prosodic features to obtain semantic information of the children's audio information.
[0013] In an optional embodiment, the method further comprises:
[0014] determining child audio data in the speech recognition model;
[0015] Using the children's audio data and the children's audio information as samples to be trained for the speech recognition model;
[0016] Performing enhancement processing on the to-be-trained samples to obtain target to-be-trained samples;
[0017] The speech recognition model is trained based on the target training sample to obtain a trained speech recognition model.
[0018] In an optional embodiment, the step of determining a first weight of a trained speech recognition model based on the child audio information includes:
[0019] determining a child's voice feature of the child's audio information;
[0020] determining a target child voice feature corresponding to the child voice feature from a plurality of preset child voice features, wherein each of the preset child voice features corresponds to a different weight;
[0021] The weight corresponding to the target child's speech feature is used as the first weight of the trained speech recognition model.
[0022] In an optional embodiment, the target speech recognition model includes an acoustic module, and the step of determining the current prosodic features corresponding to the child audio information based on the target speech recognition model includes:
[0023] inputting the child audio information into the acoustic module;
[0024] When the acoustic module cannot identify the current prosodic feature corresponding to the child audio information, searching for a target pronunciation corresponding to the child voice feature from a voice dictionary;
[0025] Determining a standard vocabulary corresponding to the target pronunciation;
[0026] The standard vocabulary and the children's audio information are used as inputs of the acoustic module to obtain current prosodic features corresponding to the children's audio information.
[0027] In an optional embodiment, when the predicted prosodic feature does not match the current prosodic feature, the step of adjusting the current prosodic feature based on the predicted prosodic feature to obtain a target prosodic feature includes:
[0028] determining whether the sound quality of the target prosodic feature is clear;
[0029] When the sound quality of the target prosodic feature is clear, determining whether the target prosodic feature is complete;
[0030] In the case that the target prosodic feature is incomplete, the target prosodic feature is completed to obtain a new target prosodic feature.
[0031] In an optional embodiment, the method further comprises:
[0032] determining contextual information corresponding to the child audio information;
[0033] Using the semantic information and the context information of the child's audio information as input to a large language model, and obtaining response information corresponding to the child's audio information;
[0034] Verifying the response information, and synthesizing the verified response information into a response audio;
[0035] The response audio is output.
[0036] In an optional embodiment, the step of determining the context information corresponding to the children's audio information includes:
[0037] Determining time information of receiving the child audio information;
[0038] Comparing the time information with each preset context information, wherein each preset context information corresponds to different time information;
[0039] Determining target context information matching the time information from each of the preset context information;
[0040] The target context information is used as the context information corresponding to the children's audio information.
[0041] In an optional embodiment, the step of determining the context information corresponding to the children's audio information includes:
[0042] determining each word in the child's audio information;
[0043] Determining whether each of the words contains a preset keyword;
[0044] When the preset keyword is included in each of the words, the context corresponding to the preset keyword is determined as the context information corresponding to the children's audio information.
[0045] In a second aspect, an embodiment of the present application provides a speech processing device, the device comprising:
[0046] A collection module, used to collect children's audio information within a preset frequency band;
[0047] a determination module, configured to determine a first weight of a trained speech recognition model based on the child audio information;
[0048] an adjustment module, configured to adjust the weights of the trained speech recognition model based on the first weights to obtain a target speech recognition model;
[0049] The determination module is further configured to determine, based on the target speech recognition model, a predicted prosodic feature corresponding to the child's audio information and a current prosodic feature corresponding to the child's audio information;
[0050] The adjustment module is further configured to adjust the current prosodic feature based on the predicted prosodic feature to obtain a target prosodic feature when the predicted prosodic feature does not match the current prosodic feature;
[0051] The determination module is further configured to perform speech recognition based on the target prosodic features to obtain semantic information of the children's audio information.
[0052] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the speech processing method when executing the computer program.
[0053] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the speech processing method when executed by a processor.
[0054] This application has the following beneficial effects:
[0055] The present application collects children's audio information within a preset frequency band, determines a first weight of a trained speech recognition model based on the children's audio information, adjusts the weight of the trained speech recognition model based on the first weight to obtain a target speech recognition model, determines predicted prosodic features corresponding to the children's audio information and current prosodic features corresponding to the children's audio information based on the speech recognition model, and when the predicted prosodic features do not match the current prosodic features, adjusts the current prosodic features based on the predicted prosodic features to obtain target prosodic features, performs speech recognition based on the target prosodic features, and obtains semantic information of the children's audio information. The application can accurately identify the semantic information corresponding to the children's audio information, so as to provide accurate feedback on questions raised by children in a targeted manner, thereby improving the children's interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0057] Figure 1 A block diagram of an electronic device provided by an embodiment of the present invention;
[0058] Figure 2 One of the flow charts of a speech processing method provided by an embodiment of the present invention;
[0059] Figure 3 A second flow chart of a speech processing method provided by an embodiment of the present invention;
[0060] Figure 4 A third flow chart of a speech processing method provided by an embodiment of the present invention;
[0061] Figure 5 A fourth flowchart of a speech processing method provided by an embodiment of the present invention;
[0062] Figure 6 A fifth flow chart of a speech processing method provided by an embodiment of the present invention;
[0063] Figure 7 A sixth flow chart of a speech processing method provided by an embodiment of the present invention;
[0064] Figure 8 This is a structural block diagram of a speech processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0066] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0067] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0068] In the description of the present invention, it should be noted that if the terms "upper", "lower", "inside", "outside", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the accompanying drawings, or is the orientation or position relationship in which the product of the invention is usually placed when in use. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be understood as a limitation on the present invention.
[0069] In addition, the terms "first", "second", etc., if used, are merely used to distinguish and describe, and should not be understood as indicating or implying relative importance.
[0070] It should also be noted that, in the description of this application, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.
[0071] After extensive research, the inventors found that the designs of smart terminal devices used for human-computer interaction in the existing technology did not take the needs of children into consideration, the feedback on children's questions was not accurate enough, and the feedback on some questions did not provide appropriate guidance for this special group of children, resulting in poor voice interaction effects for children.
[0072] In view of the discovery of the above problems, the present embodiment provides a speech processing method, device and electronic device, which can collect children's audio information within a preset frequency band, determine the first weight of a trained speech recognition model based on the children's audio information, adjust the weight of the trained speech recognition model based on the first weight to obtain a target speech recognition model, determine the predicted prosodic features corresponding to the children's audio information and the current prosodic features corresponding to the children's audio information based on the speech recognition model, and when the predicted prosodic features do not match the current prosodic features, adjust the current prosodic features based on the predicted prosodic features to obtain the target prosodic features, perform speech recognition based on the target prosodic features, and obtain the semantic information of the children's audio information. The semantic information corresponding to the children's audio information can be accurately identified, so that accurate feedback can be given to the questions raised by the children in a targeted manner, thereby improving the children's interactive experience. The solution provided by this embodiment is elaborated in detail below.
[0073] This embodiment provides an electronic device capable of processing speech. In one possible implementation, the electronic device may be a user terminal, such as, but not limited to, a server, a smartphone, a personal computer (PC), a tablet computer, a personal digital assistant (PDA), a mobile internet device (MID), etc.
[0074] Please refer to Figure 1 , Figure 1 The electronic device 100 provided in the embodiment of the present application is shown in FIG. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0075] The electronic device 100 includes a speech processing device 110 , a memory 120 , and a processor 130 .
[0076] The memory 120 and the processor 130 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The voice processing device 110 includes at least one software function module that can be stored in the memory 120 in the form of software or firmware or solidified in the operating system (OS) of the electronic device 100. The processor 130 is used to execute the executable modules stored in the memory 120, such as the software function modules and computer programs included in the voice processing device 110.
[0077] The memory 120 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory 120 is used to store a program, and the processor 130 executes the program after receiving an execution instruction.
[0078] Please refer to Figure 2 , Figure 2 For application Figure 1 The flowchart of a voice processing method of the electronic device 100 is shown in FIG. , and the method including each step is described in detail below.
[0079] S201: Collecting children's audio information within a preset frequency band.
[0080] S202: Determine a first weight of the trained speech recognition model based on the child's audio information.
[0081] S203: Adjust the weight of the trained speech recognition model based on the first weight to obtain a target speech recognition model.
[0082] S204: Determine, based on the target speech recognition model, predicted prosodic features corresponding to the child's audio information and current prosodic features corresponding to the child's audio information.
[0083] S205: When the predicted prosodic feature does not match the current prosodic feature, the current prosodic feature is adjusted based on the predicted prosodic feature to obtain a target prosodic feature.
[0084] S206: Perform speech recognition based on the target prosodic features to obtain semantic information of the child's audio information.
[0085] Because children's vocalizations typically have higher frequencies than adults, we can capture children's audio information within a preset frequency band. This allows us to design a filter bank with higher resolution in the high-frequency band, capturing the high-frequency details in children's speech. This is achieved by adjusting the filter bank's bandwidth and center frequency distribution, allowing for clearer signal separation and analysis within the primary frequency range of children's speech, such as 200-8000Hz.
[0086] As children age, their language development progresses through different stages. In the early stages of language development, their vocabulary is small and their sentence structure is simple. As their language skills develop, their vocabulary increases and their sentence structure becomes more complex. Therefore, different weights can be assigned to different stages of language development.
[0087] In the early stages of children's language development, the weights of short words and simple grammatical structures in the speech recognition model can be appropriately increased, so that the speech recognition model is more inclined to recognize results that are consistent with the language habits of children at this stage. As children's language ability develops, the weights of the speech recognition model are gradually adjusted so that it can adapt to more complex vocabulary and grammatical structures, thereby achieving accurate recognition of the speech of children of different ages so that they can adapt to more complex vocabulary and grammatical structures.
[0088] The child's audio information is used as input to the target speech recognition model to obtain the predicted prosodic features and current prosodic features corresponding to the child's audio information. The predicted prosodic features corresponding to the child's audio information indicate the possible prosodic features of the child's next audio information. The current prosodic features corresponding to the child's audio information are the prosodic features corresponding to the audio information.
[0089] Exemplarily, the speech recognition model may include a prosodic modeling module and an acoustic module. The prosodic modeling module may be designed and constructed based on a recurrent neural network (RNN) or its variants (such as a long short-term memory network LSTM). RNN and LSTM can effectively process temporal information in speech. By training these models to learn the prosodic patterns in children's speech, the possible prosodic features of the next speech segment are predicted, which is the predicted prosodic features. When performing speech recognition, the predicted prosodic features are used to post-process the current prosodic features output by the acoustic module to correct recognition errors that may be caused by complex intonation. For example, if the prosodic modeling module predicts that a speech segment should have a higher fundamental frequency rising trend, and the result recognized by the acoustic module does not match the prosodic pattern, the current prosodic features can be adjusted to improve the accuracy of recognition.
[0090] The predicted prosodic features corresponding to the children's audio information are compared with the current prosodic features. When the predicted prosodic features are inconsistent with the current prosodic features, the current prosodic features are adjusted based on the predicted prosodic features to obtain the target prosodic features, so that the target prosodic features are consistent with the predicted prosodic features, and speech recognition is performed based on the target prosodic features to obtain semantic information. Since the current children's audio information may be relatively vague, if recognition is performed directly based on the current children's audio information, it will lead to incorrect recognition results. By adjusting the current prosodic features through the predicted prosodic features, the vague children's audio information can be completed, and speech recognition is performed based on the completed target prosodic features to improve the recognition accuracy.
[0091] For example, if the speech recognition model predicts that a child's audio information should have a higher fundamental frequency rising trend, and the recognized current prosodic features do not match the predicted prosodic features, the current prosodic features can be adjusted.
[0092] The essence of the mismatch between the predicted prosodic features and the current prosodic features is that the acoustic features are inconsistent with the statistical laws of children's prosodic patterns.
[0093] The match between the two can be determined by comparing the current rhythmic features with the predicted rhythmic features / sequence with the temporal pattern of the predicted rhythmic features (fundamental frequency, duration, energy, etc.), combining numerical differences, rule conflicts and context dependencies, and system quantification.
[0094] When the predicted prosodic features of the child's audio information are consistent with the current prosodic features, the current prosodic features are used as target prosodic features, and speech recognition is performed based on the target prosodic features to obtain semantic information of the child's audio information.
[0095] In order to improve the accuracy of speech recognition by the speech recognition model, it is necessary to train the speech recognition model to obtain a trained speech recognition model, such as Figure 3 As shown, the following steps are included:
[0096] S301: Determine children's audio data in a speech recognition model.
[0097] S302: Using the children's audio data and the children's audio information as samples to be trained for a speech recognition model.
[0098] S303: Perform enhancement processing on the training samples to obtain target training samples.
[0099] S304: Train the speech recognition model based on the target training sample to obtain a trained speech recognition model.
[0100] The speech recognition model needs to be continuously trained. Therefore, the speech recognition model contains the original training samples, namely children's audio data. During the use of the speech recognition model, children's audio information will be collected. The real-time collected children's audio information and the original children's audio data will be used as training samples for the speech recognition model.
[0101] Since there are fewer training samples for children, the training samples need to be expanded first. The expansion method can be to enhance the training samples, for example, adding noise to the training samples, adjusting the volume of the training samples, and stretching the time of the training samples.
[0102] Multimodal training can also be conducted based on the child's language learning environment. Background sounds in the child's environment, such as classroom noise or outdoor birdsong, can be incorporated into the training sample as a modality. This training enables the speech recognition model to accurately recognize children's speech in complex environments. This not only improves the model's applicability in real-world scenarios but also leverages background information to aid in understanding children's speech. For example, when background sounds include classroom sounds, the speech recognition model can better identify classroom-related child speech content, enhancing its adaptability to children's speech in specific scenarios.
[0103] Specific enhancement operations can also be performed on the training samples based on the characteristics of children's speech, so that the training samples are closer to the children's actual speech output and the model's adaptability to the diversity of children's speech is enhanced.
[0104] For example: simulate the common slips of the tongue, pauses and repetitions in children's pronunciation, randomly insert some common slips of the tongue, pauses or repeated pronunciations into the original speech data of the training samples, generate target training samples, and train the speech recognition model based on the target training samples, so that the trained speech recognition model can more accurately recognize children's audio information.
[0105] There are many ways to determine the first weight of the trained speech recognition model based on the child's audio information. In one implementation, Figure 4 As shown, the following steps are included:
[0106] S401: Determine child voice features of child audio information.
[0107] S402: Determine a target child voice feature corresponding to the child voice feature from a plurality of preset child voice features.
[0108] Among them, each preset child voice feature corresponds to a different weight.
[0109] S403: The weight corresponding to the target child's speech feature is used as the first weight of the trained speech recognition model.
[0110] One implementation method for determining the first weight of the trained speech recognition model is to determine the child speech feature of the child audio information, and match the child speech feature with multiple preset child speech features, each preset child speech feature having a corresponding different weight.
[0111] Exemplarily, the preset child speech features include a first preset child speech feature, a second preset child speech feature, and a third preset child speech feature. The weight of simple vocabulary recognition corresponding to the first preset child speech feature is 0.4, the weight of simple grammar recognition is 0.4, the weight of complex vocabulary recognition is 0.1, and the weight of complex grammar recognition is 0.1. The weight of simple vocabulary recognition corresponding to the second preset child speech feature is 0.3, the weight of simple grammar recognition is 0.3, the weight of complex vocabulary recognition is 0.2, and the weight of complex grammar recognition is 0.2. The weight of simple vocabulary recognition corresponding to the third preset child speech feature is 0.1, the weight of simple grammar recognition is 0.1, the weight of complex vocabulary recognition is 0.4, and the weight of complex grammar recognition is 0.4.
[0112] The child's voice feature is compared with the first preset child's voice feature, the second preset child's voice feature and the third preset child's voice feature respectively, so as to determine the target child's voice feature that matches the child's voice feature from multiple preset child's voice features, thereby determining the first weight of the trained voice recognition model that matches the child.
[0113] When the child's speech feature matches the first preset child's speech feature, the first weight of the trained speech recognition model is determined to be 0.4 for simple vocabulary recognition, 0.4 for simple grammar recognition, 0.1 for complex vocabulary recognition, and 0.1 for complex grammar recognition.
[0114] When the child's speech feature matches the second preset child's speech feature, the first weight of the trained speech recognition model is determined to be 0.3 for simple vocabulary recognition, 0.3 for simple grammar recognition, 0.2 for complex vocabulary recognition, and 0.2 for complex grammar recognition.
[0115] When the child's speech feature matches the third preset child's speech feature, the first weight of the trained speech recognition model is determined to be 0.1 for simple vocabulary recognition, 0.1 for simple grammar recognition, 0.4 for complex vocabulary recognition, and 0.4 for complex grammar recognition.
[0116] It should be noted that the first preset child voice feature, the second preset child voice feature, and the third preset child voice feature represent the stage characteristics of children's language development.
[0117] Based on the first weight, the weight of the trained speech recognition model is adjusted to obtain the target speech recognition model:
[0118] Exemplarily, the weights of the trained speech recognition model are: the weight of simple vocabulary recognition is 0.5, the weight of simple grammar recognition is 0.2, the weight of complex vocabulary recognition is 0.2, and the weight of complex grammar recognition is 0.1. When the first weight indicates that the weight of simple vocabulary recognition is 0.3, the weight of simple grammar recognition is 0.3, the weight of complex vocabulary recognition is 0.2, and the weight of complex grammar recognition is 0.2, the weight of simple vocabulary recognition of the trained speech recognition model is adjusted from 0.5 to 0.3, the weight of simple grammar recognition of the trained speech recognition model is adjusted from 0.2 to 0.3, and the weight of complex grammar recognition of the trained speech recognition model is adjusted from 0.1 to 0.2. The trained speech recognition model after weight adjustment is used as the target speech recognition model.
[0119] There are many ways to implement the process of determining the predicted prosodic features corresponding to the child's audio information based on the target speech recognition model. In one implementation, for example, Figure 5 As shown, the following steps are included:
[0120] S501: Input the children's audio information into the acoustic module.
[0121] S502: When the acoustic module cannot recognize the current prosodic features corresponding to the child's audio information, searching for a target pronunciation corresponding to the child's voice features from a voice dictionary.
[0122] S503: Determine the standard vocabulary corresponding to the target pronunciation.
[0123] S504: Using the standard vocabulary and the child audio information as inputs to the acoustic module to obtain current prosodic features corresponding to the child audio information.
[0124] Children at different stages of language development have limited vocabulary and may have inaccurate pronunciation. For young children, pronunciation simplification and substitution often occur.
[0125] In order to adapt to the above situation, a pronunciation dictionary is constructed based on the characteristics of children's pronunciation. The pronunciation dictionary not only contains the pronunciation of standard vocabulary, but also includes common children's pronunciation variants.
[0126] In order to further improve the effect of children's speech recognition, during the speech recognition process, when the pronunciation matched by the acoustic module does not find a complete match in the standard dictionary, a search can be performed in the pronunciation dictionary to find the closest target pronunciation, and the standard vocabulary and child audio information corresponding to the target pronunciation can be used as input to the acoustic module to obtain the current prosodic features corresponding to the child audio information.
[0127] In order to further improve the effect of children's speech recognition, in one implementation, as Figure 6 As shown, the following steps are included:
[0128] S601: Determine whether the sound quality of the target prosodic feature is clear.
[0129] The means for determining whether the sound quality of the target prosodic feature is clear include but are not limited to one of the following:
[0130] Determine the fundamental frequency of the target prosodic feature and detect the continuity of the fundamental frequency trajectory using the autocorrelation function or YIN algorithm. Discontinuities or jitters may indicate unclearness. Determine the harmonic-to-noise ratio corresponding to the target prosodic feature. When the harmonic-to-noise ratio is higher than the preset harmonic-to-noise ratio, it indicates that the periodicity is obvious and the sound is clear. When the harmonic-to-noise ratio is lower than the preset harmonic-to-noise ratio, it indicates hoarseness or breathiness, and the sound quality of the target prosodic feature is unclear. Determine the change in the fundamental frequency of the target prosodic feature. If the change is within the fundamental frequency change range of the whole child's voice, it indicates clarity. If it exceeds this range, it may be distorted.
[0131] S602: When the sound quality of the target prosodic feature is clear, determine whether the target prosodic feature is complete.
[0132] The means of determining whether the target prosodic feature is complete can be by checking whether there are empty values or abnormal filling values in the target prosodic feature sequence, and counting the number of control or abnormal filling values. If the ratio of the number to the total value of the target prosodic feature sequence exceeds a threshold, it is considered incomplete.
[0133] S603: When the target prosodic feature is incomplete, the target prosodic feature is completed to obtain a new target prosodic feature.
[0134] There are various methods for completing the target prosodic features. For example, probabilistic modeling of the target prosodic features based on a Gaussian mixture model can be used to generate a smooth output. Bidirectional models can be used to leverage contextual information to complete the correct portion. Linear or spline interpolation can be used to fill in missing values for the number of bottles or energy. This embodiment of the present application does not impose any specific limitations on this.
[0135] In order to further improve the effect of children's speech recognition, contextual information can also be considered, such as Figure 7 As shown, the following steps are included:
[0136] S701: Determine context information corresponding to the children's audio information.
[0137] S702: Using the semantic information and contextual information of the child's audio information as input to the large language model to obtain response information corresponding to the child's audio information.
[0138] S703: Verify the response information and synthesize the verified response information into a response audio.
[0139] S704: Output the response audio.
[0140] Because the responses to children's audio messages generated based on a large language model are sometimes complex, with intricate expressions, excessive details, and specialized vocabulary, making them difficult for children to understand, we verify these responses by extracting keywords, performing semantic clustering, simplifying sentences, and deleting non-critical information. The final output is clear, understandable, and positive. We also conduct content review on the final output to ensure it is healthy and positive, helping children absorb and apply the information.
[0141] Perform high-quality audio synthesis on the verified response information to make it match the audio output specifications on the terminal, such as audio frequency, channel, etc.
[0142] There are many ways to determine the context information corresponding to the children's audio information. In one implementation method, in one implementation method: determine the time information of receiving the children's audio information, compare the time information with each preset context information, wherein each preset context information corresponds to different time information, and from each preset context information, determine the target context information that matches the time information, and use the target context information as the context information corresponding to the children's audio information.
[0143] For example, the time information of the children's audio information can be used to determine that the context information is school time when the time information of the children's audio information is consistent with the school time period.
[0144] Specifically, when the time information of the children's audio information is 7:05 in the morning, and the preset context information is that the time information for going to school is between 6:00-9:00 in the morning, the time information of the children's audio information matches the preset context information for going to school, then the context information is determined to be going to school.
[0145] The setting of the time information corresponding to the preset context information can be set according to actual conditions and is not specifically limited here.
[0146] In another implementation: determine each word in the children's audio information, determine whether each word contains a preset keyword, and when each word contains the preset keyword, determine the context corresponding to the preset keyword as the context information corresponding to the children's audio information.
[0147] For example: the preset keywords are going to school, schoolbag, classmates, class, homework, etc., and the context information corresponding to going to school, schoolbag, classmates, class, homework is going to school. Then, it is determined whether each word in the children's audio information contains the above preset keywords. If so, it is determined that the context information corresponding to the children's audio information is going to school.
[0148] For example: the preset keywords are breakfast, keys, coat, bus, subway, goodbye mom, etc. The context information corresponding to the above keywords is going out in the morning. Then, it is determined whether the words in the children's audio information contain the above preset keywords. If so, it is determined that the context information corresponding to the children's audio information is going out in the morning.
[0149] The context information can also be determined by determining the sentence structure and time information in the children's audio information. For example, during school hours, if the sentence structure is a question ("What time do you arrive at school today?") or a reminder ("Pack your schoolbag quickly!"), the context information is determined to be going to school.
[0150] In another implementation, a geographic location where the children's audio information is received is determined, and contextual information is determined based on the geographic location.
[0151] For example, the geographic location information is determined through GPS or IP, and it is determined whether the geographic location information is near a school. If so, the context information is determined to be going to school.
[0152] It should be noted that the present application does not impose any specific restrictions on the method of determining the context, and those skilled in the art can make specific settings based on actual conditions.
[0153] Please refer to Figure 8 The embodiment of the present application also provides an application Figure 1The speech processing device 110 of the electronic device 100 includes:
[0154] The acquisition module 111 is used to acquire children's audio information within a preset frequency band;
[0155] a determination module 112 for determining a first weight of a trained speech recognition model based on the child audio information;
[0156] An adjustment module 113, configured to adjust the weights of the trained speech recognition model based on the first weights to obtain a target speech recognition model;
[0157] The determining module 112 is further configured to determine, based on the target speech recognition model, a predicted prosodic feature corresponding to the child audio information and a current prosodic feature corresponding to the child audio information;
[0158] The adjustment module 113 is further configured to adjust the current prosodic feature based on the predicted prosodic feature to obtain a target prosodic feature when the predicted prosodic feature does not match the current prosodic feature;
[0159] The determination module 112 is further configured to perform speech recognition based on the target prosodic features to obtain semantic information of the children's audio information.
[0160] The present application further provides an electronic device 100, which includes a processor 130 and a memory 120. The memory 120 stores computer-executable instructions, which, when executed by the processor 130, implement the speech processing method.
[0161] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor 130, the speech processing method is implemented.
[0162] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0163] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0164] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0165] The above descriptions are merely examples of various embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A speech processing method, characterized in that: The method comprises: Collect children's audio information within a preset frequency band; Determining a first weight of a trained speech recognition model based on the child audio information; Adjusting the weight of the trained speech recognition model based on the first weight to obtain a target speech recognition model; determining, based on the target speech recognition model, predicted prosodic features corresponding to the child's audio information and current prosodic features corresponding to the child's audio information; When the predicted prosodic feature does not match the current prosodic feature, adjusting the current prosodic feature based on the predicted prosodic feature to obtain a target prosodic feature; Speech recognition is performed based on the target prosodic features to obtain semantic information of the children's audio information.
2. The method according to claim 1, characterized in that The method further comprises: determining child audio data in the speech recognition model; Using the children's audio data and the children's audio information as samples to be trained for the speech recognition model; Performing enhancement processing on the to-be-trained sample to obtain a target to-be-trained sample; The speech recognition model is trained based on the target training sample to obtain a trained speech recognition model.
3. The method according to claim 1, characterized in that The step of determining a first weight of the trained speech recognition model based on the child audio information comprises: determining a child's voice feature of the child's audio information; determining a target child voice feature corresponding to the child voice feature from a plurality of preset child voice features, wherein each of the preset child voice features corresponds to a different weight; The weight corresponding to the target child's speech feature is used as the first weight of the trained speech recognition model.
4. The method according to claim 1, wherein The target speech recognition model includes an acoustic module. The step of determining the current prosodic features corresponding to the child audio information based on the target speech recognition model includes: inputting the child audio information into the acoustic module; When the acoustic module cannot identify the current prosodic feature corresponding to the child audio information, searching for a target pronunciation corresponding to the child voice feature from a voice dictionary; Determining a standard vocabulary corresponding to the target pronunciation; The standard vocabulary and the children's audio information are used as inputs of the acoustic module to obtain current prosodic features corresponding to the children's audio information.
5. The method according to claim 1, wherein The method further comprises: determining whether the sound quality of the target prosodic feature is clear; When the sound quality of the target prosodic feature is clear, determining whether the target prosodic feature is complete; In the case that the target prosodic feature is incomplete, the target prosodic feature is completed to obtain a new target prosodic feature.
6. The method according to claim 1, characterized in that The method further comprises: determining contextual information corresponding to the child audio information; Using the semantic information and the context information of the child's audio information as input to a large language model, and obtaining response information corresponding to the child's audio information; Verifying the response information, and synthesizing the verified response information into a response audio; The response audio is output.
7. The method according to claim 6, characterized in that The step of determining the context information corresponding to the children's audio information includes: Determining time information of receiving the child audio information; Comparing the time information with each preset context information, wherein each preset context information corresponds to different time information; Determining target context information matching the time information from each of the preset context information; The target context information is used as the context information corresponding to the children's audio information.
8. The method according to claim 6, characterized in that The step of determining the context information corresponding to the children's audio information includes: determining each word in the child's audio information; Determining whether each of the words contains a preset keyword; When the preset keyword is included in each of the words, the context corresponding to the preset keyword is determined as the context information corresponding to the children's audio information.
9. A speech processing device, characterized in that: The device comprises: A collection module, used to collect children's audio information within a preset frequency band; a determination module, configured to determine a first weight of a trained speech recognition model based on the child audio information; an adjustment module, configured to adjust the weights of the trained speech recognition model based on the first weights to obtain a target speech recognition model; The determination module is further configured to determine, based on the target speech recognition model, a predicted prosodic feature corresponding to the child's audio information and a current prosodic feature corresponding to the child's audio information; The adjustment module is further configured to adjust the current prosodic feature based on the predicted prosodic feature to obtain a target prosodic feature when the predicted prosodic feature does not match the current prosodic feature; The determination module is further configured to perform speech recognition based on the target prosodic features to obtain semantic information of the children's audio information.
10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of claims 1 to 8 when executing the computer program.