An audio data processing method, device and computer based on artificial intelligence

By combining voiceprint feature extraction and image data, and iteratively adjusting the repair parameters, the problem of low efficiency in repairing audio data with speech variations and semantic defects in traditional technologies has been solved. This has enabled efficient and accurate speech repair and improved the performance of the speech processing system.

CN119993206BActive Publication Date: 2026-02-27CHENGDU TREDE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510152179.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2026-02-27
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Traditional technologies are inefficient at processing audio data with speech variations and semantic defects, and lack real-time and effective repair methods.

Method used

Personalized acoustic features are obtained through voiceprint feature extraction technology, and combined with image data to repair speech variations and semantic defects. When residual defects are detected, the initial repair parameters are iteratively adjusted until speech variations and semantic defects are completely eliminated.

Benefits of technology

It significantly improves the real-time repair efficiency of audio data with speech variations and semantic defects, and enhances the performance of speech processing systems, especially in terms of clarity, naturalness and accuracy in speech recognition, speech synthesis and voice communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993206B_ABST
    Figure CN119993206B_ABST
Patent Text Reader

Abstract

The application relates to an audio data processing method, device and computer based on artificial intelligence. The method comprises the following steps: extracting voiceprint data of an object by using a voiceprint feature, then repairing speech variation and semantic defects in to-be-processed sound according to initial repair parameters and in combination with voiceprint and image data of the object; if the repaired sound data still has speech variation or semantic defects, adjusting the initial repair parameters and re-executing the repair process; the adjustment and repair process is iterated until it is detected that the sound data no longer has speech variation and semantic defects; when the problem cannot be detected any more, the repaired sound data is output as target sound data. The method can effectively improve the real-time repair efficiency of audio data with speech variation and semantic defects.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to an audio data processing method and device based on artificial intelligence and computer. BACKGROUND

[0002] In the traditional technology, when the audio data is damaged or missing, the signal can be compensated by a method based on spectral analysis, such as extracting frequency domain features by Fourier transform, and then using interpolation algorithm or least square method to restore the missing part. In addition, deep learning technology is also widely used in audio repair, which automatically identifies and fills in the missing or damaged part of the audio by training a neural network model (such as convolutional neural network or generative adversarial network). This method can restore natural and unnoticeable repaired audio data while maintaining audio quality. However, the traditional technology for processing audio data does not propose an effective real-time processing method for speech variation (such as accent) and semantic defects (such as grammatical errors), resulting in low real-time repair efficiency for audio data with speech variation and semantic defects. SUMMARY

[0003] Therefore, it is necessary to provide an audio data processing method, device and computer based on artificial intelligence, which can effectively improve the real-time repair efficiency for audio data with speech variation and semantic defects.

[0004] In a first aspect, the present application provides an audio data processing method based on artificial intelligence, comprising: obtaining to-be-processed sound data of a sound output object, object image data and initial sound repair parameters; performing voiceprint feature extraction on the to-be-processed sound data to obtain object voiceprint feature data; taking the initial sound repair parameters as repair constraints, repairing speech variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data to obtain processed sound data; in the case that the processed sound data still has the speech variation and / or the semantic defects, adjusting the initial sound repair parameters according to the processed sound data to obtain adjusted sound repair parameters; taking the adjusted sound repair parameters as the initial sound repair parameters, returning to execute the step of taking the initial sound repair parameters as repair constraints, repairing the speech variation and the semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data to obtain the processed sound data; until the case that the processed sound data does not have the speech variation and the semantic defects is detected, taking the processed sound data as target sound data.

[0005] In a second aspect, the present application also provides an audio data processing device based on artificial intelligence, comprising: an audio data acquisition module, configured to acquire to-be-processed sound data of a sound output object, object image data, and initial sound repair parameters; a voiceprint feature extraction module, configured to perform voiceprint feature extraction on the to-be-processed sound data to obtain object voiceprint feature data; an audio data repair module, configured to take the initial sound repair parameters as repair constraints, and repair speech variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data to obtain processed sound data; a repair parameter adjustment module, configured to, in a case where it is detected that the processed sound data still has the speech variation or / and the semantic defects, adjust the initial sound repair parameters according to the processed sound data to obtain adjusted sound repair parameters; the audio data repair module is further configured to take the adjusted sound repair parameters as the initial sound repair parameters, and return to perform the step of taking the initial sound repair parameters as repair constraints, and repairing speech variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data to obtain processed sound data; and an audio data determination module, configured to take the processed sound data as target sound data until it is detected that the processed sound data does not have the speech variation and the semantic defects.

[0006] In a third aspect, the present application also provides a computer, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements any step of the audio data processing method based on artificial intelligence when executing the computer program.

[0007] The above-mentioned audio data processing method, device and computer based on artificial intelligence obtain the personalized acoustic characteristics of an object through voiceprint feature extraction technology, and repair speech variation and semantic defects of to-be-processed sound data according to the characteristics in combination with image data. In the repair process, if it is detected that the sound data still has speech variation or / and semantic defects, the system adjusts the initial repair parameters through iteration to further optimize the repair effect. After multiple rounds of repair and parameter adjustment, the speech variation and semantic defects in the sound data are completely eliminated, and finally target sound data is output; the real-time repair efficiency of audio data with speech variation and semantic defects can be effectively improved, the performance of a speech processing system is further significantly improved, and especially in applications such as speech recognition, speech synthesis and speech communication, speech output with higher clarity, naturalness and accuracy is provided. BRIEF DESCRIPTION OF DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0009] Figure 1 An application environment diagram of the audio data processing method based on artificial intelligence in one embodiment;

[0010] Figure 2 A flowchart of the audio data processing method based on artificial intelligence in one embodiment;

[0011] Figure 3 A flowchart of the processed sound data obtaining method in one embodiment;

[0012] Figure 4 A flowchart of the first serial sound data processing method in one embodiment;

[0013] Figure 5 A flowchart of the second serial sound data processing method in one embodiment;

[0014] Figure 6 A flowchart of the third serial sound data processing method in one embodiment;

[0015] Figure 7 A flowchart of the parallel sound data processing method in one embodiment;

[0016] Figure 8 A flowchart of the first sound repair parameter adjusting method in one embodiment;

[0017] Figure 9 A flowchart of the second sound repair parameter adjusting method in one embodiment;

[0018] Figure 10 A structural block diagram of the audio data processing device based on artificial intelligence in one embodiment;

[0019] Figure 11 An internal structure diagram of the computer device in one embodiment. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0021] The application embodiment provides an audio data processing method based on artificial intelligence, which can be applied to the application environment as shown in the figure. Figure 1 The terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The server 104 can be realized by an independent server or a server cluster composed of multiple servers.

[0022] In an exemplary embodiment, as shown in the figure Figure 2 An audio data processing method based on artificial intelligence is provided. The method is applied to the server in Figure 1 for example, including the following steps 202 to 212. Among them:

[0023] Step 202, obtaining the sound data to be processed of the sound output object, the object image data and the initial sound repair parameter.

[0024] The sound output object can be an entity or individual that emits audio data. It can be a specific person, device, virtual character or any object that can produce voice output.

[0025] The sound data to be processed can be original voice data without repair, which usually contains voice input, environmental noise, speech rate, tone and other information of the sound output object.

[0026] The object image data can be visual information about the sound output object obtained by a camera or other device, including facial expressions, body language, environmental background, etc.

[0027] The initial sound repair parameter can be the repair constraint and adjustment rule set at the beginning of the repair process. These parameters can include the preliminary goal of audio repair, the range of repair, the correction standard of voice features, etc.

[0028] Specifically, the system obtains the sound data to be processed of the sound output object. These sound data are usually collected from the actual environment, which may have problems such as voice variation, noise or semantic ambiguity. At the same time, the system also obtains the object image data related to the sound output object, such as the face image or other characteristic image data of the person. These images help to understand and judge the characteristic information of the object, such as facial expression, mouth shape, emotion, scene, etc. In addition to these data, the system also needs an initial sound repair parameter, which usually includes the basic setting of the sound repair algorithm, such as repair intensity, frequency range, voice model, etc.

[0029] Step 204, voiceprint feature extraction is performed on the sound data to be processed to obtain the object voiceprint feature data.

[0030] Among them, the voiceprint feature extraction can be through analyzing the frequency, timbre, pitch, speech pattern and other information in the sound data, and extracting the individual physiological characteristics of the voice. The voiceprint of each person is unique, and the voiceprint feature extraction technology can help the system to identify and establish the personalized voice model of the object, and provide accurate reference for the subsequent sound repair, especially in repairing accent, pronunciation and other speech variations.

[0031] Among them, the object voiceprint feature data can be the voice feature data about the sound output object obtained by the voiceprint feature extraction technology. It includes the individual timbre, pitch, pronunciation habit, etc., which can reflect the unique voice characteristics of the object. In the repair process, these voiceprint data will be used as a reference basis to ensure that the repaired sound can accurately reflect the personalized voice characteristics of the object and try to retain its natural voice properties.

[0032] Specifically, the voiceprint recognition technology is used to extract the personalized voiceprint features of the object from the sound data to be processed to obtain the object voiceprint feature data. The voiceprint feature refers to the unique voice characteristics of each person, which usually includes the frequency spectrum characteristics of the voice, the pitch, the tone, the speed of the voice, etc. These features can help the system better understand the characteristics of the object when speaking.

[0033] Step 206, using the initial sound repair parameters as the repair constraint condition, according to the object voiceprint feature data and the object image data, repairing the voice variation and semantic defects of the sound data to be processed to obtain the processed sound data.

[0034] Among them, the repair constraint condition can be the rules and restrictions that must be followed in the sound repair process. It can include the initial set target (such as repairing the speed, pitch, timbre, etc.), the individual requirements of the object (such as specific timbre, emotional expression, etc.), and environmental factors, etc.

[0035] Among them, the voice variation can be the deviation in the audio data due to pronunciation, accent, speech rate, sound quality, etc. Voice variation may include local accent, unclear pronunciation, non-standard tone, etc., which usually affects the intelligibility and naturalness of the voice.

[0036] Among them, the semantic defect can be the part of the voice data that is not accurate, not clear in logic or not complete in expression, which may be caused by errors, improper expression or insufficient context understanding in the audio data output by the sound output object. For example, improper use of words, grammatical errors, information loss, etc. will affect the transmission of semantics.

[0037] Among them, the processed sound data can be sound data that has been repaired and preliminarily adjusted. These data have repaired part of the speech variation and semantic defects, but may still need further optimization.

[0038] Specifically, the system first compares and analyzes the object voiceprint feature data with the pronunciation data of standard pronunciation in the database, identifies the types of variations in the speech, such as local accent, non-standard pronunciation, too fast or too slow speed, etc. Next, using the object voiceprint feature data, combined with audio signal processing technology, the speech parameters such as pitch, timbre, and speed of the sound to be processed are adjusted specifically, so as to correct the deviation of accent or pronunciation. For example, if it is detected that the speed is too fast, the system will adjust the speech rhythm according to the normal speed range of the object to ensure clear and understandable pronunciation; if there is ambiguous pronunciation or distorted timbre, the system will restore the natural timbre of the audio through spectrum adjustment and sound quality enhancement technology. In addition, noise suppression algorithm is also applied to remove background noise in the recording and clarify the speech signal. This repair process is iterative, and after each repair, the system analyzes the repair effect and further optimizes it until the repaired sound is highly consistent with the original voiceprint features of the object, and the repair data of speech variation is output.

[0039] Further, the role of the object image data in this stage is to provide background information for speech repair, especially in terms of emotion reasoning and context understanding. The system analyzes the information of facial expressions, body language, and surrounding environment extracted from the object image. For example, through image recognition technology, the system can judge the facial expressions of the object (such as smiling, frowning, etc.), and thus infer the emotional state of the object (such as happy, angry, nervous, sad, etc.). In addition, environmental factors (such as whether there are other people present, whether the surrounding sound is noisy, etc.) can also be judged through visual clues in the image data. The system provides the context according to these information, helping to repair the emotional expression and context of the sound.

[0040] The system not only adjusts the speech features of the sound (such as pitch, speed, intonation, etc.), but also infers the emotional state of the object using the analysis results of the object image data, and then repairs the semantic defects according to the emotion reasoning. For example, when the system detects that the object expresses an angry emotion, it will not only make the intonation of the sound more powerful and urgent, but also may adjust the tone of the sentence to ensure that the repaired sound conforms to the emotional color of anger. In this process, emotion reasoning not only affects the performance of the audio, but also identifies semantic deviations in the speech, for example, some words or sentences may not be clearly expressed due to lack of emotion, and emotion reasoning technology will identify these inaccurate semantic information, appropriately adjust the emotional expression of the sentence, so that the speech in the repair process is more consistent with the actual emotional expression of the object.

[0041] In addition, emotional reasoning not only repairs variations in the sound level, but also identifies and repairs semantic defects. For example, in some emotional expressions, words may be used incorrectly, or the emotional expression of a sentence does not match the context, causing the listener to not accurately understand. Through reasoning of emotional semantics, the system adjusts the sentence structure, word selection, and ensures that the voice information accurately conveys emotions while the semantics are clear and accurate. For example, if the object is speaking in an angry mood, the original sentence may not accurately convey the emotion because it is too simple or cold. The system will repair the sentence according to the reasoning so that it can accurately express both semantics and emotions.

[0042] In this step, the system further optimizes the repair effect through context understanding technology. Specifically, the context understanding technology identifies semantic defects in the processed sound data by analyzing the language context, dialogue situation, and background information of the object based on the analysis results of the object image data. These defects may include improper use of words, grammatical errors, unclear sentence structure, or inconsistent sentence logic. For example, in a certain dialogue scenario, the use of certain words may cause information to be ambiguous or misunderstood. The system combines the analysis results of the object image data with context inference to adjust and repair these sentences to ensure that their semantics are more consistent with the true intent of the dialogue. Through context understanding, the system can not only analyze language errors in sentences themselves, but also adjust the direction of language expression to adapt to specific communication purposes and situational needs based on the current language environment of the object.

[0043] Finally, the neural cognitive repair technology performs deep repair on the processed sound data by simulating the way the brain processes language. Its core advantage is its ability to handle complex language and speech characteristics, especially at the semantic level, to identify and repair potential semantic ambiguities, misunderstandings, or missing parts. For example, when the system finds that the object's expression is not clear enough or the grammatical structure is incorrect, the neural cognitive repair technology can perform semantic corrections and reconstruct sentences based on context. This process not only includes simple grammatical corrections, but also involves deep semantic reasoning to ensure that the repaired voice not only meets the emotional and contextual requirements of the object, but also eliminates semantic ambiguities or information gaps at a deep level.

[0044] In addition, the neural cognitive repair technology can also handle some potential cognitive biases in speech, such as mishearing, misunderstanding, or unclearness in speech synthesis. Through in-depth analysis and correction of voice information, the system can make the voice clearer and more accurate, while eliminating cognitive differences and understanding biases in voice communication; the above repair process is a multi-level, multi-dimensional repair, and the final output is the processed sound data.

[0045] Step 208, in the case of detecting that the processed sound data still has speech variation or / and semantic defects, adjusting the initial sound repair parameters according to the processed sound data to obtain adjusted sound repair parameters.

[0046] Wherein, the adjusted sound repair parameters can be new parameters obtained by the system optimizing and resetting the initial repair parameters when detecting that the processed sound data still has speech variation or semantic defects.

[0047] Specifically, after the initial repair, the system will detect the processed sound data to determine whether there are still speech variations or semantic defects; if it is detected that these problems have not been completely eliminated, the system will adjust the initial repair parameters according to the feedback of the processed data. The adjustment may include changing the parameters of the repair algorithm, such as repair intensity, algorithm weight, repair precision, etc., so as to more accurately solve the remaining problems in the next round of repair process. In addition, the system may optimize the repair strategy according to the new feedback information, for example, different repair methods are used for different speech defects, to obtain the adjusted sound repair parameters.

[0048] Step 210, taking the adjusted sound repair parameters as the initial sound repair parameters, returning to execute the step of repairing the speech variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data, to obtain the processed sound data, taking the initial sound repair parameters as the repair constraints.

[0049] Specifically, the adjusted adjusted sound repair parameters will be re-used as new initial sound repair parameters, and the system will return and re-execute the repair step. At this time, the repair process will continue to be based on the updated repair parameters, and the to-be-processed sound data will be repaired again in combination with the object voiceprint feature data and the object image data of the object. Through this cyclic repair, the system can gradually optimize the repair effect and solve the possible speech variation and semantic defects in the sound data, ensuring that each step of repair is more accurate to achieve higher sound quality and accuracy.

[0050] Step 212, until the processed sound data cannot be detected to have speech variation and semantic defects, taking the processed sound data as the target sound data.

[0051] Wherein, the target sound data can be the final sound data generated after multiple rounds of repair and adjustment, which meets the requirements of the object. The target sound data not only repairs the speech variation and semantic defects, but also maintains the personalized timbre and emotional expression of the object, ensuring clear and natural speech and meeting the expression needs of the object in a specific situation.

[0052] Specifically, this cycle of repair process will continue until the system detects that there are no longer any speech variations or semantic defects in the processed sound data. As new repair parameters are obtained after each iteration of the previous repair parameters, the quality of the sound data is gradually improved through re-repairing and adjusting the original data until the predetermined repair standard is reached. Finally, when the system confirms that the processed sound data has perfectly eliminated all problems of speech variations and semantic defects, and no further defects can be detected, the repair process ends. At this time, the system regards the processed sound data as the target sound data and outputs the final repair result.

[0053] In the above-mentioned audio data processing method based on artificial intelligence, the individualized acoustic features of the object are obtained through voiceprint feature extraction technology, and the speech variations and semantic defects of the sound data to be processed are repaired according to these features combined with image data. In the repair process, if speech variations and / or semantic defects are still detected in the sound data, the system will adjust the initial repair parameters through iteration to further optimize the repair effect. After multiple rounds of repair and parameter adjustment, the speech variations and semantic defects in the sound data are completely eliminated, and the final target sound data is output; it can effectively improve the real-time repair efficiency of audio data with speech variations and semantic defects, further significantly improve the performance of the speech processing system, especially in speech recognition, speech synthesis and speech communication applications, providing speech output with higher clarity, naturalness and accuracy.

[0054] In an exemplary embodiment, as shown in Figure 3 With the initial sound repair parameters as the repair constraint condition, the speech variations and semantic defects of the sound data to be processed are repaired according to the object voiceprint feature data and the object image data to obtain the processed sound data, including steps 302 to 306. Among them:

[0055] Step 302, with the initial sound repair parameters as the repair constraint condition, the speech variations and semantic defects of the sound data to be processed are repaired in sequence according to the object voiceprint feature data and the object image data to obtain the serially processed sound data.

[0056] Among them, the serially processed sound data can be different problems in the sound to be repaired processed in a fixed order. In this mode, the repair process is divided into multiple stages, first the system will process the speech variations part, such as accent, pronunciation not standard, etc., and then gradually repair the semantic defects to ensure that the sentence is clear, logical and accurate in emotion transmission. The processing result of each stage will be used as the input of the next stage until all repair steps are completed.

[0057] Specifically, the system utilizes the initial sound repair parameters as repair constraints and combines the voiceprint feature data and image data of the subject to gradually perform repair processing. First, the voiceprint feature data of the subject is compared and analyzed with the pronunciation data of standard pronunciation in the database to identify the speech variation parts in the speech, such as accent, unclear pronunciation, inconsistent speech speed, etc. Combined with audio signal processing technology, the pronunciation deviation of the sound data to be processed is adjusted to restore normal speech features, ensuring that the repaired pronunciation conforms to the personalized features and standard pronunciation mode of the subject. After repairing the speech variation, the system will enter the semantic defect repair stage. By analyzing the image data (such as facial expressions, eye contact, body language, etc.) and voiceprint feature data (such as tone, speed, tone, etc.) of the subject, the emotional state and current context of the subject are inferred. For example, by changes in facial expressions and tone, the system can identify whether the subject is expressing anger, anxiety, or happiness, etc. emotions, so as to determine whether there are semantic defects in the sound to be processed, such as whether the subject's words are incorrect, the sentence is not smooth, or the grammar is not standard, etc. under the current emotional state and current context; when the subject's speech has the above problems, or there are unclear expressions, logical confusion or language that does not conform to common expression methods, the system will combine the specific information of the subject's emotional state and current context and repair these problems on the data after repairing the speech variation. For example, if the subject's speech uses inaccurate or inappropriate words, the system will infer more appropriate words according to the context and context and replace them on the data after repairing the speech variation; if the sentence structure is not smooth, the system will reorganize the sentence to make it more consistent with the language expression habit, ensuring that the grammar is correct and the logic is clear; if there is ambiguity or unclear expression in the speech, the system will adjust the sentence according to the subject's emotion and intention to make the semantics more clear and explicit. Through these ways, the system can effectively repair the semantic defects in the speech, so that the final speech expression is not only smooth and fluent, but also accurately conveys the real intention and emotion of the subject. This process is sequentially executed, meaning that the repair of speech variation will be completed before the repair of semantic defects, resulting in serially processed sound data.

[0058] Step 304, with the initial sound repair parameters as repair constraints, the voice variation and semantic defects of the sound data to be processed are repaired in parallel according to the voiceprint feature data of the subject and the image data of the subject, to obtain parallelly processed sound data.

[0059] In the parallel processing mode, the voice variation repair and the semantic defect repair are performed simultaneously, instead of strictly in sequence. The system utilizes the parallel computing capability to perform the repair in different processing modules respectively for the voice features (e.g. pronunciation, tone, etc.) and the semantic features (e.g. grammar, word usage, etc.), and then integrates the results of different processing modules.

[0060] Specifically, the system simultaneously repairs the voice variation and the semantic defect of the sound data to be processed. Since the problems to be repaired do not change, only the order of repair changes from serial to parallel, the system still uses the initial sound repair parameters as the constraint condition, and combines the object voiceprint feature data and the object image data to repair the voice variation and the semantic defect of the sound data to be processed. Unlike the serial repair, the system can repair both aspects at the same time through parallel processing: on the one hand, as in step 302, the object voiceprint feature data is first compared and analyzed with the pronunciation data of the standard pronunciation in the database to identify the voice variation parts in the voice, such as accent, unclear pronunciation, inconsistent speech speed, etc., and the audio signal processing technology is combined to adjust the pronunciation deviation of the sound data to be processed, restore the normal voice features, and ensure that the pronunciation after repair conforms to the personalized features of the object and the standard pronunciation mode.

[0061] On the other hand, the semantic defects are also repaired in the same way as the implementation of step 302, but the input data is different. The serial repair mode is to repair the semantic defects on the data after the voice variation repair, while the parallel repair mode is to repair the semantic defects on the sound data to be processed, that is, by analyzing the image data of the object (such as facial expression, eye contact, body language, etc.) and the voiceprint feature data (such as tone, speed, tone, etc.), the emotional state and the current context of the object are inferred. For example, by the change of facial expression and tone, the system can identify whether the object is expressing anger, anxiety or happiness, etc. emotions, so as to judge whether there are semantic defects in the sound to be processed, such as whether the words used by the object are wrong, the sentences are not smooth or the grammar is not standard, etc. under the current emotional state and the current context; when the above problems exist in the voice of the object, or there are unclear expressions, logical confusion or language that does not conform to the common expression method, the system will combine the specific information of the emotional state and the current context of the object to repair these problems on the sound data to be processed. For example, if the object's voice uses inaccurate or inappropriate words, the system will infer more appropriate words according to the context and context, and replace them on the sound data to be processed; if the sentence structure is not smooth, the system will reorganize the sentence to make it more consistent with the language expression habit, ensuring that the grammar is correct and the logic is clear; if there is ambiguity or unclear expression in the voice, the system will adjust the sentence according to the object's emotion and intention to make the semantics more clear and explicit. Through these ways, the system can effectively repair the semantic defects in the voice, so that the final voice expression is not only smooth and fluent, but also accurately conveys the real intention and emotion of the object. The data after the voice variation repair and the data after the semantic defect repair are integrated to obtain the parallel processing sound data.

[0062] Step 306, cross-pipeline fusion of the serial processing sound data and the parallel processing sound data to obtain the processed sound data.

[0063] Among them, the cross-pipeline fusion can be a process of integrating the results of serial processing and parallel processing. At this stage, the system combines and optimizes the data results from the serial repair path and the parallel repair path to achieve higher quality repair effect; serial processing can usually provide higher precision repair, while parallel processing can speed up the repair process. The system intelligently integrates the advantages of the two, combines their repair results, adjusts the weight of different repair paths, and eliminates redundant parts.

[0064] Specifically, since the serial and parallel repairs each have their own advantages, the system needs to intelligently combine the results of the two repair paths according to different repair needs. For example, in some cases, serial repair may perform better in detail repair (such as accurate repair of voice variation), while parallel repair is more efficient in handling large-scale semantic defect repair. The system will perform cross-pipeline fusion on the repair results of the two, evaluate the results from the serial and parallel repair paths during cross-pipeline fusion, analyze the advantages and disadvantages of each path in handling voice variation and semantic defects, intelligently fuse the repair results of the two, combine their respective advantages, adjust the weight of the repair effect, eliminate repeated or redundant repair steps, while ensuring that the final sound data meets the repair accuracy requirements and the repair effect is not distorted or conflicting, and finally generate processed sound data.

[0065] In this embodiment, by serially and in parallel processing the repair processes of voice variation and semantic defects, and through cross-pipeline fusion, the accuracy and efficiency of sound repair can be effectively improved. The serial repair ensures the systematicness and coherence of the repair process, with each step addressing specific problems one by one, while the parallel repair speeds up the entire repair process by simultaneously handling voice and semantic problems, minimizing processing time. In addition, cross-pipeline fusion combines the advantages of the two processing methods, retaining the fine-tuned adjustments of serial processing and benefiting from the efficiency of parallel processing, thereby obtaining more accurate and efficient sound repair results. This multi-level, multi-angle repair method can effectively improve the quality of sound data and ensure that the repaired sound achieves the best balance in terms of voice clarity and semantic accuracy.

[0066] In one exemplary embodiment, as shown in FIG. 4, the initial sound repair parameters are used as repair constraints, and the voice variation and semantic defects of the sound data to be processed are sequentially repaired in series according to the object voiceprint feature data and the object image data, to obtain serially processed sound data, including steps 402 to 406. Among them: Figure 4

[0067] Step 402, using the initial sound repair parameters as repair constraints, adjusting the pronunciation data of the sound data to be processed according to the object voiceprint feature data to obtain accent-weakened sound data.

[0068] Wherein, the pronunciation data can be various information about pronunciation extracted from the sound data to be processed by speech recognition technology, including pitch, timbre, syllable, tone, speed, etc.

[0069] Wherein, the accent-weakened sound data can be voice data after repair, reducing or eliminating non-standard accent features in the pronunciation of the object.

[0070] ​Specifically, taking the initial sound repair parameters as repair constraints, the system will compare the voiceprint features of the object (such as pitch, tone, speech rate, etc.) with the pronunciation data of the standard pronunciation in the database, identify the accent features and pronunciation deviations of the object. According to the characteristic data of these differences, the system will apply specific algorithm models or audio signal processing techniques (such as multi-dimensional speech feature joint optimization recognition algorithm, cross-domain speech feature conversion algorithm, speech synthesis and accent correction algorithm based on transformation learning, etc.) to gradually optimize and adjust the pronunciation part. For example, if the object has a certain local accent, the system will adjust the pronunciation mode to make it closer to the standard Mandarin pronunciation or the target pronunciation style. The process not only includes the correction of syllables, tones, and intonations, but also ensures that the adjusted speech remains natural and smooth, avoiding making the speech too artificial or not conforming to the object's individual voiceprint features. Finally, through these adjustments, the system generates accent-weakened sound data, making the speech clearer, standard, and meeting the target pronunciation requirements.

[0071] Step 404, according to the object image data and the accent-weakened sound data, inferring the generation environment of the to-be-processed sound data to obtain scene feature data.

[0072] Among them, the scene feature data can be analyzed by analyzing the object image data (such as facial expressions, eye contact, posture, etc.) and speech data to infer the specific situation or environment information of the object.

[0073] Specifically, the system uses object image data for situation inference to analyze and infer the environment of the object, where the object image data provides important non-verbal information about the object's facial expressions, eye contact, posture, and body language, reflecting the object's current emotional state and context. For example, facial expressions can help the system determine whether the object is in a state of tension, joy, or confusion, and eye contact and posture can reveal the object's spatial environment (such as home, office, or outdoors). The system compares the above-identified image data with known situation models, and through emotional analysis and situation inference, obtains specific scene feature data, including environmental noise (such as background music, crowd noise, etc.), conversation atmosphere (such as formal or informal), and emotional color (such as relaxed, serious, etc.).

[0074] Step 406, according to the scene feature data, modifying the language expression data of the accent-weakened sound data to obtain serial processing sound data.

[0075] Specifically, the system further repairs the language expression of the voice content after accent weakening according to the previously inferred scene feature data. In the specific implementation process, the system will infer the language expression accuracy and language style used in a specific environment according to the scene feature data. For example, in a formal meeting or business occasion, the sentences or words in the voice need to be expressed more formally and normatively, and words of profanity should be avoided; while in a friend gathering or family communication, the sentences or words in the voice should be more casual and friendly, and sentences that do not match the current scene (such as the use of inappropriate words like "deserve it" by a friend in difficulty) should be avoided. By identifying the language needs in these scenes, the system adjusts the vocabulary, tone and sentence structure in the voice to conform to the language expression in the current environment. In addition, the system also corrects the tone intensity, emotional communication and other aspects according to the context, so that the voice content is more consistent with the emotional state and situation of the object. Finally, the system generates serial processing sound data.

[0076] In this embodiment, by combining the object voiceprint feature data and the object image data, multi-level repair of the sound data to be processed is realized. Among them, the adjustment of pronunciation data based on voiceprint features can effectively weaken the difference in accent, making the sound more standardized and easy to understand. And combining with the object image data to infer the environment, the actual scene features behind the sound can be identified, providing support for further context analysis. Finally, relying on the scene feature data to modify the language expression ensures the accuracy and naturalness of the sentence in different contexts. This process optimizes pronunciation and semantics through precise steps, so that the repaired sound is not only clearer, but also better adapts to specific situations and expression needs, thereby improving the comprehensiveness and adaptability of sound repair.

[0077] In one example embodiment, as shown in Figure 5 According to the scene feature data, the language expression data of the accent-weakened sound data is modified to obtain serial processing sound data, including steps 502 to 508. Among them:

[0078] Step 502, context semantic recognition is performed on the accent-weakened sound data to obtain sound semantic recognition data.

[0079] Among them, the context semantic recognition can be that when analyzing speech or text, the meaning of a word, phrase or sentence is not only dependent on its individual word meaning, but also on its position and relationship in the whole conversation or context. For example, the word "bank" can refer to a financial institution or a riverbank, depending on the context. In actual operation, the system identifies the structure of the sentence by using natural language processing (NLP) techniques such as dependency syntax analysis and semantic role labeling, and judges the accurate meaning of the vocabulary according to the context.

[0080] The sound semantic recognition data can be a dataset that analyzes the sound data to be processed and extracts semantic information from the speech. This process typically involves parsing the meaning of each word or sentence in the speech, identifying possible semantic ambiguities, vague or unclear expressions. For example, the system may identify certain words in the speech that are semantically ambiguous due to unclear pronunciation or unclear context, and record this information as sound semantic recognition data.

[0081] Specifically, through natural language processing technology, the accent- weakened sound data is converted into text using speech recognition technology, and analyzed in combination with context information. At this time, the system not only analyzes the independent meaning of each word or phrase, but also identifies the overall structure of the sentence, grammatical relationships and logical coherence. For example, the system will identify and label key words, verbs, nouns, etc. in the sentence, and check their grammatical correctness in the current context. At the same time, the system will detect possible semantic ambiguities in the speech, such as incorrect use of vocabulary (for example, incorrect context of "improve"), and grammatical inconsistencies (such as subject-verb disagreement). Through contextual semantic recognition, the system can extract the true intent behind each sentence and generate sound semantic recognition data.

[0082] Step 504, according to the sound semantic recognition data, the objective semantic defects of the accent-weakened sound data are repaired to obtain semantic-repaired sound data.

[0083] The objective semantic defects can be a situation in which the speech or text is not clear, ambiguous or inaccurate due to grammatical, logical, vocabulary selection or sentence structure problems.

[0084] The semantic-repaired sound data can be sound data that has been repaired to eliminate objective semantic defects in the speech or text, making it more clear, accurate and contextually appropriate.

[0085] Specifically, according to the previously extracted sound semantic recognition data, the objective semantic defects in the content of the accent-weakened sound data are repaired, which include objective grammatical errors, incorrect word usage that does not conform to conventional expression methods, and logical confusion, i.e. the expression method used is not a normal expression method. For example, when the system finds that a certain word in the sound semantic recognition data is used incorrectly (for example, in the process of resource transaction, the word "buy" should be used for the object, but the system identifies that the object uses "take away"), or finds that a certain sentence has grammatical inconsistencies (such as lack of conjunctions leading to ambiguous sentence meaning), it will automatically correct it. Through natural language processing techniques such as grammar correction and synonym replacement, the system can repair these semantic defects to generate semantic-repaired sound data, making the speech content more consistent with standard grammar and more accurate in expression, avoiding any errors that may cause ambiguity.

[0086] Step 506, according to the scene feature data and the object image data, the emotion of the sound output object is recognized, and object scene recognition data is obtained.

[0087] Among them, the object scene recognition data can be the analysis of the behavior, expression, context, environment and other external information of the object, and the system infers the data of the current emotional state, intention or specific scene of the object.

[0088] Specifically, the system will identify and infer the emotional state of the object according to the information from the object image data and the scene feature data. Specifically, by analyzing the emotional signals (such as tone, speed, pitch change) in the facial expression, eye contact, body language and voice of the object in the object image data, the system can infer the emotional state of the object in the current conversation. For example, if the facial expression of the object shows nervousness or anxiety, the system may judge that the object is experiencing a state of nervousness or anxiety, and vice versa, if the facial expression is relaxed and happy, the system may infer that the object feels relaxed or happy. Combined with the scene feature data (such as the type of background noise, the environment atmosphere, etc.) of the current object, it also provides strong support for emotion recognition, helping the system to capture the emotional expression and situation of the object, and generates object scene recognition data, that is, the system's evaluation of the object's emotional state, which provides a guide for the emotional expression of the next voice repair.

[0089] Step 508, according to the object scene recognition data, the scene semantic defects of the semantic repair sound data are repaired, and serial processing sound data is obtained.

[0090] Specifically, the system corrects the emotion and context of the sound data that has repaired the objective semantic defects according to the generated object scene recognition data, that is, repairs the scene semantic defects; where the scene semantic defects refer to the places where the words or sentences of the voice in expression may not match the current emotion and scene. These defects often manifest as mismatching of words or sentences in emotional transmission or incoordination of voice tone. For example, in a serious occasion, if the words or the expression of the whole sentence used by the voice are too casual or relaxed, the system will correct the words or adjust the expression of the whole sentence of the voice through the emotional recognition data, so that it is more consistent with the formal or tense situation; if the object is expressing sad or depressed emotions, the system will correct the sentence expression and sentence structure in the voice accordingly, so that it is consistent with the emotional state. At the same time, the system will also consider other factors in the context, such as whether formal language is needed or whether some words should be avoided for too harsh expression. Through these adjustments, the system ensures that the final voice is not only grammatically and logically correct, but also uses reasonable expression and reasonable words in the emotional state of the object and the situation, generating serial processing sound data.

[0091] In this embodiment, through context semantic recognition, the intelligibility of the voice after accent weakening is ensured at the semantic level, and the possible expression ambiguity of the sentence is solved. The objective semantic defects are repaired, and the accuracy and fluency of the language are further optimized. By combining scene feature data and object image data, the system can accurately identify the emotional state of the object, and give the voice data more rich emotional color in language expression. Finally, the scene context defects in the semantics are repaired using emotional and scene recognition data, so that the language expression of the repaired voice not only conforms to the context more, but also accurately conveys emotions and intentions. Overall, through the multi-dimensional repair process, the language expression of the voice is more natural, accurate and emotional, and the transmission effect and user experience of the voice data are comprehensively improved.

[0092] In one exemplary embodiment, as shown in Figure 6 According to the object scene recognition data, the scene semantic defects of the semantic repair voice data are repaired to obtain serial processing voice data, including steps 602 to 606. Among them:

[0093] Step 602, according to the emotional reasoning data of the object scene recognition data, the voice emotional expression of the semantic repair voice data is repaired to obtain emotional enhancement recognition data.

[0094] Among them, the voice emotional expression can be the features such as words, sentences, paragraphs and structures in the voice to convey the emotional state or attitude of the speaker.

[0095] Among them, the emotional enhancement recognition data can be to analyze and process the voice data to be processed, and to enhance the emotional information in the voice data by modifying the expression of the features such as words, sentences, paragraphs and structures, so that the repaired voice can better convey the real emotions of the speaker.

[0096] Specifically, the problem of insufficient emotional expression in speech is repaired based on sentiment reasoning data in object scenario recognition data. Since sentiment reasoning data is derived by analyzing the emotional characteristics of the subject's facial expressions, body language, speech tone, and speaking content, it helps the system understand the subject's current emotional state (such as joy, anger, sadness, etc.). The system uses sentiment reasoning data to adjust the sound data that has already undergone semantic repair. Specifically, the system evaluates the use of words, sentence expression, paragraph expression, overall structure, and other parameters in the speech, and optimizes them based on the inferred sentiment reasoning data, i.e., combines these sentiment reasoning data with the semantic repair sound data to improve the emotional expression of words, sentences, paragraphs, and other parts of the speech that are insufficient or incorrect. For example, the original sentence may be written as "I feel a little uncomfortable", but this expression is relatively flat and does not adequately reflect the subject's emotional intensity. The system will modify the sentence to "I feel very uneasy in my heart, almost to the point of collapse", and through more intense emotional vocabulary and tone correction, the intensity and accuracy of emotional expression are improved. Finally, the system outputs sentiment-enhanced recognition data, which makes the emotional expression of the speech more accurately aligned with the subject's emotional state, and reflects more rich emotional colors in the speech.

[0097] Step 604, according to the scene reasoning data of the object scenario recognition data, the dialogue theme scene of the sentiment-enhanced recognition data is repaired to obtain scene-enhanced recognition data.

[0098] Wherein, the dialogue theme scene can be the context, background, discussion theme or context involved in the dialogue.

[0099] Wherein, the scene-enhanced recognition data can be to analyze and infer the context, background and dialogue theme scene of the subject, by modifying the expression of words, sentences, paragraphs and structure, to enhance these scene information, so as to better understand the potential semantic requirements in the speech.

[0100] Specifically, the system further repairs the voice according to the scene reasoning data of the object context recognition data, ensuring that it conforms to the actual scene of the dialogue. The scene reasoning data is inferred by analyzing the environmental background, context and content of the dialogue of the object, for example, if the system detects that the object is in a formal meeting situation, the expression in the voice should be more precise and standard; if the object is in a casual family gathering, the tone and words should be more relaxed and casual. The system first determines from the emotion enhancement recognition data whether these contents are consistent with the current scene, and repairs the words, sentences, paragraphs or structures that are not consistent with the current scene; assuming that in a business meeting scenario, the original sentence may express "we can consider this matter later", which is too casual and not precise enough in a formal setting. The system determines through scene reasoning to recommend modifying it to "we can further discuss and make decisions in the later meeting". Similarly, if the dialogue is in a casual gathering, the system may suggest adjusting the tone to "let's talk about it later, don't rush, let's discuss it together". This repair ensures the consistency of the scene enhancement recognition data in content and tone, ensuring that the language can express emotions and match the background and social occasion of the dialogue.

[0101] Step 606, simulating neural cognitive optimization of the linguistic features of the scene enhancement recognition data to obtain serial processing sound data.

[0102] Among them, the simulation of neural cognitive optimization can be a technology that optimizes voice processing and understanding by simulating the human brain's cognitive process. It combines the principles of neural networks and cognitive psychology to simulate how humans understand, analyze and generate voice or language. Through simulation of neural cognitive optimization, the system can more intelligently identify subtle emotional changes, semantic defects and potential problems in language expression in the voice. For example, when the system identifies grammatical errors or unnatural sentences in the voice, the simulation of neural cognitive optimization can optimize the sentence structure, adjust the expression method, and make it more consistent with human cognitive patterns, thereby improving the naturalness and accuracy of voice output.

[0103] Specifically, the system conducts further linguistic optimization on the voice data after emotional and scene repair, focusing on correcting deficiencies in language structure and improving grammatical correctness, fluency, and logicality. Through simulated neural cognitive optimization technology, the system simulates how the human brain processes language, adjusting grammatical errors, unclear expressions, or logically inconsistent parts in the voice data. Assuming there is a sentence in the original voice: "At the beginning of the meeting, I felt completely unprepared, so I was late, which was really embarrassing.", where "which was really embarrassing" is slightly abrupt and logically slightly confusing, easily causing understanding barriers, the system can modify it to: "At the beginning of the meeting, I did not prepare well enough, resulting in missing the start time, which made me feel very embarrassed." by reorganizing the sentence structure, fixing the grammar and logic problems, while ensuring the fluency and readability of the sentence. In addition, the system will analyze the redundant parts of the sentence according to the cognitive model, simplifying lengthy or repetitive expressions, such as simplifying "I felt completely unprepared" to "I was not prepared", which not only improves the clarity of the language, but also makes the language expression more consistent with the human cognitive language processing mode. Through this simulated optimization, the generated serial processing voice data becomes more refined, natural in grammar, semantics, and logical structure, and can better convey the adaptation of emotions and scenes.

[0104] In this embodiment, through emotional reasoning, the system can repair the language in the voice for emotional expression according to the emotional state of the object, so that the voice is more consistent with the current emotional background of the object. For example, if the emotional state of the object is anxiety or happiness, the change in language will enhance the emotional color of the voice, making it more expressive. Scene reasoning repairs the dialogue theme and scene background of the voice, ensuring the semantic consistency and natural fluency of the language in the voice content in a specific environment and context. Finally, through simulated neural cognitive optimization, the linguistic features in the voice are further refined, improving the auditory effect and cognitive acceptability of the voice. Overall, through the dual repair of emotions and scenes, the final voice data is more vivid, immersive, and can accurately convey emotions and context, improving the interactive experience between the voice and the listener.

[0105] In one exemplary embodiment, as shown in Figure 7 the initial sound repair parameters are used as repair constraints, and the voice variation and semantic defects of the voice data to be processed are repaired in parallel according to the object voiceprint feature data and the object image data, to obtain parallel processing voice data, including steps 702 to 718. Among them:

[0106] Step 702, according to the object image data, the emotion of the voice output object is identified, and the object scene recognition data is obtained.

[0107] Specifically, the technical principle of identifying the object context recognition data from the object image data is consistent with step 506 for identifying the object image data, but step 506 finally obtains the object context recognition data which needs to be combined with the information in the scene feature data, but here is only obtained by identifying the object image data. Specifically, by analyzing the emotional signals (such as tone, speed, and pitch changes) in the facial expressions, eye contact, body language, and speech of the object in the object image data, the system can infer the emotional state of the object in the current conversation. For example, if the facial expression of the object shows tension or anxiety, the system may judge that the object is experiencing a tense or anxious emotional state, and vice versa, if the facial expression is relaxed and happy, the system may infer that the object feels relaxed or happy, and finally the object context recognition data is generated.

[0108] Step 704, context semantic recognition is performed on the to-be-processed sound data to obtain sound semantic recognition data.

[0109] Specifically, the technical principle of identifying the object context recognition data from the object image data is consistent with step 506 for identifying the object image data, but step 506 finally obtains the object context recognition data which needs to be combined with the information in the scene feature data, but here is only obtained by identifying the object image data. Specifically, by analyzing the emotional signals (such as tone, speed, and pitch changes) in the facial expressions, eye contact, body language, and speech of the object in the object image data, the system can infer the emotional state of the object in the current conversation. For example, if the facial expression of the object shows tension or anxiety, the system may judge that the object is experiencing a tense or anxious emotional state, and vice versa, if the facial expression is relaxed and happy, the system may infer that the object feels relaxed or happy, and finally the object context recognition data is generated.

[0110] Step 706, copy isolation is performed on the to-be-processed sound data to obtain isolated sound data.

[0111] Among them, the isolated sound data can be data obtained by analyzing and processing the to-be-processed sound data and copying the voice signal into multiple data.

[0112] Specifically, the system replicates the sound data to be processed and isolates it into multiple independent sound data isolation units, each containing different repair tasks, resulting in multiple isolated sound data. For example, the system may replicate the sound data to be processed into multiple identical data, each of which focuses on a different dimension of speech: for example, one segment only processes accent, another segment processes grammatical errors, and another segment processes emotional expression, etc.

[0113] Step 708: Adjust the pronunciation data of any selected isolated sound data according to the object voiceprint feature data, with the initial sound repair parameters as the repair constraints, to obtain accent-weakened sound data.

[0114] Specifically, the technical principle of adjusting the pronunciation data of the isolated sound data according to the object voiceprint feature data is consistent with the identification of the accent-weakened sound data in step 402 (the isolated sound data is the replicated sound data to be processed). Specifically, with the initial sound repair parameters as the repair constraints, the system will compare the object's voiceprint features (such as pitch, timbre, speech rate, etc.) with the pronunciation data of standard pronunciation in the database to identify the object's accent features and pronunciation deviations. According to these feature data, the system will apply specific algorithm models or audio signal processing techniques (such as multi-dimensional speech feature joint optimization identification algorithm, cross-domain speech feature conversion algorithm, speech synthesis and accent correction algorithm based on transformation learning, etc.) to gradually optimize and adjust the pronunciation part. For example, if the object has a certain local accent, the system will adjust the pronunciation mode to make it closer to the standard Mandarin pronunciation or the target pronunciation style. The process not only includes the correction of syllables, tones, and intonations, but also ensures that the adjusted speech remains natural and smooth according to the object's speech characteristics, avoiding making the speech too artificial or not conforming to the object's individual voiceprint features. Finally, through these adjustments, the system generates accent-weakened sound data, making the speech clearer, standard, and meeting the target pronunciation requirements.

[0115] Step 710: Repair the objective semantic defects of any unselected isolated sound data according to the sound semantic recognition data to obtain semantic repair sound data.

[0116] Specifically, by the same token, the technical principle and steps 504 of repairing the objective semantic defects of isolated sound data according to the sound semantic recognition data are consistent for repairing the objective semantic defects of accent-weakened sound data, except that the object to be repaired changes from accent-weakened sound data to isolated sound data (i.e. the sound data to be processed). Specifically, according to the previously extracted sound semantic recognition data, the objective semantic defects in the content of the isolated sound data are repaired, wherein the objective semantic defects include objectively existing grammatical errors, word usage that does not conform to conventional expression methods, logical confusion and other problems, i.e. the expression method adopted is not a normal expression method. For example, when the system finds that a certain word is used incorrectly in the sound semantic recognition data (for example, in the process of resource transaction, the object should use the word "buy", but the system recognizes that the object uses "take away"), or finds that a certain sentence has a grammatical error (such as lack of conjunctions leading to ambiguous sentence meaning), it will automatically correct it. Through natural language processing techniques (such as grammar correction, synonym replacement, etc.), the system can repair these semantic defects to generate semantic repair sound data, so that the voice content is more in line with standard grammar and the expression is more accurate, avoiding any errors that may cause ambiguity.

[0117] Step 712, according to the emotional reasoning data of the object scene recognition data, repairing the emotional expression of the voice of the sound data pronunciation data of any selected isolated sound data, obtaining emotional enhancement recognition data.

[0118] Specifically, by the same logic, the technical principle and steps 602 of repairing the speech emotional expression of isolated sound data according to the emotional inference data of the object scene recognition data are consistent with repairing the speech emotional expression of semantic repair sound data, except that the repaired object changes from the speech emotional expression of semantic repair sound data to the speech emotional expression of isolated sound data. Specifically, based on the emotional inference data in the object scene recognition data, the problem of insufficient emotional expression in the voice is repaired. Since the emotional inference data is obtained by analyzing the emotional features in the object's facial expressions, body language, voice tone and speech content, it helps the system to understand the current emotional state of the object (such as joy, anger, sadness, etc.). The system adjusts the isolated sound data using the emotional inference data. Specifically, the system will evaluate the parameters such as word choice, sentence expression, paragraph expression, overall structure in the voice, and optimize it according to the inferred emotional inference data, that is, combine these emotional inference data with the isolated sound data to improve the parts of the voice that are insufficient or incorrect in emotional expression. For example, the original sentence may be written as "I feel a little uncomfortable", but this expression is relatively flat and cannot reflect the emotional intensity of the object. The system will modify the sentence to "I feel very uneasy in my heart, almost to the point of collapse", and through more intense emotional vocabulary and tone correction, the intensity and accuracy of emotional expression are improved. Finally, the system outputs emotional enhancement recognition data, which makes the emotional expression of the voice more accurately aligned with the emotional state of the object, and reflects more rich emotional colors in the voice.

[0119] Step 714, according to the scene inference data of the object scene recognition data, the dialogue theme scene of any isolated sound data pronunciation data is repaired to obtain scene enhancement recognition data.

[0120] Specifically, the technical principle and steps of repairing the dialogue theme scene of isolated sound data according to the scene reasoning data of object context recognition data are consistent with those of repairing the dialogue theme scene of emotion enhancement recognition data, except that the repaired object changes from the speech emotional expression of emotion enhancement recognition data to the speech emotional expression of isolated sound data. Specifically, the system further repairs the speech according to the scene reasoning data of object context recognition data to ensure that it conforms to the actual scene of the dialogue. The scene reasoning data is inferred by analyzing the environmental background, context and content of the dialogue of the object, for example, if the system detects that the object is in a formal meeting, the expression in the speech should be more precise and standard; if the object is in a casual family gathering, the tone and words should be more casual and casual. The system first determines from the isolated sound data whether these contents are consistent with the current scene, and repairs the words, sentences, paragraphs or structures that are not consistent with the current scene; assuming that in a business meeting scene, the original sentence may express "we can consider this matter later", which is too casual and not precise enough in a formal setting. The system determines through scene reasoning and recommends modifying it to "we can further discuss and make decisions in the later meeting". Similarly, if the dialogue is in a casual gathering, the system may suggest adjusting the tone to "let's talk about it later, don't rush, let's discuss it together". This repair ensures the consistency of scene enhancement recognition data in content and tone, ensuring that the language can express emotion and match the background and social occasion of the dialogue.

[0121] Step 716, fuse the accent weakening sound data, semantic repair sound data, emotion enhancement recognition data and scene enhancement recognition data to obtain fused sound repair data.

[0122] The fused sound repair data can be the integration of sound data processed by different steps to generate the final repair result. In the process of voice repair, the system may perform multiple processing on different aspects of the same voice, such as weakening the accent in the voice, repairing the semantics, enhancing the emotional expression, etc. These processing steps generate different data sets, such as "accent weakening sound data", "semantic repair sound data" and "emotion enhancement recognition data", etc. The process of fusing sound repair data is to combine these independently processed data to ensure their continuity in the time axis and voice stream, while retaining the optimization effect brought by each repair step.

[0123] Specifically, the accent weakening sound data, the semantic repair sound data, the emotion enhancement recognition data, and the scene enhancement recognition data are merged into a unified audio file. This fusion process can use weighted averaging, feature splicing, or other audio synthesis algorithms to combine the results of each repair module, ensuring that the repair in each dimension is presented in the final speech, obtaining the fused sound repair data.

[0124] At step 718, the linguistic features of the fused sound repair data are simulated and neuro-cognitively optimized to obtain parallel processing sound data.

[0125] Specifically, the technical principle of simulating and neuro-cognitively optimizing the linguistic features of the fused sound repair data is consistent with that of simulating and neuro-cognitively optimizing the linguistic features of the scene enhancement recognition data at step 606, except that the object being repaired changes from the linguistic features of the scene enhancement recognition data to the linguistic features of the fused sound repair data. Specifically, the system performs more in-depth linguistic optimization on the fused data after emotion and scene repair, focusing on correcting deficiencies in language structure and improving grammatical correctness, fluency, and logicality. Through the simulated neuro-cognitive optimization technology, the system simulates how the human brain processes language, adjusts grammatical errors, unclear expressions, or logically incoherent parts in the voice data. Assuming that the original voice has a sentence: "At the beginning of the meeting, I felt completely unprepared, so I was late, and the situation was really embarrassing.", the expression "the situation is really embarrassing" is slightly abrupt, and the logic is slightly confusing, which can easily cause understanding barriers. The system can modify this sentence to: "At the beginning of the meeting, I did not prepare well enough, resulting in missing the start time, which made me feel very embarrassed.", by reorganizing the sentence structure, fixing the grammar and logic problems, and ensuring the fluency and readability of the sentence. In addition, the system will analyze the redundant parts of the sentence according to the cognitive model, simplify long or repetitive expressions, such as simplifying "I feel completely unprepared" to "I am not prepared", which not only improves the clarity of the language, but also makes the language expression more consistent with the human cognitive language processing mode. Through this simulated optimization, the generated parallel processing sound data becomes more refined, natural, and better conveys the adaptation of emotions and scenes in terms of grammar, semantics, and logical structure.

[0126] In this embodiment, through emotion recognition and context semantic recognition based on object image data, the emotional state and context of the sound output object can be accurately grasped, ensuring that the subsequent repair operation can be consistent with the real emotions and context of the object. Replication isolation and targeted repair independently optimize pronunciation, semantics and emotional expression, solving the problems of accent, unclear semantics and insufficient emotional expression, and improving the naturalness and accuracy of speech expression. Especially in terms of voice emotional expression and scene reasoning, the system can repair the emotional deficiency and context inconsistency of language expression in the voice according to the emotion reasoning and scene reasoning data, making the voice more expressive in emotion and more natural in sentence expression. Finally, the simulation of neural cognitive optimization for linguistic features further improves the sound hearing effect and cognitive acceptance. Overall, the repair process not only improves the clarity and semantic accuracy of the sound, but also better conveys the emotions and intentions of the object, significantly enhancing the interactive experience between the sound and the listener.

[0127] In one exemplary embodiment, as shown in Figure 8 According to the processed sound data, the initial sound repair parameters are adjusted to obtain adjusted sound repair parameters, including steps 802 to 806. Among them:

[0128] Step 802, determining the speech data adjustment rule corresponding to the voice variation and the semantic data adjustment rule corresponding to the semantic defect from the processed sound data.

[0129] Among them, the speech data adjustment rule can be a series of adjustment specifications and strategies followed when repairing the voice variation (such as accent, unclear pronunciation, too fast speech, etc.) in the sound data again. These rules determine how to improve or optimize the output of the voice through the analysis of voiceprint features, pronunciation patterns, speech quality, etc. For example, if the system detects that the pronunciation of a certain syllable is unclear, the adjustment rule may include increasing the volume of that syllable, changing the pitch or speed, or even adjusting the pronunciation method to ensure that the syllable can be clearly understood.

[0130] Among them, the semantic data adjustment rule can be a set of specifications and methods for repairing semantic defects (such as sentence disorder, inappropriate word use or grammatical errors, etc.) in the voice. These rules repair and optimize through the analysis of sentence structure, grammatical consistency, accuracy of vocabulary, and clarity of expression. For example, if the words in a sentence are used improperly or expressed ambiguously, the semantic adjustment rule will suggest replacing them with more accurate and appropriate words; if the sentence structure is complex or does not conform to grammatical norms, the rule may suggest simplifying the sentence or reorganizing the word order.

[0131] Specifically, since the processed sound data still has speech variation and / or semantic defects after repair is detected, the system detects the accent or pronunciation pattern in the processed sound data through voiceprint analysis technology, further determines which pronunciation features still need to be adjusted, such as whether the pronunciation is unclear after some syllable features are adjusted, or whether the standardization of speech is affected after the characteristics of the accent of a certain region are adjusted. Then the semantic defect analysis will further identify the unclear expression of the sentence in the processed sound data, such as the problem of inaccurate vocabulary, incoherent sentence structure, or insufficient emotional expression. For example, a sentence may have ambiguous semantics due to unclear grammatical structure, or the listener cannot accurately understand the speaker's intention due to insufficient emotional expression. The system will set corresponding repair rules for these problems that are more stringent than before, or increase the sensitivity of the repair rules, where the speech data adjustment rules are adjusted by adjusting the accuracy of pronunciation, the clarity of speech, etc., and the semantic data adjustment rules include modifying vocabulary, sentence structure, and grammar to ensure the accuracy and clarity of semantics.

[0132] Step 804, identifying the adjustment strategy correlation information between the speech data adjustment rules and the semantic data adjustment rules.

[0133] Wherein, the adjustment strategy correlation information can be the mutual relationship and influence between the repaired speech data and semantic data. These information help the system determine how to coordinate the speech and semantic repair to avoid conflict between the two repair processes. For example, when repairing pronunciation in speech, it may affect the original semantic expression, especially when non-standard pronunciation causes vocabulary confusion; similarly, modifying semantics may affect the way sentences are pronounced and the tone.

[0134] Specifically, the system needs to identify the interaction and influence between the speech and semantic adjustment rules. For example, accent variation in speech may not only affect the clarity of pronunciation, but also cause semantic misunderstanding, especially in homophones, synonyms, and sentence structure. The system needs to identify the correlation between these adjustment rules to ensure that the repair of speech and semantics can be complementary and coordinated. If the pronunciation problem of speech is repaired, but the semantic expression is not corrected accordingly, it may cause deviation in the listener's understanding. Therefore, the system will use the interaction between the clarity of speech and the accuracy of semantics as part of the adjustment strategy to make corrections. For example, if there is an accent problem in speech, it may cause some vocabulary misunderstanding, so speech repair not only needs to adjust pronunciation, but also needs to combine semantic repair rules to ensure that the repaired speech does not have ambiguity in meaning.

[0135] The implementation is to identify the relationship between the speech variation (such as accent, unclear pronunciation, etc.) and the semantic defect (such as grammatical error, ambiguous expression, etc.) through in-depth analysis of the processed sound data. For example, when the accent in the speech causes the pronunciation of certain words to be non-standard, it may cause semantic misunderstanding or ambiguity. The system will evaluate this mutual relationship and consider this coupling of speech and semantics when using speech data adjustment rules and semantic data adjustment rules. The adjustment strategy correlation information is based on this evaluation result, that is, when using speech data adjustment rules and semantic data adjustment rules, the repair of speech and semantics is considered to ensure that there is no conflict when repairing both. For example, if the pronunciation repair may cause changes in semantic expression, the system will adjust the semantic repair rules accordingly to avoid logical inconsistency. By analyzing the relationship between these rules, the adjustment strategy correlation information between the speech data adjustment rules and the semantic data adjustment rules is obtained.

[0136] Step 806, adjusting the initial sound repair parameters according to the speech data adjustment rules, the semantic data adjustment rules, and the adjustment strategy correlation information to obtain adjusted sound repair parameters.

[0137] Specifically, the system will integrate the speech data adjustment rules, the semantic data adjustment rules, and the correlation information between them to optimize and adjust the initial sound repair parameters. Since the initial repair parameters usually include speech-related model repair parameters (such as model repair parameters for tone, speech rate, pitch, volume, pronunciation clarity, etc.) and semantic-related model repair parameters (such as model repair parameters for sentence patterns, accuracy of vocabulary use, grammatical structure, etc.), according to the restrictions of the speech and semantic data adjustment rules obtained in the first two steps, combined with the adjustment strategy correlation information of both, the system will adjust these model parameters at the same time. For example, if the system detects that the accent problem in the speech affects the expression of semantics, then when adjusting the pronunciation clarity as the target and adjusting the parameters of the audio processing model that meets the adjustment rules, it may also need to adjust the grammatical structure at the semantic level as the adjustment target and adjust the parameters of the audio processing model that meets the adjustment rules to ensure that the repaired speech is clear and meets the grammatical specifications. In this process, the system will continuously iterate and optimize through algorithms until the most suitable adjusted sound repair parameters are obtained to achieve the maximum repair effect of speech variation and semantic defects.

[0138] In this embodiment, by accurately analyzing the speech variation and semantic defects in the processed sound data, the corresponding adjustment rules can be effectively determined, thereby providing clear guidance for sound repair. In this process, the system not only identifies the speech data and semantic data adjustment rules, but also analyzes the adjustment strategy correlation between the two, ensuring that the accuracy of the semantics or the integrity of the emotional expression is not damaged when repairing the speech variation. This adjustment strategy correlation analysis provides a global optimization perspective for sound repair, making the repair process more coordinated and efficient. By adjusting the speech and semantic parameters comprehensively, the system can improve the clarity and accuracy of expression on the basis of ensuring the natural and smooth flow of speech, ultimately obtaining a repaired sound data that is more consistent with the object's real situation, emotions and context, significantly improving the quality of the repair result and the listener's experience.

[0139] In an exemplary embodiment, as shown in Figure 9 According to the speech data adjustment rule, the semantic data adjustment rule, and the adjustment strategy correlation information, the initial sound repair parameters are adjusted to obtain adjusted sound repair parameters, including steps 902 to 906. Among them:

[0140] Step 902, according to the speech data adjustment rule to determine the speech parameter adjustment initial value, and according to the semantic data adjustment rule to determine the semantic parameter adjustment initial value.

[0141] Among them, the speech parameter adjustment initial value can be a preliminary adjustment value of a parameter set based on the analysis of the processed sound data. These parameter values are derived from the speech data adjustment rule, such as adjusting the initial parameters of the pitch, speech rate, pronunciation clarity, and sound quality.

[0142] Among them, the semantic parameter adjustment initial value can be a preliminary adjustment value of a parameter corresponding to a preliminary adjustment scheme set for the vocabulary, syntactic structure, or syntax error in the sentence in the repair process based on the analysis of the processed sound data combined with the semantic data adjustment rule.

[0143] Specifically, the system sets initial adjustment values for voice and semantic repair parameters according to defined voice data adjustment rules and semantic data adjustment rules. In the specific implementation process, the voice data adjustment rules will be based on the analysis of the characteristics of the processed sound data, such as pronunciation mode, voice clarity, accent, intonation, pitch, and speech rate, and further determine the specific initial parameters of the model for implementing the repair scheme of voice variation, i.e. voice parameter adjustment initial value; for example, for some unclear pronunciation, the rules may suggest making it clearer by increasing the volume or adjusting the voice frequency; and for the accent or dialect in the voice, rules may be set to weaken these accents so that the listener can understand more easily. Secondly, the semantic data adjustment rules will be based on the analysis of the semantic defects of the processed sound data (such as inappropriate vocabulary, grammatical errors, and incoherent sentences), to set the preliminary direction of semantic repair; for example, if a sentence has grammatical errors, the rules may suggest restructuring the sentence; if the statement is ambiguous, the rules may suggest replacing it with more precise words. Finally, according to the preliminary direction of semantic repair, the specific initial parameters of the model for implementing the repair scheme of semantic defect repair are determined, i.e. semantic parameter adjustment initial value.

[0144] Step 904, according to the adjustment strategy correlation information, jointly optimize the voice parameter adjustment initial value and the semantic parameter adjustment initial value respectively with the voice data adjustment rules and the semantic data adjustment rules as constraint conditions, to obtain the voice parameter adjustment determined value and the semantic parameter adjustment determined value.

[0145] Among them, the voice parameter adjustment determined value can be the final model parameter value for adjusting voice variation after joint optimization and verification on the basis of the preliminary adjustment value of voice variation repair.

[0146] Among them, the semantic parameter adjustment determined value can be the final semantic repair parameter value for adjusting semantic defects after joint optimization and verification on the basis of the preliminary adjustment value of semantic defect repair.

[0147] Specifically, the system will adjust the voice parameter adjustment initial value and the semantic parameter adjustment initial value through joint optimization according to the voice data adjustment rules and the semantic data adjustment rules, so the system will not only focus on voice repair (such as pronunciation, voice quality), but also repair semantic problems (such as grammar, vocabulary) at the same time, and ensure that the repair between the two can be coordinated. For example, if the voice needs to be adjusted when repairing, the adjustment of the pronunciation frequency or the intonation may affect the semantic understanding, especially when the voice intonation is too high or too low, which may cause ambiguity in emotional or semantic transmission.

[0148] In the process of joint optimization, the system analyzes the correlation information of the adjustment strategies under the constraints of the voice data adjustment rules and the semantic data adjustment rules, and evaluates whether each adjustment will cause inconsistency or conflict. For example, the adjustment of voice parameters (such as pitch, speed, volume, etc.) may affect the emotional expression and semantic delivery of the sentence. If the pitch of the voice is too high or too low, it may change the semantics of the sentence or cause misunderstanding, so the system must ensure that the voice adjustment does not affect the clarity of the semantics. Semantic adjustment (such as modifying vocabulary, adjusting syntax) may change the fluency or structure of the language, which may also affect the voice performance, such as adjusting the speed or accent position; further calculate and combine the coupling between the adjustments of various parameters, and develop an optimization scheme through optimization algorithm, so as to realize the dual optimization of voice and semantics. Through this joint optimization system can determine the optimal voice parameter adjustment determination value and semantic parameter adjustment determination value.

[0149] Step 906, using standard verification sound data, the voice parameter adjustment determination value and the semantic parameter adjustment determination value are checked for compliance and robustness, and the parameter adjustment value verification result is obtained.

[0150] Specifically, the voice parameter adjustment determination value and the semantic parameter adjustment determination value are checked for compliance and robustness using standard verification sound data. Among them, the compliance check mainly focuses on whether the voice parameter adjustment determination value and the semantic parameter adjustment determination value meet certain parameter standards, such as whether the parameters are negatively optimized to produce voice distortion or sound quality decline, and whether the parameter expression is accurate and conforms to the syntax rules. Robustness check tests the performance of voice parameter adjustment determination value and semantic parameter adjustment determination value in different environments and contexts, such as whether it can still ensure clear pronunciation in noisy environments, or whether it can correctly convey the intention in various contexts. The system will conduct a comprehensive evaluation of the compliance and robustness of the adjusted voice and semantic parameters by simulating different actual application scenarios, to ensure that they can maintain good performance in a wide range of use cases. Through this process, the system can obtain the parameter adjustment value verification result of the parameter adjustment value.

[0151] Step 908, in the case where the parameter adjustment value verification result is not abnormal, the initial sound repair parameter is adjusted according to the voice parameter adjustment determination value and the semantic parameter adjustment determination value, and the adjusted sound repair parameter is obtained.

[0152] Specifically, in the case where neither the parameter adjustment value verification result compliance check nor the robustness check finds abnormalities, the system will adjust the initial sound repair parameters according to the final voice parameter adjustment determination value and semantic parameter adjustment determination value. Because if the verification result shows that the adjusted parameters can effectively solve the problems of voice variation and semantic defects and do not produce side effects, the system applies these adjustment values to the actual initial repair parameters, updates the repair scheme, and obtains adjusted sound repair parameters. At this time, the initial repair parameters have been optimized to the adjusted sound repair parameters, ensuring that more accurate and natural sound repair effects can be provided in subsequent processing.

[0153] In this embodiment, by accurately analyzing the adjustment rules of voice data and semantic data, the initial adjustment values of voice and semantic parameters are determined, laying a foundation for subsequent optimization. Based on the correlation information of the adjustment strategy, the system jointly optimizes the initial adjustment values of voice and semantics, ensuring that the two remain coordinated in the repair process and avoiding the influence of single-dimensional adjustment on the overall repair effect. Further, the standard verification sound data is used for compliance and robustness checks, and the system can effectively verify whether the adjusted parameters meet the predetermined repair standards and ensure the stability and consistency of the repair results in different environments. Finally, in the case where the verification result is normal, the repair parameters are adjusted, and the initial sound repair parameters are accurately optimized. More accurate and stable sound repair can be achieved, the quality of the repair results is improved, and the sound in terms of emotion, context, and semantics is optimized to achieve the best effect, enhancing the naturalness and clarity of the auditory experience.

[0154] Based on the same inventive concept, the embodiments of the present application also provide an artificial intelligence-based audio data processing device for implementing the above-mentioned artificial intelligence-based audio data processing method. The implementation scheme of the device for solving the problem is similar to the implementation scheme described in the above method, so the specific limitations in one or more artificial intelligence-based audio data processing device embodiments provided below can refer to the limitations of the artificial intelligence-based audio data processing method in the above text, which will not be repeated here.

[0155] In one exemplary embodiment, as shown in Figure 10 An artificial intelligence-based audio data processing device is provided, which includes an audio data acquisition module 1002, a voiceprint feature extraction module 1004, an audio data repair module 1006, a repair parameter adjustment module 1008, and an audio data determination module 1010.

[0156] In one exemplary embodiment, a computer, which can be a server, is provided, and its internal structure diagram can be as shown in Figure 11The computer includes a memory and a processor, the memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program. A computer readable storage medium is also provided, which stores a computer program. The computer program is executed by the processor to implement the steps in the above method embodiments. A computer program product or computer program is also provided, which includes computer instructions stored in a computer readable storage medium. The processor of the computer reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer execute the steps in the above method embodiments.

Claims

1. An artificial intelligence-based audio data processing method, characterized by, The method comprises: acquiring to-be-processed sound data of a sound output object, object image data, and initial sound repair parameters; extracting voiceprint features from the to-be-processed sound data to obtain object voiceprint feature data; repairing speech variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data, with the initial sound repair parameters as repair constraints, to obtain processed sound data; wherein in the case of parallel repair of the repairing manner; identifying the emotion of the sound output object according to the object image data to obtain object scene identification data; performing context semantic identification on the to-be-processed sound data to obtain sound semantic identification data; performing copy isolation on the to-be-processed sound data to obtain isolated sound data; adjusting pronunciation data of any selected isolated sound data according to the object voiceprint feature data, with the initial sound repair parameters as repair constraints, to obtain accent-weakened sound data; and repairing objective semantic defects of any unselected isolated sound data according to the sound semantic identification data, to obtain semantic-repaired sound data; and repairing speech emotional expression of pronunciation data of any unselected isolated sound data according to emotional reasoning data of the object scene identification data, to obtain emotion-enhanced identification data; and repairing dialogue theme scenes of pronunciation data of any isolated sound data according to scene reasoning data of the object scene identification data, to obtain scene-enhanced identification data; fusing the accent-weakened sound data, the semantic-repaired sound data, the emotion-enhanced identification data, and the scene-enhanced identification data to obtain fused sound repair data; performing simulated neural cognitive optimization on linguistic features of the fused sound repair data to obtain parallel-processed sound data; in the case of detecting that the processed sound data still has the speech variation and / or the semantic defects, adjusting the initial sound repair parameters according to the processed sound data to obtain adjusted sound repair parameters; taking the adjusted sound repair parameters as the initial sound repair parameters, returning to perform the step of repairing speech variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data, with the initial sound repair parameters as repair constraints, to obtain processed sound data; until the case where the processed sound data no longer has the speech variation and the semantic defects is detected, taking the processed sound data as target sound data.

2. The method of claim 1, wherein, the step of repairing speech variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data, with the initial sound repair parameters as repair constraints, to obtain processed sound data, comprises: Using the initial sound restoration parameters as restoration constraints, the speech variations and semantic defects of the sound data to be processed are sequentially restored according to the object's voiceprint feature data and the object's image data to obtain serially processed sound data. Using the initial sound restoration parameters as restoration constraints, the speech variations and semantic defects of the sound data to be processed are simultaneously restored in parallel based on the object's voiceprint feature data and the object's image data, resulting in parallel processed sound data. The serially processed audio data and the parallel processed audio data are fused across pipelines to obtain the processed audio data.

3. The method of claim 2, wherein, The process involves using the initial sound restoration parameters as restoration constraints, and sequentially restoring the speech variations and semantic defects of the sound data to be processed based on the object's voiceprint feature data and the object's image data, to obtain serially processed sound data, including: Using the initial sound restoration parameters as restoration constraints, the pronunciation data of the sound data to be processed is adjusted according to the voiceprint feature data of the object to obtain accent-weakened sound data. Based on the object image data and the accent-weakened sound data, the generation environment of the sound data to be processed is inferred to obtain scene feature data; Based on the scene feature data, the language representation data of the accent-weakened sound data is modified to obtain the serially processed sound data.

4. The method of claim 3, wherein, The step of modifying the language representation data of the accent-weakened sound data based on the scene feature data to obtain the serially processed sound data includes: Contextual semantic recognition is performed on the weakened accent sound data to obtain sound semantic recognition data; Based on the sound semantic recognition data, the objective semantic defects of the accent-weakened sound data are repaired to obtain semantically repaired sound data. Based on the scene feature data and the object image data, the emotion of the sound output object is identified to obtain object scene recognition data; Based on the object scene recognition data, the scene semantic defects of the semantic repair sound data are repaired to obtain the serially processed sound data.

5. The method of claim 4, wherein, The step of repairing scene semantic defects in the semantically repaired sound data based on the object scene recognition data to obtain the serially processed sound data includes: Based on the sentiment inference data of the object scene recognition data, the speech sentiment expression of the semantic repair sound data is repaired to obtain sentiment enhancement recognition data; Based on the scene reasoning data of the object scene recognition data, the dialogue theme scene of the emotion enhancement recognition data is repaired to obtain scene enhancement recognition data; The linguistic features of the scene enhancement recognition data are optimized using simulated neurocognitive methods to obtain the serially processed sound data.

6. The method of claim 1, wherein, The step of adjusting the initial sound restoration parameters based on the processed sound data to obtain adjusted sound restoration parameters includes: Determine the speech data adjustment rules corresponding to the speech variations and the semantic data adjustment rules corresponding to the semantic defects from the processed sound data; identify adjustment strategy correlation information between the voice data adjustment rule and the semantic data adjustment rule; adjust the initial sound restoration parameter according to the voice data adjustment rule, the semantic data adjustment rule, and the adjustment strategy correlation information, to obtain the adjusted sound restoration parameter.

7. The method of claim 6, wherein, The adjusting the initial sound restoration parameter according to the voice data adjustment rule, the semantic data adjustment rule, and the adjustment strategy correlation information, to obtain the adjusted sound restoration parameter, comprises: determining a voice parameter adjustment initial value according to the voice data adjustment rule, and determining a semantic parameter adjustment initial value according to the semantic data adjustment rule; using the voice data adjustment rule and the semantic data adjustment rule as constraint conditions, performing joint optimization on the voice parameter adjustment initial value and the semantic parameter adjustment initial value according to the adjustment strategy correlation information, to obtain a voice parameter adjustment determined value and a semantic parameter adjustment determined value; performing compliance checking and robustness checking on the voice parameter adjustment determined value and the semantic parameter adjustment determined value using standard verification sound data, to obtain a parameter adjustment value verification result; in a case where the parameter adjustment value verification result is not abnormal, adjusting the initial sound restoration parameter according to the voice parameter adjustment determined value and the semantic parameter adjustment determined value, to obtain the adjusted sound restoration parameter.

8. An artificial intelligence-based audio data processing apparatus, characterized by comprising: The device comprises: an audio data acquisition module configured to acquire to-be-processed sound data of a sound output object, object image data, and an initial sound restoration parameter; a voiceprint feature extraction module configured to perform voiceprint feature extraction on the to-be-processed sound data, to obtain object voiceprint feature data; an audio data restoration module configured to use the initial sound restoration parameter as a restoration constraint condition, and perform restoration on voice variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data, to obtain processed sound data; in a case where the restoration is performed in parallel; identify the emotion of the sound output object according to the object image data, to obtain object scene recognition data; perform context semantic recognition on the to-be-processed sound data, to obtain sound semantic recognition data; perform copy isolation on the to-be-processed sound data, to obtain isolated sound data; use the initial sound restoration parameter as a restoration constraint condition, and adjust pronunciation data of any selected isolated sound data according to the object voiceprint feature data, to obtain accent-weakened sound data; and, according to the sound semantic recognition data, perform restoration on objective semantic defects of any unselected isolated sound data, to obtain semantic restoration sound data; and, according to emotion reasoning data of the object scene recognition data, perform restoration on voice emotion expression of pronunciation data of any unselected isolated sound data, to obtain emotion-enhanced recognition data; And, according to the object scene recognition data, scene inference data, the dialogue theme scene of any of the isolated sound data pronunciation data is repaired, and scene enhanced recognition data is obtained; The accent weakening sound data, the semantic repair sound data, the emotion enhancement recognition data and the scene enhancement recognition data are fused to obtain fused sound repair data; The linguistic features of the fused sound repair data are simulated and neuro-cognitive optimized to obtain parallel processing sound data; The repair parameter adjustment module is configured to, in a case where it is detected that the processed sound data still has the speech variation and / or the semantic defect, adjust the initial sound repair parameter according to the processed sound data to obtain an adjusted sound repair parameter; The audio data repair module is further configured to return to execute the step of taking the initial sound repair parameter as a repair constraint condition, repairing the speech variation and the semantic defect of the to-be-processed sound data according to the object voiceprint feature data and the object image data to obtain processed sound data, by taking the adjusted sound repair parameter as the initial sound repair parameter. The audio data determination module is configured to, until it is detected that the processed sound data does not have the speech variation and the semantic defect, take the processed sound data as target sound data.

9. A computer comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor implements the steps of the method of any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Voice quality inspection method and device, terminal and computer readable storage medium

    CN111210842A

  • Chinese speech enhancement recognition and text error correction and correction method

    CN115602161A