Audio data processing method and device based on artificial intelligence and computer
Through the audio data processing method based on artificial intelligence, combined with voiceprint feature extraction and image data, voice variation and semantic defects are repaired in real time, solving the problem of inefficient repair in the existing technology and achieving a more efficient voice repair effect.
Patent Information
- Application Number
- CN202510152179.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The prior art is difficult to effectively repair audio data with speech variations and semantic defects in real time, resulting in ineffective repair.
The audio data processing method based on artificial intelligence is adopted to repair voice variation and semantic defects through combining voiceprint feature extraction and image data, and to optimize the repair effect through iterative adjustment of the initial repair parameters.
It significantly improves the real-time repair efficiency of audio data with speech variations and semantic defects, improves the performance of speech processing systems, and provides higher clarity, naturalness and accuracy of speech output.
Smart Images

Figure CN119993206A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an audio data processing method, device and computer based on artificial intelligence. Background Art
[0002] In traditional technology, when audio data is damaged or missing, the signal can be compensated through a method based on spectrum analysis, such as using Fourier transform to extract frequency domain features, and then using interpolation algorithms or least squares methods to restore the lost parts. In addition, deep learning technology has also been widely used in audio restoration, by training neural network models (such as convolutional neural networks or generative adversarial networks) to automatically identify and fill in missing or damaged parts in the audio. This method can restore natural audio data without audible signs of restoration while maintaining audio quality. However, in the traditional technology for processing audio data, no effective real-time processing method has been proposed for speech variations (such as accents) and semantic defects (such as grammatical errors), resulting in low efficiency in real-time repair of audio data with speech variations and semantic defects. Summary of the invention
[0003] Based on this, it is necessary to provide an audio data processing method, device and computer based on artificial intelligence that can effectively improve the efficiency of real-time repair of audio data with speech variations and semantic defects in response to the above technical problems.
[0004] In a first aspect, the present application provides an audio data processing method based on artificial intelligence, comprising: obtaining to-be-processed sound data, object image data and initial sound repair parameters of a sound output object; performing voiceprint feature extraction on the to-be-processed sound data to obtain object voiceprint feature data; using the initial sound repair parameters as repair constraints, and repairing the voice variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data to obtain processed sound data; when it is detected that the processed sound data still has the voice variation and / or the semantic defect, adjusting the initial sound repair parameters according to the processed sound data to obtain adjusted sound repair parameters; using the adjusted sound repair parameters as the initial sound repair parameters, returning to execute the step of using the initial sound repair parameters as repair constraints, and repairing the voice variation and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and the object image data to obtain processed sound data; until the processed sound data can no longer be detected to have the voice variation and the semantic defect, the processed sound data is used as the target sound data.
[0005] In a second aspect, the present application also provides an audio data processing device based on artificial intelligence, including: an audio data acquisition module, used to obtain the to-be-processed sound data, object image data and initial sound repair parameters of the sound output object; a voiceprint feature extraction module, used to extract the voiceprint features of the to-be-processed sound data to obtain the object voiceprint feature data; an audio data repair module, used to use the initial sound repair parameters as repair constraints, and repair the voice variations and semantic defects of the to-be-processed sound data according to the object voiceprint feature data and object image data to obtain processed sound data; a repair parameter adjustment module, used to repair the voice variations and semantic defects of the to-be-processed sound data when it is detected that the processed sound data still has the voice variations or / and the semantic defects, according to the processed sound data, the initial sound repair parameters are adjusted to obtain adjusted sound repair parameters; the audio data repair module is further used to use the adjusted sound repair parameters as the initial sound repair parameters, return to execute the step of using the initial sound repair parameters as repair constraints, and repairing the voice variation and semantic defects of the sound data to be processed according to the object voiceprint feature data and the object image data to obtain the processed sound data; the audio data determination module is used to use the processed sound data as the target sound data until the voice variation and the semantic defects in the processed sound data cannot be detected.
[0006] In a third aspect, the present application also provides a computer, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements any step of the audio data processing method based on artificial intelligence when executing the computer program.
[0007] The above-mentioned audio data processing method, device and computer based on artificial intelligence obtain the personalized acoustic features of the object through voiceprint feature extraction technology, and repair the voice variations and semantic defects of the sound data to be processed based on these features and combined with image data. During the repair process, if it is detected that the sound data still has voice variations and / or semantic defects, the system will iteratively adjust the initial repair parameters to further optimize the repair effect. After multiple rounds of repair and parameter adjustment, until the voice variations and semantic defects in the sound data are completely eliminated, the target sound data is finally output; it can effectively improve the real-time repair efficiency of audio data with voice variations and semantic defects, and further significantly improve the performance of the speech processing system, especially in applications such as speech recognition, speech synthesis and voice communication, providing speech output with higher clarity, naturalness and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0009] Figure 1 is an application environment diagram of an audio data processing method based on artificial intelligence in one embodiment;
[0010] Figure 2 1 is a flow chart of an audio data processing method based on artificial intelligence in one embodiment;
[0011] Figure 3 A schematic diagram of a flow chart of a method for obtaining processed sound data in one embodiment;
[0012] Figure 4 It is a flowchart of a first method for serially processing sound data in one embodiment;
[0013] Figure 5 It is a flowchart of a second method for serially processing sound data in one embodiment;
[0014] Figure 6 It is a flowchart of a third method for serially processing sound data in one embodiment;
[0015] Figure 7 A schematic flow chart of a method for processing sound data in parallel in one embodiment;
[0016] Figure 8 A schematic flow chart of a first method for adjusting sound restoration parameters in an embodiment;
[0017] Fig. 9 It is a flowchart diagram of a second method for adjusting sound restoration parameters in one embodiment;
[0018] Fig.10 is a structural block diagram of an audio data processing device based on artificial intelligence in one embodiment;
[0019] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0021] The audio data processing method based on artificial intelligence provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the server 104 can be implemented with an independent server or a server cluster composed of multiple servers.
[0022] In an exemplary embodiment, Figure 2 As shown, an audio data processing method based on artificial intelligence is provided, and the method is applied to Figure 1 The server in the example is used to illustrate, including the following steps 202 to 212. Among them:
[0023] Step 202: Acquire the to-be-processed sound data, the object image data and the initial sound restoration parameters of the sound output object.
[0024] The sound output object may be an entity or individual that emits audio data. It may be a specific person, device, virtual character, or any object that can generate voice output.
[0025] The sound data to be processed may be original voice data that has not been repaired, and usually includes information on the voice input, environmental noise, speaking speed, intonation, etc. of the sound output object.
[0026] The object image data may be visual information about the sound output object acquired by a camera or other device, including facial expressions, body language, environmental background, etc.
[0027] The initial sound restoration parameters may be the restoration constraints and adjustment rules set when the restoration process begins. These parameters may include the initial goal of audio restoration, the scope of restoration, the correction criteria for speech features, and the like.
[0028] Specifically, the system obtains the unprocessed sound data of the sound output object, which is usually collected from the actual environment and may have problems such as voice variation, noise or semantic ambiguity. At the same time, the system also obtains object image data related to the sound output object, such as facial images of people or other characteristic image data. These images help to understand and judge the characteristic information of the object, such as facial expressions, mouth shapes, emotions, scenes, etc. when speaking. In addition to these data, the system also requires an initial sound restoration parameter, which usually includes the basic settings of the sound restoration algorithm, such as restoration strength, frequency range, speech model, etc.
[0029] Step 204: extract voiceprint features from the sound data to be processed to obtain object voiceprint feature data.
[0030] Among them, voiceprint feature extraction can be to extract individual voice physiological characteristics by analyzing the frequency, timbre, pitch, voice pattern and other information in the sound data. Each person's voiceprint is unique. Voiceprint feature extraction technology can help the system identify and establish a personalized voice model for the object, providing an accurate reference for subsequent sound restoration, especially when repairing voice variations such as accents and pronunciations.
[0031] The object voiceprint feature data can be the voice feature data of the sound output object obtained through voiceprint feature extraction technology. It includes the individual's timbre, pitch, pronunciation habits, etc., which can reflect the unique voice characteristics of the object. During the restoration process, these voiceprint data will be used as a reference to ensure that the restored sound can accurately reflect the personalized voice characteristics of the object and retain its natural voice attributes as much as possible.
[0032] Specifically, voiceprint recognition technology is used to extract the personalized voiceprint features of the object from the sound data to be processed to obtain the object's voiceprint feature data. The voiceprint features refer to each person's unique voice characteristics, usually including the spectral characteristics of the voice, the pitch, tone, speed of the voice, etc. These features can help the system better understand the characteristic performance of the object when speaking.
[0033] Step 206, using the initial sound restoration parameters as restoration constraints, and according to the object voiceprint feature data and the object image data, repairing the speech variation and semantic defects of the sound data to be processed, to obtain processed sound data.
[0034] Restoration constraints can be rules and restrictions that must be followed during the sound restoration process. They can include the initially set goals (such as restoring speech speed, pitch, timbre, etc.), the personalized requirements of the subject (such as specific timbre, emotional expression, etc.), and environmental factors.
[0035] Among them, speech variation can be deviations in audio data due to factors such as pronunciation, accent, speaking speed, sound quality, etc. Speech variation may include regional accents, unclear pronunciation, non-standard intonation, etc., which usually affect the intelligibility and naturalness of speech.
[0036] Semantic defects can be inaccurate content, unclear logic or incomplete expression in the voice data, which may be caused by errors in the audio data output by the sound output object, improper expression or insufficient understanding of the context. For example, improper use of words, grammatical errors, information loss, etc. will affect the transmission of semantics.
[0037] The processed sound data may be sound data that has been repaired and preliminarily adjusted, and these data have already repaired some speech variations and semantic defects, but may still need further optimization.
[0038] Specifically, the system first compares and analyzes the object's voiceprint feature data with the pronunciation data of the standard pronunciation in the database to identify the types of variation in the voice, such as local accents, non-standard pronunciation, too fast or too slow speech speed, etc. Next, the object's voiceprint feature data is used in combination with audio signal processing technology to specifically adjust the pitch, timbre, speech speed and other speech parameters of the sound to be processed, thereby correcting the deviation in accent or pronunciation. For example, if the speech speed is detected to be too fast, the system will adjust the speech rhythm according to the normal speech speed range of the object to ensure clear and easy-to-understand pronunciation; if there is unclear pronunciation or timbre distortion, the system will restore the natural timbre of the audio through spectrum adjustment and sound quality enhancement technology. In addition, a noise suppression algorithm will be applied to remove background noise in the recording and clarify the voice signal. This repair process is iterative. After each repair, the system will analyze the repair effect and further optimize it until the repaired sound is highly consistent with the original voiceprint features of the object, and output the repair data of the voice variation.
[0039] Furthermore, the role of the object image data at this stage is to provide background information for speech restoration, especially in terms of emotional reasoning and contextual understanding. The system analyzes facial expressions, body language, and surrounding information extracted from the object image. For example, through image recognition technology, the system can judge the facial expression of the object (such as smiling, frowning, etc.), and thus infer the emotional state of the object (such as happiness, anger, tension, sadness, etc.). In addition, environmental factors (such as whether there are other people present, whether the surrounding sounds are noisy, etc.) can also be judged through visual clues in the image data. Based on this information, the system provides contextual background to help repair the emotional expression and context of the sound.
[0040] The system not only adjusts the voice features of the sound (such as pitch, speaking speed, intonation, etc.), but also uses the analysis results of the object's image data to infer the object's emotional state, and then repairs the semantic defects based on emotional reasoning. For example, when the system detects that the object expresses anger, it will not only make the tone of the voice more powerful and urgent, but may also adjust the tone of the sentence to ensure that the repaired voice conforms to the emotional color of anger. In this process, emotional reasoning not only affects the performance of the audio, but also identifies semantic deviations in the speech. For example, some words or sentences may not be clearly expressed due to insufficient emotions. Emotional reasoning technology will identify these inaccurate semantic information and appropriately adjust the emotional expression of the sentence, so that the speech is more in line with the actual emotional expression of the object during the repair process.
[0041] In addition, emotional reasoning not only repairs variations at the sound level, but can also identify and repair semantic defects. For example, in some emotional expressions, words may be used incorrectly, or the emotional expression of the sentence may not match the context, resulting in the listener being unable to accurately understand. The system adjusts sentence structure and word selection by reasoning about emotional semantics to ensure that the voice information conveys emotions while also being clear and accurate in meaning. For example, if the subject speaks in anger, the original sentence may not be able to accurately convey the emotion because it is too simple or indifferent. The system will repair the sentence based on reasoning so that it can be accurately expressed in both semantics and emotion.
[0042] In this step, the system further optimizes the repair effect through contextual understanding technology. Specifically, the contextual understanding technology identifies semantic defects in the sound data to be processed by analyzing the language context, conversation situation and background information of the object based on the analysis results of the object image data. These defects may include improper use of words, grammatical errors, unclear sentence structure or inconsistent sentence logic. For example, in a certain conversation scene, the use of certain words may lead to unclear information or misunderstanding. The system combines the analysis results of the object image data with contextual inference to adjust and repair these sentences to ensure that their semantics are more in line with the true intention of the conversation. Through contextual understanding, the system can not only analyze the language errors of the sentence itself, but also adjust the direction of language expression according to the current speech environment of the object to adapt to specific communication purposes and situational needs.
[0043] Finally, neurocognitive repair technology deeply repairs the sound data to be processed by simulating the way the brain processes language. Its core advantage lies in its ability to process complex language and speech features, especially at the semantic level, and can identify and repair potential semantic ambiguities, misunderstandings or missing parts. For example, when the system finds that the meaning expressed by the object is not clear enough or the grammatical structure is inappropriate, neurocognitive repair technology can perform semantic correction and reconstruct the sentence according to the context. This process includes not only simple grammatical corrections, but also deep semantic reasoning to ensure that the repaired speech not only meets the emotional and contextual requirements of the object, but also eliminates semantic ambiguity or information missing problems at a deep level.
[0044] In addition, neurocognitive repair technology can also handle potential cognitive biases in speech, such as mishearing, misunderstanding, or ambiguity in speech synthesis. Through in-depth analysis and correction of speech information, the system can make speech clearer and more accurate, while eliminating cognitive differences and understanding biases in speech communication; the above repair process is a multi-level and multi-dimensional repair, and finally outputs processed sound data.
[0045] Step 208: When it is detected that the processed sound data still has speech variations and / or semantic defects, the initial sound restoration parameters are adjusted according to the processed sound data to obtain adjusted sound restoration parameters.
[0046] The adjustment of the sound restoration parameters may be that when it is detected that there are still voice variations or semantic defects in the processed sound data, the system optimizes and resets the initial restoration parameters to obtain new parameters.
[0047] Specifically, after the initial repair, the system will detect the processed sound data to determine whether there are still voice variations or semantic defects; if it detects that these problems have not been completely eliminated, the system will adjust the initial repair parameters based on the feedback from the processed data. The adjustments may include changing the parameters of the repair algorithm, such as repair strength, algorithm weight, repair accuracy, etc., so as to more accurately solve the remaining problems in the next round of repair. In addition, the system may optimize the repair strategy based on new feedback information, such as using different repair methods for different voice defects to adjust the sound repair parameters.
[0048] Step 210, using the adjusted sound restoration parameters as initial sound restoration parameters, returning to the step of using the initial sound restoration parameters as restoration constraints, repairing the speech variations and semantic defects of the sound data to be processed according to the object voiceprint feature data and the object image data, and obtaining the processed sound data.
[0049] Specifically, the adjusted sound restoration parameters will be used as new initial sound restoration parameters, and the system will return and re-execute the restoration step. At this point, the restoration process will continue based on the updated restoration parameters, and the object voiceprint feature data and object image data of the object will be combined to restore the processed sound data. Through this cyclic restoration, the system can gradually optimize the restoration effect, solve the possible voice variations and semantic defects in the sound data, and ensure that each step of the restoration is more precise to achieve higher sound quality and accuracy.
[0050] Step 212, until no phonetic variation or semantic defect can be detected in the processed sound data, the processed sound data is used as the target sound data.
[0051] The target sound data can be the sound data that meets the requirements of the subject after multiple rounds of repair and adjustment. The target sound data not only repairs the variation and semantic defects in the voice, but also maintains the personalized timbre and emotional expression of the subject, ensuring that the voice is clear and natural, and meets the expression needs of the subject in a specific situation.
[0052] Specifically, this cyclic repair process will continue until the system detects that there are no more phonetic variations or semantic defects in the processed sound data. With each iteration of the previous repair parameters, new repair parameters are obtained, and the original data is repaired and adjusted again. The quality of the sound data will gradually improve until it reaches the predetermined repair standard. Finally, when the system confirms that the processed sound data has perfectly eliminated all problems with phonetic variations and semantic defects, and no further defects can be detected, the repair process will end. At this time, the system regards the processed sound data as the target sound data and outputs the final repair result.
[0053] In the above-mentioned audio data processing method based on artificial intelligence, the personalized acoustic features of the object are obtained by voiceprint feature extraction technology, and these features are combined with image data to repair the speech variations and semantic defects of the sound data to be processed. During the repair process, if it is detected that the sound data still has speech variations and / or semantic defects, the system will iteratively adjust the initial repair parameters to further optimize the repair effect. After multiple rounds of repair and parameter adjustment, until the speech variations and semantic defects in the sound data are completely eliminated, the target sound data is finally output; it can effectively improve the real-time repair efficiency of audio data with speech variations and semantic defects, and further significantly improve the performance of the speech processing system, especially in applications such as speech recognition, speech synthesis and voice communication, providing speech output with higher clarity, naturalness and accuracy.
[0054] In an exemplary embodiment, Figure 3 As shown, the initial sound restoration parameters are used as restoration constraints, and the voice variation and semantic defects of the sound data to be processed are repaired according to the object voiceprint feature data and the object image data to obtain the processed sound data, including steps 302 to 306. Among them:
[0055] Step 302, using the initial sound restoration parameters as restoration constraints, and according to the object voiceprint feature data and the object image data, serially repairing the speech variations and semantic defects of the sound data to be processed, to obtain serially processed sound data.
[0056] Serial processing of sound data means processing different problems in the sound to be repaired in a fixed order. In this mode, the repair process is divided into multiple stages. First, the system will process the voice variation part, such as accent, non-standard pronunciation, etc., and then gradually repair the semantic defects to ensure that the sentences are clear, logical and emotional. The processing results of each stage will be used as input for the next stage until all repair steps are completed.
[0057] Specifically, the system uses the initial sound restoration parameters as the restoration constraints, and combines the voiceprint feature data and image data of the object to gradually perform the restoration process. First, the voiceprint feature data of the object is compared and analyzed with the pronunciation data of the standard pronunciation in the database to identify the voice variation parts in the speech, such as accent, unclear pronunciation, inconsistent speech speed, etc., and combined with audio signal processing technology, the pronunciation deviation of the sound data to be processed is adjusted to restore the normal voice characteristics, ensuring that the restored pronunciation conforms to the personalized characteristics and standard pronunciation mode of the object. After the voice variation is repaired, the system will enter the stage of repairing semantic defects, and infer the emotional state and current context of the object by analyzing the object's image data (such as facial expressions, eyes, body language, etc.) and voiceprint feature data (such as pitch, speech speed, tone, etc.). For example, through changes in facial expressions and voice tones, the system can identify whether the subject is expressing emotions such as anger, anxiety or happiness, and thus determine whether there are semantic defects in the speech to be processed, such as whether the subject's wording is wrong, whether the sentence is incoherent or whether the grammar is irregular in the current emotional state and current context; when the above problems appear in the subject's speech, or there is unclear expression, logical confusion or language that does not conform to common expressions, the system will combine the subject's emotional state and specific information of the current context, and repair these problems on the data after the voice variation is repaired. For example, if the subject's speech uses inaccurate or inappropriate words, the system will infer more appropriate words based on the context and replace them on the data after the voice variation is repaired; if the sentence structure is not smooth, the system will reorganize the sentence to make it more in line with the language expression habits, ensure correct grammar and clear logic; if there is linguistic ambiguity or unclear expression in the speech, the system will adjust the sentence structure according to the subject's emotions and intentions to make the semantics clearer. In these ways, the system can effectively repair the semantic defects in the speech, so that the final speech expression is both smooth and fluent, and can accurately convey the subject's true intentions and emotions. This process is performed sequentially, which means that the repair of phonetic variations is completed before the repair of semantic defects, resulting in serially processed sound data.
[0058] Step 304, using the initial sound restoration parameters as restoration constraints, and according to the object voiceprint feature data and the object image data, simultaneously perform parallel restoration on the speech variation and semantic defects of the sound data to be processed, thereby obtaining parallel processed sound data.
[0059] Parallel processing of sound data means processing multiple problems of the sound data to be repaired at the same time. In this mode, voice variation repair and semantic defect repair are carried out in parallel, and no longer strictly in sequence. The system uses the power of parallel computing to repair voice features (such as pronunciation, tone, etc.) and semantic features (such as grammar, word usage, etc.) in different processing modules, and then integrates the results of different processing modules.
[0060] Specifically, the system will repair the voice variation and semantic defects of the sound data to be processed at the same time. Since the problem to be repaired has not changed, but the order of repair has changed from serial to parallel, the system still uses the initial sound repair parameters as constraints, and repairs the voice variation and semantic defects of the sound data to be processed in combination with the object voiceprint feature data and the object image data. Different from serial repair, the system can repair two aspects at the same time through parallel processing: on the one hand, the same as step 302, first compare and analyze the object voiceprint feature data with the pronunciation data of the standard pronunciation in the database, identify the voice variation parts in the speech, such as accent, unclear pronunciation, inconsistent speech speed, etc., and combine audio signal processing technology to adjust the pronunciation deviation of the sound data to be processed, restore normal voice characteristics, and ensure that the repaired pronunciation meets the personalized characteristics and standard pronunciation mode of the object.
[0061] On the other hand, the semantic defect is also the same as the specific implementation process of step 302, but the input data is different. The serial repair method is to repair the semantic defect on the data after the voice variation repair, while the parallel repair method is to repair the semantic defect on the sound data to be processed, that is, by analyzing the object image data (such as facial expressions, eyes, body language, etc.) and voiceprint feature data (such as pitch, speech speed, tone, etc.), the emotional state and current context of the object are inferred. For example, through the changes in facial expressions and pitch, the system can identify whether the object is expressing emotions such as anger, anxiety or happiness, so as to determine whether there are semantic defects in the speech to be processed, such as whether the object's wording is wrong, whether the sentence is not smooth or whether the grammar is not standardized in the current emotional state and current context; when the above problems appear in the object's voice, or there is unclear expression, logical confusion or language does not conform to the commonly used expression, the system will combine the specific information of the object's emotional state and the current context to repair these problems on the sound data to be processed. For example, if the subject uses inaccurate or inappropriate words in their speech, the system will infer more appropriate words based on the context and replace them in the sound data to be processed; if the sentence structure is not smooth, the system will reorganize the sentence to make it more in line with the expression habits of the language, ensuring correct grammar and clear logic; if there are linguistic ambiguities or unclear expressions in the speech, the system will adjust the sentence structure according to the subject's emotions and intentions to make the semantics clearer. In these ways, the system can effectively repair the semantic defects in the speech, so that the final speech expression is both smooth and fluent, and can accurately convey the subject's true intentions and emotions. The data after the speech variation is repaired and the data after the semantic defects are repaired are integrated to obtain parallel processing sound data.
[0062] Step 306 , cross-pipeline fusion of the serially processed sound data and the parallel processed sound data to obtain processed sound data.
[0063] Among them, cross-pipeline fusion can be the process of integrating the results of serial processing and parallel processing. At this stage, the system merges and optimizes the data results from the serial repair path and the parallel repair path to achieve a higher quality repair effect; serial processing can usually provide higher precision repairs, while parallel processing can speed up the repair process. The system intelligently integrates the advantages of the two, combines their repair results, adjusts the weights of different repair paths, and eliminates redundant parts.
[0064] Specifically, since serial and parallel repair each have their own advantages, the system needs to intelligently combine the results of the two repair paths according to different repair requirements. For example, in some cases, serial repair may perform better in detail repair (such as accurate repair of speech variations), while parallel repair is more efficient in handling large-scale semantic defect repair. The system will cross-pipeline fuse the repair results of the two. During the cross-pipeline fusion process, the results from the two paths of serial repair and parallel repair are evaluated, and the advantages and disadvantages of each path in dealing with speech variations and semantic defects are analyzed. By intelligently fusing the repair results of the two, combining their respective advantages, the weight of the repair effect is adjusted, and repeated or redundant repair steps are eliminated. At the same time, it is ensured that the final sound data meets the repair accuracy requirements and the repair effects are not distorted or conflicting, and finally the processed sound data is generated.
[0065] In this embodiment, by processing the repair processes of speech variations and semantic defects in series and in parallel respectively, and by cross-pipeline fusion, the accuracy and efficiency of sound restoration can be effectively improved. Among them, serial repair ensures the systematicness and coherence of the repair process, and each step solves specific problems one by one, while parallel repair speeds up the entire repair process, and minimizes the processing time by processing speech and semantic problems at the same time. In addition, cross-pipeline fusion combines the advantages of the two processing methods, not only retaining the refined adjustment of serial processing, but also relying on the high efficiency of parallel processing, thereby obtaining a more accurate and efficient sound restoration result. This multi-level, multi-angle repair method can effectively improve the quality of sound data and ensure that the repaired sound achieves the best balance in speech clarity and semantic accuracy.
[0066] In an exemplary embodiment, Figure 4 As shown, the initial sound restoration parameters are used as restoration constraints, and the voice variation and semantic defects of the sound data to be processed are serially restored in sequence according to the object voiceprint feature data and the object image data, to obtain serially processed sound data, including steps 402 to 406. Among them:
[0067] Step 402, using the initial sound restoration parameters as restoration constraints, and adjusting the pronunciation data of the sound data to be processed according to the object voiceprint feature data, to obtain accent-weakened sound data.
[0068] The pronunciation data may be various types of information about pronunciation extracted from the sound data to be processed by speech recognition technology, including features such as pitch, timbre, syllables, intonation, and speaking speed.
[0069] The accent-weakened sound data may be restored speech data that reduces or eliminates non-standard accent features in the subject's pronunciation.
[0070] Specifically, with the initial sound restoration parameters as the restoration constraints, the system will analyze the voiceprint features (such as pitch, timbre, speaking speed, etc.) of the object and compare them with the pronunciation data of the standard pronunciation in the database to identify the accent features and pronunciation deviations of the object. Based on the feature data of these differences, the system will apply specific algorithm models or audio signal processing technologies (such as multi-dimensional speech feature joint optimization recognition algorithm, cross-domain speech feature conversion algorithm, speech synthesis and accent correction algorithm based on transformation learning, etc.) to gradually optimize and adjust the pronunciation part. For example, if the object has a certain local accent, the system will adjust the pronunciation mode to make it closer to the standard Mandarin pronunciation or the target pronunciation style. The process not only includes the correction of syllables, tones, and intonations, but also ensures that the adjusted speech remains natural and fluent according to the speech characteristics of the object, avoiding making the speech appear too artificial or inconsistent with the personalized voiceprint characteristics of the object. Finally, through these adjustments, the system generates accent-weakened sound data, making the speech clearer and more standard, meeting the target pronunciation requirements.
[0071] Step 404 , based on the object image data and the accent-weakened sound data, infer the generation environment of the sound data to be processed to obtain scene feature data.
[0072] Among them, scene feature data can be analyzed by analyzing the object's image data (such as facial expressions, eyes, body posture, etc.) and voice data to infer the specific situation or environmental information of the object.
[0073] Specifically, the system uses the object image data to perform situational reasoning to analyze and infer the environment in which the object is located. The object image data provides important non-verbal information about the object's facial expression, eyes, posture, body shape, etc., reflecting the object's current emotional state and context. For example, facial expressions can help the system determine whether the object is in a state of tension, pleasure, or confusion, and eyes and posture may reveal the spatial environment in which the object is located (such as home, office, or outdoors). The system compares the above-mentioned recognized image data with the known situational model, and through sentiment analysis and situational reasoning, obtains specific scene feature data, including environmental noise (such as background music, crowd noise, etc.), conversation atmosphere (such as formal or informal), and emotional color (such as relaxed, serious, etc.).
[0074] Step 406, modifying the language expression data of the accent-weakened sound data according to the scene feature data to obtain serially processed sound data.
[0075] Specifically, the system further repairs the language expression of the voice content after the accent is weakened based on the previously inferred scene feature data. In the specific implementation process, the system will infer the accuracy of the language expression and the language style used in a specific environment based on the scene feature data. For example, in formal meetings or business occasions, the sentences or words in the voice need to be more formal and standardized, and avoid the appearance of foul language; in friends' gatherings or family exchanges, the sentences or words in the voice should be more casual and friendly, and avoid the appearance of sentences that do not match the current scene (friends encounter difficulties and use inappropriate terms such as "you deserve it"). By identifying the language needs in these scenes, the system will adjust the vocabulary, tone and sentence structure in the voice to make it consistent with the language expression in the current environment. In addition, the system will also correct the tone intensity, emotional communication and other aspects according to the context, so that the voice content is more consistent with the emotional state and situation of the object, and finally the system generates serial processing sound data.
[0076] In this embodiment, by combining the object voiceprint feature data and the object image data, multi-level repair of the sound data to be processed is achieved. Among them, adjusting the pronunciation data based on the voiceprint features can effectively weaken the accent differences, making the sound more standardized and easy to understand. And combining the object image data for environmental reasoning can identify the actual scene characteristics behind the sound, providing support for further contextual analysis. Finally, the language expression is modified based on the scene feature data to ensure the accuracy and naturalness of the sentence in different contexts. This process optimizes pronunciation and semantics through precise steps, making the repaired sound not only clearer, but also better adapted to specific situations and expression needs, thereby improving the comprehensiveness and adaptability of sound repair.
[0077] In an exemplary embodiment, Figure 5 As shown, according to the scene feature data, the language expression data of the accent weakened sound data is modified to obtain the serially processed sound data, including steps 502 to 508. Among them:
[0078] Step 502: Perform contextual semantic recognition on the accent-weakened sound data to obtain sound semantic recognition data.
[0079] Among them, contextual semantic recognition can be that when analyzing speech or text, the meaning of words, phrases or sentences is understood not only based on their individual meanings, but also based on their position and relationship in the entire conversation or context. For example, the word "bank" can refer to a financial institution or a river bank, depending on the context. In actual operation, the system uses natural language processing (NLP) techniques such as dependency parsing and semantic role labeling to identify the structure of the sentence and determine the exact meaning of the vocabulary based on the context.
[0080] Among them, the sound semantic recognition data can be a data set that analyzes the sound data to be processed and extracts the semantic information in the speech. This process usually includes parsing the meaning of each word or sentence in the speech and identifying possible semantic problems, ambiguous or unclear expressions. For example, the system may recognize that some words in the speech are semantically ambiguous due to unclear pronunciation or unclear context, and record this information as sound semantic recognition data.
[0081] Specifically, through natural language processing technology, the accent-weakened sound data is converted into text using speech recognition technology and analyzed in combination with contextual information. At this time, the system not only analyzes the independent meaning of each word or phrase, but also identifies the overall structure, grammatical relations, and logical coherence of the sentence. For example, the system will identify and mark the keywords, verbs, nouns, etc. in the sentence, and check their grammatical correctness in the current context. At the same time, the system will detect semantic ambiguities that may appear in the speech, such as incorrect vocabulary usage (such as the wrong context of using "improve"), and grammatical incoherence (such as inconsistency between subject and predicate). Through contextual semantic recognition, the system can extract the true intention behind each sentence and generate sound semantic recognition data.
[0082] Step 504: repair the objective semantic defects of the accent-weakened sound data according to the sound semantic recognition data to obtain semantically repaired sound data.
[0083] Among them, objective semantic defects can be situations in which the expression is unclear, ambiguous or inaccurate due to problems with grammar, logic, vocabulary choice or sentence structure in speech or text.
[0084] Among them, semantically repaired sound data can be sound data that, after repair processing, can eliminate objective semantic defects in speech or text, making its expression clearer, more accurate and in line with the context.
[0085] Specifically, based on the previously extracted sound semantic recognition data, the objective semantic defects in the accent-weakened sound data content are repaired, where the objective semantic defects include objective grammatical errors, words that do not conform to conventional expressions, logical confusion and other problems, that is, the expression used is not a normal expression. For example, when the system finds that a certain word in the sound semantic recognition data is used incorrectly (for example, in the process of resource trading, the word "buy" should be used for the object, but the system recognizes that the object uses "take away"), or finds that a sentence is grammatically incoherent (such as the lack of conjunctions leading to ambiguous sentence meaning), it will automatically correct it. Through natural language processing technology (such as grammatical correction, synonym replacement, etc.), the system can repair these semantic defects and generate semantically repaired sound data, so that the voice content is more in line with standard grammar, the expression is more accurate, and any errors that may cause ambiguity are avoided.
[0086] Step 506, identifying the emotion of the sound output object based on the scene feature data and the object image data to obtain object scene recognition data.
[0087] Among them, object scenario recognition data can be data obtained by analyzing the object's behavior, expression, context, environment and other external information, and the system infers the object's current emotional state, intention or specific scenario.
[0088] Specifically, the system will identify and infer the emotional state of the object based on the information from the object image data and scene feature data. Specifically, by analyzing the object's facial expressions, eyes, body posture and emotional signals in the voice (such as intonation, speaking speed, and pitch changes) in the object image data, the system can infer the emotional state of the object in the current conversation. For example, if the object's facial expression shows tension or uneasiness, the system may judge that the object is experiencing a tense or anxious emotional state. Conversely, if the facial expression is relaxed and happy, the system may infer that the object feels relaxed or happy. Combining the scene feature data of the current object (such as the type of background noise, environmental atmosphere, etc.) also provides strong support for emotion recognition, helping the system to assist in capturing the object's emotional expression and situational state, and generating object scenario recognition data, that is, the system's evaluation result of the object's emotional state, providing guidance for emotional expression for subsequent voice restoration.
[0089] Step 508: repair the scene semantic defects of the semantic repair sound data according to the object scene recognition data to obtain serially processed sound data.
[0090] Specifically, the system corrects the emotion and context of the sound data that has been repaired for objective semantic defects based on the generated object scene recognition data, that is, repairs the scene semantic defects; the scene semantic defects refer to the places where the words or sentences of the voice may not match the current emotion and scene in expression. These defects often show the mismatch of words or sentences in emotional transmission or the incoordination of voice tone. For example, in a serious occasion, if the words used in the voice or the expression of the whole sentence are too casual or relaxed, the system will use the emotion recognition data to correct the words of the voice or adjust the expression of the whole sentence to make it more in line with the formal or tense situation; if the object is expressing sadness or frustration, the system will correct the sentence expression and sentence structure in the voice accordingly to make it consistent with the emotional state. At the same time, the system will also take into account other factors in the context, such as whether formal terms are needed or whether certain words should be avoided from being too blunt. Through these adjustments, the system ensures that the final voice is not only grammatically and logically correct, but also uses reasonable expressions and reasonable words in the emotional state of the object and the situation in which it is located, and generates serially processed sound data.
[0091] In this embodiment, through contextual semantic recognition, the clarity of the sound at the semantic level after the accent is weakened is ensured, and the problem of ambiguous expression that may exist in the sentence is solved; and the objective semantic defects are repaired to further optimize the accuracy and fluency of the language. By combining scene feature data and object image data, the system can accurately identify the emotional state of the object and give the sound data a richer emotional color in the expression of language. Finally, the scene context defects in the semantics are repaired using emotion and scene recognition data, so that the language expression of the repaired sound is not only more in line with the context, but also can accurately convey emotions and intentions. Overall, through a multi-dimensional repair process, the language expression of the sound is made more natural, accurate and emotional, which comprehensively improves the transmission effect of the sound data and the user experience.
[0092] In an exemplary embodiment, Figure 6 As shown, according to the object scene recognition data, the scene semantic defects of the semantic repair sound data are repaired to obtain the serially processed sound data, including steps 602 to 606. Among them:
[0093] Step 602, repairing the speech emotion expression of the semantically repaired sound data according to the emotion reasoning data of the object scene recognition data, to obtain emotion enhanced recognition data.
[0094] Among them, speech emotion expression can be the words, sentences, paragraphs and structures in the speech to convey the speaker's emotional state or attitude.
[0095] Among them, emotion enhancement recognition data can be the analysis and processing of speech data to be processed, by modifying the expression of features such as words, sentences, paragraphs and structures to enhance the emotional information therein, so that the repaired speech can better convey the speaker's true emotions.
[0096] Specifically, the problem of insufficient emotional expression in speech is repaired based on the emotional reasoning data in the object scene recognition data. Since the emotional reasoning data is obtained by analyzing the emotional features in the facial expressions, posture, voice intonation and speech content of the object, it helps the system understand the current emotional state of the object (such as joy, anger, sadness, etc.). The system uses the emotional reasoning data to adjust the sound data that has been semantically repaired. Specifically, the system will evaluate the parameters such as the words, sentence expressions, paragraph expressions, and overall structure in the speech, and optimize them according to the inferred emotional reasoning data, that is, combine these emotional reasoning data with the semantically repaired sound data to improve the words, sentences, paragraphs, etc. in the speech that are insufficient or wrong in expressing emotions. For example, the original sentence may be written as "I feel a little uncomfortable", but this expression is relatively bland and is not enough to reflect the emotional intensity of the object. The system will modify the sentence to "I feel very uneasy and almost collapsed", and improve the intensity and accuracy of emotional expression through stronger emotional vocabulary and tone correction. Finally, the system outputs emotional enhancement recognition data, which makes the emotional expression of the speech more accurately aligned with the emotional state of the object, and reflects richer emotional colors in the speech.
[0097] Step 604, repairing the dialogue subject scene of the emotion enhancement recognition data according to the scene inference data of the object scenario recognition data, to obtain the scene enhancement recognition data.
[0098] The dialogue theme scene may be the situation, background, topic of discussion or context involved in the dialogue.
[0099] Among them, scene enhancement recognition data can be the context, background and dialogue topic scene of the analysis and reasoning object. By modifying the expression of features such as words, sentences, paragraphs and structures, these scene information can be enhanced to help better understand the potential semantic needs in the speech.
[0100] Specifically, the system further repairs the speech based on the scene reasoning data of the object scene recognition data to ensure that it conforms to the actual scene of the conversation. The scene reasoning data is inferred by analyzing the environmental background, context and context of the object's conversation content. For example, if the system detects that the object is in a formal meeting, the expression in the voice should be more rigorous and standardized; if the object is in a casual family gathering, the tone and words should be more relaxed and casual. The system first determines whether the content is consistent with the current scene from the emotion enhancement recognition data, and repairs the words, sentences, paragraphs or structures that do not match the current scene; suppose in a business meeting scene, the original sentence may be expressed as "We can think about this later", which seems too casual and imprecise in formal occasions. The system recommends modifying it to "We can discuss and make a decision in a later meeting" through scene reasoning. Similarly, if the conversation is in a relaxed gathering, the system may suggest adjusting the tone to "Let's talk later, don't worry, let's discuss it together." This restoration ensures the consistency of the content and tone of the scene-enhanced recognition data, ensuring that the language is both emotionally expressive and matches the context and social setting of the conversation.
[0101] Step 606, simulate neurocognitive optimization of the linguistic features of the scene enhancement recognition data to obtain serially processed sound data.
[0102] Among them, simulated neurocognitive optimization can be a technology that optimizes speech processing and understanding by simulating the cognitive process of the human brain. It combines the principles of neural networks and cognitive psychology to simulate how humans understand, analyze and generate speech or language. Through simulated neurocognitive optimization, the system can more intelligently identify subtle emotional changes in speech, semantic defects and potential problems in language expression. For example, when the system identifies grammatical errors or unnatural sentences in speech, simulated neurocognitive optimization can optimize sentence structure, adjust expression methods, and make them more in line with human cognitive patterns by simulating the human language understanding process, thereby improving the naturalness and accuracy of speech output.
[0103] Specifically, the system performs a deeper linguistic optimization on the voice data after emotion and scene repair, focusing on correcting the deficiencies in the language structure and improving its grammatical correctness, fluency and logic. Through the simulation of neurocognitive optimization technology, the system simulates how the human brain processes language and adjusts grammatical errors, unclear expressions or logical incoherence in the voice data. Suppose there is a sentence in the original voice: "At the beginning of the meeting, I felt that I was completely unprepared, so I was late, which was really embarrassing." The expression of "this situation is really embarrassing" here is slightly abrupt and the logic is slightly confusing, which is easy to cause understanding difficulties. The system can modify this sentence to: "At the beginning of the meeting, I was not fully prepared, and I missed the start time, which made me feel very embarrassed." By reorganizing the sentence structure, grammatical and logical problems are fixed, while ensuring the fluency and readability of the sentence. In addition, the system will analyze the redundant parts of the sentence based on the cognitive model and simplify lengthy or repetitive expressions, such as simplifying the original "I feel completely unprepared" to "I am not ready", which not only improves the clarity of the language, but also makes the language expression more in line with the language processing mode of human cognition. Through this simulation optimization, the generated serial processing sound data becomes more refined and natural in terms of grammar, semantics and logical structure, and can better convey emotions and scene adaptation.
[0104] In this embodiment, through emotional reasoning, the system can repair the language used for emotional expression in the speech according to the emotional state of the object, so that the speech is more in line with the current emotional background of the object. For example, if the emotional state of the object is anxious or happy, the change in language will enhance the emotional color of the speech, making it more expressive. The scene reasoning repairs the conversation topic and scene background of the speech to ensure the semantic consistency and natural fluency of the language of the speech content in a specific environment and context. Finally, after simulation neurocognitive optimization, the linguistic features in the speech are further refined, and the auditory effect and cognitive acceptability of the sound are improved. Overall, through the dual repair of emotions and scenes, the final sound data is more vivid and immersive, and can accurately convey emotions and context, which improves the interactive experience between the speech and the listener.
[0105] In an exemplary embodiment, Figure 7 As shown, the initial sound restoration parameters are used as restoration constraints, and according to the object voiceprint feature data and the object image data, the voice variation and semantic defects of the processed sound data are restored in parallel, and the parallel processed sound data is obtained, including steps 702 to 718. Among them:
[0106] Step 702: Identify the emotion of the sound output object based on the object image data to obtain object scene recognition data.
[0107] Specifically, the technical principle of identifying the object scene recognition data from the object image data is consistent with the recognition of the object image data in step 506, but the object scene recognition data finally obtained in step 506 needs to be obtained in combination with the information in the scene feature data, but here it is only obtained by identifying the object image data. Specifically, by analyzing the object's facial expressions, eyes, body posture and emotional signals in the voice (such as intonation, speaking speed, pitch changes) in the object image data, the system can infer the emotional state of the object in the current conversation. For example, if the object's facial expression shows tension or uneasiness, the system may judge that the object is experiencing a tense or anxious emotional state. On the contrary, if the facial expression is relaxed and happy, the system may infer that the object feels relaxed or happy, and finally the object scene recognition data is summarized and generated.
[0108] Step 704: Perform context semantic recognition on the sound data to be processed to obtain sound semantic recognition data.
[0109] Specifically, the technical principle of performing contextual semantic recognition on the processed sound data to obtain the sound semantic recognition data is consistent with the recognition of the accent weakened sound data in step 502, but the input data for obtaining the sound semantic recognition data in step 502 is the accent weakened sound data, but the input data for obtaining the sound semantic recognition data here is the sound data to be processed. Specifically, through natural language processing technology, the processed sound data is converted into text using speech recognition technology and analyzed in combination with context information. At this time, the system not only analyzes the independent meaning of each word or phrase, but also recognizes the overall structure, grammatical relationship and logical coherence of the sentence. For example, the system will identify and mark the keywords, verbs, nouns, etc. in the sentence, and check their grammatical correctness in the current context. At the same time, the system will detect semantic ambiguities that may appear in the speech, such as vocabulary usage errors (such as "improvement" using the wrong context), and grammatical incoherence (such as subject-predicate inconsistency). Through contextual semantic recognition, the system can extract the true intention behind each sentence and generate sound semantic recognition data.
[0110] Step 706, copy and isolate the sound data to be processed to obtain each isolated sound data.
[0111] The isolated sound data may be data obtained by copying a voice signal into multiple data when analyzing and processing the sound data to be processed.
[0112] Specifically, the system copies and isolates the sound data to be processed, and copies the audio data into multiple independent sound data isolation units, each of which contains different repair tasks, to obtain multiple isolated sound data. For example, the system may copy the sound data to be processed into multiple identical data, each of which focuses on different dimensions of speech: for example, one segment only processes accents, another segment processes grammatical errors, and another segment processes emotional expressions, etc.
[0113] Step 708, using the initial sound restoration parameters as restoration constraints, and adjusting the pronunciation data of any selected isolated sound data according to the object voiceprint feature data, to obtain accent-weakened sound data.
[0114] Specifically, the technical principle of adjusting the pronunciation data of the isolated sound data according to the voiceprint feature data of the object is consistent with the recognition of the accent-weakened sound data in step 402 (the isolated sound data is the copied sound data to be processed). Specifically, the initial sound repair parameters are used as the repair constraint conditions. The system analyzes the voiceprint features of the object (such as pitch, timbre, speech speed, etc.) and compares them with the pronunciation data of the standard pronunciation in the database to identify the accent characteristics and pronunciation deviations of the object. According to the feature data of these differences, the system will apply specific algorithm models or audio signal processing technologies (such as multi-dimensional speech feature joint optimization recognition algorithm, cross-domain speech feature conversion algorithm, speech synthesis and accent correction algorithm based on transformation learning, etc.) to gradually optimize and adjust the pronunciation part. For example, if the object has a certain local accent, the system will adjust the pronunciation mode to make it closer to the standard Mandarin pronunciation or the target pronunciation style. The process not only includes the correction of syllables, tones, and intonations, but also ensures that the adjusted voice remains natural and fluent according to the voice characteristics of the object, so as to avoid making the voice appear too artificial or inconsistent with the personalized voiceprint characteristics of the object. Ultimately, through these adjustments, the system generated accent-weakened sound data, making the speech clearer, more standard, and in line with the target pronunciation requirements.
[0115] Step 710, based on the sound semantic recognition data, the objective semantic defects of any isolated sound data that are not selected are repaired to obtain semantically repaired sound data.
[0116] Specifically, in the same way, the technical principle of repairing the objective semantic defects of isolated sound data according to the sound semantic recognition data is consistent with the repair of the objective semantic defects of the accent weakened sound data in step 504, except that the object to be repaired is changed from the accent weakened sound data to the isolated sound data (i.e., the sound data to be processed). Specifically, according to the sound semantic recognition data extracted previously, the objective semantic defects in the content of the isolated sound data are repaired, wherein the objective semantic defects include objective grammatical errors, words that do not conform to conventional expressions, logical confusion, and other problems, that is, the expression adopted is not a normal expression. For example, when the system finds that a certain word in the sound semantic recognition data is used incorrectly (for example, in the process of resource trading, the word "buy" should be used for the object, and the system recognizes that the object uses "take away"), or finds that a sentence is grammatically incoherent (such as lack of conjunctions leading to ambiguous sentence meaning), it will automatically correct it. Through natural language processing technology (such as grammatical error correction, synonym replacement, etc.), the system can repair these semantic defects and generate semantic repair sound data, so that the voice content is more in line with standard grammar, the expression is more accurate, and any errors that may cause ambiguity are avoided.
[0117] Step 712, based on the emotional inference data of the object scenario recognition data, the speech emotional expression of the pronunciation data of any unselected isolated sound data is repaired to obtain the emotional enhancement recognition data.
[0118] Specifically, in the same way, the technical principle of repairing the speech emotional expression of the isolated sound data according to the emotional reasoning data of the object scene recognition data is consistent with the step 602 for repairing the speech emotional expression of the semantic repair sound data, except that the repaired object changes from the speech emotional expression of the semantic repair sound data to the speech emotional expression of the isolated sound data. Specifically, the problem of insufficient emotional expression in the voice is repaired based on the emotional reasoning data in the object scene recognition data. Since the emotional reasoning data is obtained by analyzing the emotional features in the facial expression, posture, voice intonation and speech content of the object, it helps the system understand the current emotional state of the object (such as joy, anger, sadness, etc.). The system uses the emotional reasoning data to adjust the isolated sound data. Specifically, the system will evaluate the parameters such as the words, sentence expression, paragraph expression, and overall structure in the voice, and optimize it according to the inferred emotional reasoning data, that is, combine these emotional reasoning data with the isolated sound data to enhance the insufficient or wrong emotional expression of words, sentences, paragraphs, etc. in the voice. For example, the original sentence may be written as "I feel a little uncomfortable", but this expression is relatively bland and is not enough to reflect the emotional intensity of the object. The system will modify the sentence to "I feel very uneasy and almost collapsed", and improve the intensity and accuracy of emotional expression through stronger emotional vocabulary and tone correction. Finally, the system outputs emotional enhancement recognition data, which makes the emotional expression of the voice more accurately aligned with the emotional state of the object, and reflects richer emotional colors in the voice.
[0119] Step 714, based on the scene inference data of the object scene recognition data, repair the dialogue subject scene of any isolated sound data pronunciation data to obtain scene enhanced recognition data.
[0120] Specifically, similarly, the technical principle of repairing the dialogue theme scene of the isolated sound data according to the scene reasoning data of the object scene recognition data is consistent with the repair of the dialogue theme scene of the emotion enhancement recognition data in step 604, except that the object to be repaired changes from the voice emotion expression of the emotion enhancement recognition data to the voice emotion expression of the isolated sound data. Specifically, the system further repairs the voice according to the scene reasoning data of the object scene recognition data to ensure that it conforms to the actual scene of the conversation. The scene reasoning data is inferred by analyzing the environmental background, context and context of the conversation content of the object. For example, if the system detects that the object is in a formal meeting, the expression in the voice should be more rigorous and standardized; and if the object is in a casual family gathering, the tone and wording should be more relaxed and casual. The system first determines whether these contents are consistent with the current scene from the isolated sound data, and repairs the words, sentences, paragraphs or structures that do not match the current scene; assuming that in a business meeting scene, the original sentence may be expressed as "We can consider this later", which is too casual and not precise in formal occasions. The system recommends modifying it to "We can discuss and make decisions in a later meeting" through scene reasoning. Similarly, if the conversation is in a relaxed gathering, the system may suggest adjusting the tone to "Let's talk later, don't worry, let's discuss it together." This repair ensures the consistency of the content and tone of the scene-enhanced recognition data, ensuring that the language can express emotions and match the context and social occasion of the conversation.
[0121] Step 716, merging the accent weakened sound data, the semantically restored sound data, the emotion enhanced recognition data, and the scene enhanced recognition data to obtain fused sound restoration data.
[0122] Among them, the fusion of sound restoration data can be to integrate the sound data that has undergone different processing steps to generate the final restoration result. During the speech restoration process, the system may process different aspects of the same speech multiple times, such as weakening the accent in the speech, repairing the semantics, enhancing the emotional expression, etc. These processing steps generate different data sets, such as "accent weakened sound data", "semantic repair sound data", and "emotion enhancement recognition data". The process of fusing sound restoration data is to merge these independently processed data to ensure their coherence in the timeline and voice stream, while retaining the optimization effect brought by each restoration step.
[0123] Specifically, the accent weakened sound data, semantic repair sound data, emotion enhancement recognition data, and scene enhancement recognition data are merged into a unified audio file. This fusion process can combine the results of each repair module using weighted average, feature splicing, or other audio synthesis algorithms to ensure that the repair of each dimension can be presented in the final speech to obtain fused sound repair data.
[0124] Step 718, simulated neurocognitive optimization is performed on the linguistic features of the fused sound restoration data to obtain parallel processed sound data.
[0125] Specifically, similarly, the technical principle of simulating neurocognitive optimization of the linguistic features of the fused sound repair data is consistent with the simulating neurocognitive optimization of the linguistic features of the scene enhancement recognition data in step 606, except that the repaired object is changed from the linguistic features of the scene enhancement recognition data to the linguistic features of the fused sound repair data. Specifically, the system performs a deeper linguistic optimization on the fused data after emotion and scene repair, focusing on correcting the deficiencies in the language structure and improving its grammatical correctness, fluency and logic. Through the simulation neurocognitive optimization technology, the system simulates how the human brain processes language and adjusts grammatical errors, unclear expressions or logically incoherent parts in the voice data. Suppose there is a sentence in the original voice: "At the beginning of the meeting, I felt that I was not fully prepared, so I was late, which was really embarrassing." The expression of "this situation is really embarrassing" here is slightly abrupt and slightly confusing in logic, which is easy to cause comprehension difficulties. The system can modify this sentence to: "At the beginning of the meeting, I was not fully prepared, and I missed the start time, which made me feel very embarrassed." By reorganizing the sentence structure, grammatical and logical problems are fixed, while ensuring the fluency and readability of the sentence. In addition, the system will analyze the redundant parts of the sentence based on the cognitive model and simplify lengthy or repetitive expressions, such as simplifying the original "I feel completely unprepared" to "I am not ready", which not only improves the clarity of the language, but also makes the language expression more in line with the language processing mode of human cognition. Through this simulation optimization, the generated parallel processing sound data has become more refined and natural in terms of grammar, semantics and logical structure, and can better convey emotions and scene adaptation.
[0126] In this embodiment, through emotion recognition and contextual semantic recognition based on object image data, the emotional state and context of the sound output object can be accurately grasped, ensuring that the subsequent repair operation can be consistent with the object's real emotion and context. Copy isolation and targeted repair independently optimize pronunciation, semantics and emotional expression, respectively, solve the problems of accent, unclear semantics and insufficient emotional expression, and improve the naturalness and expression accuracy of speech. Especially in terms of speech emotion expression and scene reasoning, the system can repair the emotional loss and contextual incoordination of language expression in the sound based on emotional reasoning and scene reasoning data, making the speech more emotionally expressive and more natural in sentence expression. Finally, the simulation neurocognitive optimization refines the linguistic features, further improving the auditory effect and cognitive acceptance of the sound. Overall, the repair process not only improves the clarity and semantic accuracy of the sound, but also better conveys the emotions and intentions of the object, significantly enhancing the interactive experience between the sound and the listener.
[0127] In an exemplary embodiment, Figure 8 As shown, the initial sound restoration parameters are adjusted according to the processed sound data to obtain the adjusted sound restoration parameters, including steps 802 to 806. Among them:
[0128] Step 802: determining, from the processed sound data, a speech data adjustment rule corresponding to the speech variation and a semantic data adjustment rule corresponding to the semantic defect.
[0129] Among them, the voice data adjustment rules can be a series of adjustment specifications and strategies followed when repairing voice variations in sound data (such as accents, unclear pronunciation, too fast speaking speed, etc.). These rules determine how to improve or optimize the output of voice by analyzing factors such as voiceprint features, pronunciation patterns, and voice quality. For example, if the system detects that the pronunciation of a syllable is unclear, the adjustment rules may include increasing the volume of the syllable, changing the pitch or speaking speed, or even adjusting the pronunciation method to ensure that the syllable can be clearly understood.
[0130] Among them, semantic data adjustment rules can be a set of norms and methods to repair semantic defects in speech (such as incoherent sentences, inappropriate words or grammatical errors, etc.). These rules repair and optimize by analyzing sentence structure, grammatical consistency, vocabulary accuracy and clarity of expression. For example, if a word in a sentence is used improperly or the expression is unclear, the semantic adjustment rules will suggest replacing it with a more precise and appropriate word; if the sentence structure is complex or does not conform to grammatical norms, the rules may suggest simplifying the sentence or reorganizing the word order.
[0131] Specifically, since it is detected that the processed sound data still has voice variations and / or semantic defects after being repaired, the system detects the accent or pronunciation pattern in the processed sound data through voiceprint analysis technology, and further determines which pronunciation features need to be adjusted, such as whether the pronunciation is unclear after certain syllable signs, or whether the characteristics of a certain region's accent still affect the standardized expression of speech after adjustment. Then the semantic defect analysis will further identify the unclear expression of sentences in the processed sound data, such as inaccurate vocabulary, incoherent sentences, or insufficient emotional expression. For example, a sentence may be semantically ambiguous due to unclear grammatical structure, or the listener may not be able to accurately understand the speaker's intention due to insufficient emotional expression. The system will set corresponding more stringent repair rules for these problems, or increase the sensitivity of the repair rules. The voice data adjustment rules are to adjust the accuracy of pronunciation, the clarity of speech, etc., while the semantic data adjustment rules include modifying vocabulary, sentence patterns and grammatical structures to ensure the accuracy and clarity of semantics.
[0132] Step 804: Identify adjustment strategy correlation information between the voice data adjustment rule and the semantic data adjustment rule.
[0133] Among them, the adjustment strategy relevance information can be to identify and understand the relationship and influence between voice data and semantic data when repairing them. This information helps the system determine how to coordinate voice and semantic repair to avoid conflicts between the two repair processes. For example, when repairing the pronunciation of speech, it may affect the original semantic expression, especially when the pronunciation is not standard and causes vocabulary confusion; similarly, modifying the semantics may affect the pronunciation and intonation of the sentence.
[0134] Specifically, the system needs to identify the interaction and influence between speech and semantic adjustment rules. For example, accent variation in speech may not only affect the clarity of pronunciation, but may also lead to semantic misunderstandings, especially in polyphones, synonyms, and sentence structures. The system needs to identify the correlation between these adjustment rules to ensure that speech and semantic repairs can complement and coordinate. If the pronunciation problem of speech is repaired, but the semantic expression is not corrected accordingly, it may lead to deviations in the listener's understanding. Therefore, the system will use the influence relationship between speech adjustment and semantic repair as part of the adjustment strategy, and use the interaction between speech clarity and semantic accuracy to make corrections. For example, if there is an accent problem in speech, it may lead to some vocabulary misunderstandings. At this time, speech repair not only needs to adjust the pronunciation, but also needs to combine semantic repair rules to ensure that the repaired speech will not cause ambiguity in the meaning conveyed.
[0135] The specific implementation is to identify the relationship between speech variations (such as accents, unclear pronunciation, etc.) and semantic defects (such as grammatical errors, ambiguous expressions, etc.) through in-depth analysis of the processed sound data. For example, when the accent in the speech causes the pronunciation of certain words to be non-standard, it may cause semantic misunderstanding or ambiguity. The system will evaluate this relationship and consider the coupling of speech and semantics when using the speech data adjustment rules and semantic data adjustment rules. The adjustment strategy correlation information is based on this evaluation result, that is, when using the speech data adjustment rules and semantic data adjustment rules, the repair of speech and semantics is considered to ensure that there is no conflict when the two are repaired. For example, if the repair of pronunciation may cause changes in semantic expression, the system will adjust the semantic repair rules accordingly to avoid logical inconsistencies. By analyzing the relationship between these rules, the adjustment strategy correlation information between the speech data adjustment rules and the semantic data adjustment rules is obtained.
[0136] Step 806, adjusting the initial sound restoration parameters according to the voice data adjustment rule, the semantic data adjustment rule and the adjustment strategy correlation information to obtain adjusted sound restoration parameters.
[0137] Specifically, the system will optimize and adjust the initial sound repair parameters by integrating the voice data adjustment rules, semantic data adjustment rules, and the correlation information between them. Since the initial repair parameters usually include model repair parameters related to voice (such as model repair parameters for tone, speech speed, pitch, volume, pronunciation clarity, etc.) and model repair parameters related to semantics (such as model repair parameters for sentence pattern, accuracy of vocabulary use, grammatical structure, etc.), according to the constraints of the voice and semantic data adjustment rules obtained in the first two steps, combined with the adjustment strategy correlation information of the two, the system will adjust the above model parameters at the same time. For example, if the system detects that the accent problem in the voice affects the expression of semantics, then while adjusting the parameters of the audio processing model with pronunciation clarity as the target and satisfying the adjustment rules, it may be necessary to adjust the parameters of the audio processing model with the grammatical structure at the semantic level as the adjustment target and satisfying the adjustment rules to ensure that the repaired voice is both clear and grammatical. In this process, the system will continuously iterate and optimize through algorithms until the most appropriate adjustment of the sound repair parameters is obtained to achieve the maximum repair effect of voice variation and semantic defects.
[0138] In this embodiment, by accurately analyzing the speech variations and semantic defects in the processed sound data, the corresponding adjustment rules can be effectively determined, thereby providing clear guidance for sound restoration. In this process, the system can not only identify the adjustment rules for speech data and semantic data, but also analyze the correlation between the adjustment strategies of the two, to ensure that the accuracy of semantics or the integrity of emotional expression will not be destroyed when repairing speech variations. This correlation analysis of adjustment strategies provides a global optimization perspective for sound restoration, making the repair process more coordinated and efficient. Through the adjustment of comprehensive speech and semantic parameters, the system can improve the clarity of semantics and the accuracy of expression while ensuring the natural fluency of speech, and finally obtain a repaired sound data that is more in line with the actual situation, emotion and context of the object, significantly improving the quality of the repair results and the listener's experience.
[0139] In an exemplary embodiment, Fig. 9 As shown, the initial sound restoration parameters are adjusted according to the voice data adjustment rules, the semantic data adjustment rules and the adjustment strategy correlation information to obtain the adjusted sound restoration parameters, including steps 902 to 906. Among them:
[0140] Step 902: determining initial adjustment values of speech parameters according to speech data adjustment rules, and determining initial adjustment values of semantic parameters according to semantic data adjustment rules.
[0141] Among them, the initial value of the speech parameter adjustment can be a preliminary parameter adjustment value set first based on the analysis of the processed sound data. These parameter values are based on the speech data adjustment rules, such as adjusting the preliminary parameters of the speech pitch, speaking speed, pronunciation clarity, sound quality, etc.
[0142] Among them, the initial value of the semantic parameter adjustment can be the initial parameter adjustment value corresponding to the initial adjustment plan set by the system for vocabulary, syntactic structure or grammatical errors in the sentence based on the analysis of the processed sound data and the semantic data adjustment rules during the repair process.
[0143] Specifically, the system sets initial adjustment values for the speech and semantic repair parameters respectively according to the defined speech data adjustment rules and semantic data adjustment rules. In the specific implementation process, the speech data adjustment rules will be based on the analysis of the pronunciation pattern, speech clarity, accent, intonation, pitch, speech speed and other features in the processed sound data, and further determine the specific initial parameters of the model for implementing the repair solution for speech variation, that is, the initial value of speech parameter adjustment; for example, for some unclear pronunciations, the rules may suggest increasing the volume or adjusting the speech frequency to make it clearer; and for the accent or dialect in the speech, rules may be set to weaken these accents so that the listener can understand more easily. Secondly, the semantic data adjustment rules will set the initial direction of semantic repair based on the analysis of the semantic defects of the processed sound data (such as inappropriate vocabulary, grammatical errors, incoherent sentences, etc.); for example, if a sentence is grammatically incorrect, the rules may suggest reconstructing the sentence; if the sentence is ambiguous, the rules may suggest replacing it with more precise words. Finally, the specific initial parameters of the model for implementing the repair solution for semantic defects, that is, the initial value of semantic parameter adjustment, are determined according to the initial direction of semantic repair.
[0144] Step 904, using the voice data adjustment rules and the semantic data adjustment rules as constraints, according to the adjustment strategy correlation information, jointly optimize the initial values of voice parameter adjustment and the initial values of semantic parameter adjustment, respectively, to obtain the determined values of voice parameter adjustment and the determined values of semantic parameter adjustment.
[0145] The speech parameter adjustment determined value may be a model parameter value for adjusting speech variation that is finally determined after joint optimization and verification based on a preliminary adjustment value for speech variation repair.
[0146] The determined value of the semantic parameter adjustment may be the final semantic repair parameter value for adjusting the semantic defect after joint optimization and verification based on the preliminary adjustment value of the semantic defect repair.
[0147] Specifically, the system will adjust the initial values of speech parameter adjustment and semantic parameter adjustment by joint optimization according to the speech data adjustment rules and semantic data adjustment rules. Therefore, the system will not only focus on speech repair (such as pronunciation and sound quality), but also repair semantic problems (such as grammar, vocabulary, etc.) at the same time, and ensure that the repairs between the two can be coordinated. For example, if the pronunciation frequency or pitch needs to be adjusted when repairing speech, this adjustment may affect semantic understanding, especially when the speech pitch is too high or too low, which may cause ambiguity in emotional or semantic transmission.
[0148] In the process of joint optimization, the system uses the voice data adjustment rules and semantic data adjustment rules as constraints, and analyzes the adjustment strategy correlation information to evaluate whether each adjustment will lead to inconsistency or conflict. For example, the adjustment of voice parameters (such as voice pitch, speech rate, volume, etc.) may affect the emotional expression and semantic transmission of the sentence. If the voice pitch is too high or too low, it may change the semantics of the sentence or cause misunderstanding. Therefore, the system must ensure that the voice adjustment does not affect the clarity of the semantics. Semantic adjustments (such as modifying vocabulary and adjusting syntax) may change the fluency or structure of the language, which may also affect the voice performance, such as adjusting the speech rate or stress position; further calculate and combine the coupling between the adjustments of each parameter, and formulate an optimization plan through the optimization algorithm, so as to achieve dual optimization of voice and semantics. Through this joint optimization system, the optimal voice parameter adjustment values and semantic parameter adjustment values can be determined.
[0149] Step 906, using the standard to verify the sound data, performing compliance check and robustness check on the speech parameter adjustment determined values and the semantic parameter adjustment determined values, and obtaining the parameter adjustment value verification result.
[0150] Specifically, standard verification sound data will be used to perform compliance checks and robustness checks on the speech parameter adjustment values and semantic parameter adjustment values. Among them, the compliance check mainly focuses on whether the speech parameter adjustment values and semantic parameter adjustment values meet certain parameter standards, such as whether the parameters are negatively optimized to produce speech distortion or sound quality degradation, and whether the parameter expression is accurate and conforms to grammatical rules. The robustness check tests the performance of the speech parameter adjustment values and semantic parameter adjustment values in different environments and situations, such as whether clear pronunciation can still be guaranteed in a noisy environment, or whether the intention can be correctly conveyed in various contexts. The system will simulate different actual application scenarios to conduct a comprehensive assessment of the compliance check and robustness check of the adjusted speech and semantic parameters to ensure that they can maintain good results in a wide range of usage situations. Through this process, the system can obtain the parameter adjustment value verification results of the parameter adjustment value.
[0151] Step 908, when there is no abnormality in the parameter adjustment value verification result, the initial sound restoration parameters are adjusted according to the speech parameter adjustment determination value and the semantic parameter adjustment determination value to obtain the adjusted sound restoration parameters.
[0152] Specifically, if no abnormalities are found in the compliance check and robustness check of the parameter adjustment value verification results, the system will adjust the initial sound repair parameters according to the final speech parameter adjustment determination values and semantic parameter adjustment determination values. Because if the verification results show that the adjusted parameters can effectively solve the problems of speech variation and semantic defects without side effects, the system applies these adjustment values to the actual initial repair parameters, updates the repair plan, and obtains the adjusted sound repair parameters. At this point, the initial repair parameters have been optimized to the adjusted sound repair parameters, ensuring that a more accurate and natural sound repair effect can be provided in subsequent processing.
[0153] In this embodiment, by accurately analyzing the adjustment rules of speech data and semantic data, the initial adjustment values of speech and semantic parameters are determined, laying the foundation for subsequent optimization. Based on the correlation information of the adjustment strategy, the system jointly optimizes the initial adjustment values of speech and semantics to ensure that the two remain coordinated during the repair process and avoid the adjustment of a single dimension affecting the overall repair effect. Further using standard verification sound data for compliance and robustness checks, the system can effectively verify whether the adjusted parameters meet the predetermined repair standards and ensure the stability and consistency of the repair results in different environments. Finally, when there is no abnormality in the verification results, the repair parameters are adjusted to complete the precise optimization of the initial sound repair parameters. It can achieve more accurate and stable sound repair, improve the quality of the repair results, ensure that the sound achieves the best effect in terms of emotion, context, semantics, etc., and enhance the naturalness and clarity of the auditory experience.
[0154] Based on the same inventive concept, the embodiment of the present application also provides an audio data processing device based on artificial intelligence for implementing the audio data processing method based on artificial intelligence involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the one or more audio data processing device embodiments based on artificial intelligence provided below can refer to the limitations of an audio data processing method based on artificial intelligence above, and will not be repeated here.
[0155] In an exemplary embodiment, Fig.10 As shown, an audio data processing device based on artificial intelligence is provided, including: an audio data acquisition module 1002, a voiceprint feature extraction module 1004, an audio data repair module 1006, a repair parameter adjustment module 1008 and an audio data determination module 1010.
[0156] In an exemplary embodiment, a computer is provided, which may be a server, and its internal structure diagram may be as follows: Fig.11As shown. The computer includes a memory and a processor, the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program. A computer-readable storage medium is also provided, which stores a computer program, and the computer program implements the steps in the above-mentioned method embodiments when executed by the processor. A computer program product or a computer program is also provided, the computer program product or the computer program includes a computer instruction, and the computer instruction is stored in a computer-readable storage medium. The processor of the computer reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer executes the steps in the above-mentioned method embodiments.
Claims
1. An audio data processing method based on artificial intelligence, characterized in that: The method comprises: Acquire the to-be-processed sound data, the object image data and the initial sound restoration parameters of the sound output object; Extracting voiceprint features from the sound data to be processed to obtain object voiceprint feature data; Using the initial sound restoration parameters as restoration constraints, and according to the object voiceprint feature data and the object image data, repairing the speech variation and semantic defects of the sound data to be processed, to obtain processed sound data; When it is detected that the processed sound data still has the speech variation and / or the semantic defect, adjusting the initial sound restoration parameter according to the processed sound data to obtain an adjusted sound restoration parameter; Using the adjusted sound restoration parameters as the initial sound restoration parameters, returning to the step of using the initial sound restoration parameters as restoration constraints, and repairing the speech variation and semantic defects of the sound data to be processed according to the object voiceprint feature data and the object image data to obtain processed sound data; Until the processed sound data cannot be detected to have the phonetic variation and the semantic defect, the processed sound data is used as the target sound data.
2. The method according to claim 1, characterized in that: The method uses the initial sound restoration parameters as restoration constraints, and according to the object voiceprint feature data and the object image data, restores the voice variation and semantic defects of the sound data to be processed to obtain the processed sound data, including: Using the initial sound restoration parameters as restoration constraints, and according to the object voiceprint feature data and the object image data, serially repairing the speech variations and semantic defects of the sound data to be processed, to obtain serially processed sound data; Using the initial sound restoration parameters as restoration constraints, and according to the object voiceprint feature data and the object image data, simultaneously performing parallel restoration on the speech variation and semantic defects of the sound data to be processed, to obtain parallel processed sound data; The serially processed sound data and the parallel processed sound data are fused across pipelines to obtain the processed sound data.
3. The method according to claim 2, characterized in that The method uses the initial sound restoration parameters as restoration constraints, and sequentially restores the voice variations and semantic defects of the sound data to be processed according to the object voiceprint feature data and the object image data to obtain serially processed sound data, including: Using the initial sound restoration parameters as restoration constraints, and adjusting the pronunciation data of the sound data to be processed according to the object voiceprint feature data, to obtain accent-weakened sound data; Inferring the generation environment of the sound data to be processed based on the object image data and the accent-weakened sound data to obtain scene feature data; The language expression data of the accent-weakened sound data is modified according to the scene feature data to obtain the serially processed sound data.
4. The method according to claim 3, characterized in that The step of modifying the language expression data of the accent weakened sound data according to the scene feature data to obtain the serially processed sound data comprises: Performing contextual semantic recognition on the accent-weakened sound data to obtain sound semantic recognition data; According to the sound semantic recognition data, the objective semantic defects of the accent weakened sound data are repaired to obtain semantically repaired sound data; Identify the emotion of the sound output object according to the scene feature data and the object image data to obtain object scene recognition data; The scene semantic defects of the semantic repair sound data are repaired according to the object scene recognition data to obtain the serially processed sound data.
5. The method according to claim 4, characterized in that The step of repairing the scene semantic defects of the semantic repair sound data according to the object scene recognition data to obtain the serially processed sound data includes: Repairing the speech emotion expression of the semantically repaired sound data according to the emotion inference data of the object scene recognition data to obtain emotion enhanced recognition data; Repairing the dialogue subject scene of the emotion enhancement recognition data according to the scene inference data of the object scene recognition data to obtain scene enhancement recognition data; The linguistic features of the scene enhanced recognition data are subjected to simulated neurocognitive optimization to obtain the serially processed sound data.
6. The method according to claim 2, characterized in that The method uses the initial sound restoration parameters as restoration constraints, and according to the object voiceprint feature data and the object image data, simultaneously performs parallel restoration on the voice variation and semantic defects of the sound data to be processed to obtain parallel processed sound data, including: Recognize the emotion of the sound output object according to the object image data to obtain object scene recognition data; Performing context semantic recognition on the sound data to be processed to obtain sound semantic recognition data; Copying and isolating the sound data to be processed to obtain each isolated sound data; Using the initial sound restoration parameters as restoration constraints, and adjusting the pronunciation data of any selected isolated sound data according to the object voiceprint feature data, to obtain accent-weakened sound data; and, based on the sound semantic recognition data, repairing the objective semantic defects of any of the isolated sound data that are not selected to obtain semantically repaired sound data; and, based on the emotion inference data of the object scene recognition data, repairing the speech emotion expression of any of the pronunciation data of the isolated sound data that is not selected, to obtain emotion enhanced recognition data; and, repairing the dialogue subject scene of any of the isolated sound data pronunciation data according to the scene inference data of the object scene recognition data to obtain scene enhanced recognition data; fusing the accent weakened sound data, the semantically restored sound data, the emotion enhanced recognition data, and the scene enhanced recognition data to obtain fused sound restoration data; The linguistic features of the fused sound restoration data are subjected to simulated neurocognitive optimization to obtain the parallel processed sound data.
7. The method according to claim 1, characterized in that The step of adjusting the initial sound restoration parameter according to the processed sound data to obtain the adjusted sound restoration parameter comprises: Determining, from the processed sound data, a voice data adjustment rule corresponding to the voice variation and a semantic data adjustment rule corresponding to the semantic defect; Identifying adjustment strategy correlation information between the voice data adjustment rule and the semantic data adjustment rule; The initial sound restoration parameters are adjusted according to the voice data adjustment rule, the semantic data adjustment rule and the adjustment strategy correlation information to obtain the adjusted sound restoration parameters.
8. The method according to claim 7, characterized in that The adjusting the initial sound restoration parameter according to the voice data adjustment rule, the semantic data adjustment rule and the adjustment strategy correlation information to obtain the adjusted sound restoration parameter includes: Determining an initial value of voice parameter adjustment according to the voice data adjustment rule, and determining an initial value of semantic parameter adjustment according to the semantic data adjustment rule; Taking the voice data adjustment rule and the semantic data adjustment rule as constraint conditions, and according to the adjustment strategy correlation information, jointly optimizing the voice parameter adjustment initial value and the semantic parameter adjustment initial value, respectively, to obtain a voice parameter adjustment determined value and a semantic parameter adjustment determined value; Verify the sound data using the standard, perform compliance check and robustness check on the speech parameter adjustment determination value and the semantic parameter adjustment determination value, and obtain a parameter adjustment value verification result; When there is no abnormality in the parameter adjustment value verification result, the initial sound restoration parameter is adjusted according to the speech parameter adjustment determination value and the semantic parameter adjustment determination value to obtain the adjusted sound restoration parameter.
9. An audio data processing device based on artificial intelligence, characterized in that: The device comprises: An audio data acquisition module, used to acquire the to-be-processed sound data of the sound output object, the object image data and the initial sound restoration parameters; A voiceprint feature extraction module, used to extract voiceprint features from the sound data to be processed to obtain object voiceprint feature data; An audio data repair module, configured to repair the voice variation and semantic defects of the sound data to be processed based on the object voiceprint feature data and the object image data, using the initial sound repair parameters as repair constraints, to obtain processed sound data; A restoration parameter adjustment module, configured to adjust the initial sound restoration parameter according to the processed sound data to obtain an adjusted sound restoration parameter when it is detected that the processed sound data still has the speech variation and / or the semantic defect; The audio data repair module is further used to use the adjusted sound repair parameters as the initial sound repair parameters, return to execute the step of using the initial sound repair parameters as repair constraints, and repairing the voice variation and semantic defects of the sound data to be processed according to the object voiceprint feature data and the object image data to obtain the processed sound data; The audio data determination module is used to use the processed sound data as target sound data until the voice variation and the semantic defect cannot be detected in the processed sound data.
10. A computer comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Voice quality inspection method and device, terminal and computer readable storage medium
CN111210842A
Film restoration method and device based on speech synthesis, equipment and medium
CN113345414A
Chinese speech enhancement recognition and text error correction and correction method
CN115602161A
Voice processing method and device, electronic equipment and computer readable medium
CN117174067A
Audio processing method and system
CN117524259A
Cited By
Voice instruction dynamic acoustic simulation test method and system
CN120510868A
Target voice regulation and control method, device and equipment based on intelligent glasses
CN120544552A
Multi-language voice test method and device, computer equipment and storage medium
CN120932685A