An intelligent simulation method and system based on neural algorithm multi-source audio features

By employing an intelligent simulation method based on multi-source audio features using neural algorithms, and utilizing MPEG audio acquisition equipment and Gaussian neural networks, accurate identification and quality assessment of audio source instruments were achieved. This solves the problems of limited coverage and high resource consumption in existing technologies, and improves the stability and data availability of the simulation model.

CN116705060BActive Publication Date: 2026-05-29HEAD DIRECT (KUNSHAN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEAD DIRECT (KUNSHAN) CO LTD
Filing Date
2023-04-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing audio simulation methods cannot effectively cover the composition of test subjects, the total number of test subjects is limited, the experimental process consumes a lot of manpower and resources, and traditional experimental environments lack stability and practicality, which cannot meet the needs of audio research.

Method used

An intelligent simulation method based on multi-source audio features using neural algorithms is adopted. User metadata information and audio frames are collected through MPEG audio acquisition equipment, and users are identified by using Gaussian neural networks. Combined with the adaptive adjustment strategy of the simulation model, the accurate identification and quality assessment of the audio source instruments are achieved.

Benefits of technology

It achieves accurate identification and quality assessment of audio source instruments, improves the recall and stability of simulation models, reduces manpower and material resources consumption, and increases the coverage and data availability of experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116705060B_ABST
    Figure CN116705060B_ABST
Patent Text Reader

Abstract

The application claims a kind of intelligent simulation method and system based on neural algorithm multi-source audio feature, by collecting simulation audio source data, based on Gaussian neural network method, the identification user of current frame is extracted, the real-time average user metadata information in the preselected pitch center and interruption segment is calculated as the basis for determining the recall information of simulation model, then the real-time recall of simulation model is determined, the adaptive adjustment strategy of real-time recall of simulation model is combined, the corresponding position of simulation model and audio source data is matched;Finally, the periodic manager of audio source instrument feeds back the audio source instrument through timing signal and stores in the audio source instrument white list library in the audio receiver end.The scheme accurately identifies different levels of audio source instrument content by combining the quality performance value of accurate audio frame identification, to achieve the effect of adaptive audio source instrument level accurate collection of audio source instrument.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of user identification technology and multimedia technology, specifically to an intelligent simulation method and system based on multi-source audio features of neural algorithms. Background Technology

[0002] The 21st century is an era of rapid internet development. With the widespread adoption of the internet, the demand for transmitting audio signals over the network has increased significantly. The emergence of streaming media technology has, to some extent, alleviated the difficulties of transmitting audio over the internet, transforming the traditional "push" style of media dissemination into a "pull" style of audience engagement and real-time transmission. Because streaming media technology has, to a certain extent, overcome the limitations of network bandwidth on multimedia information transmission, it has been widely used in various fields such as online live streaming, online conferencing, distance education, and corporate training. To ensure better streaming media transmission, it is usually necessary to evaluate the quality of streaming media, which also presents new challenges for evaluating the quality of streaming media audio.

[0003] Existing simulation methods impose limitations on the time and location of participants due to the requirement for concentrated testing at specific times and locations. This restricts the selection of participants and prevents the coverage of all necessary components. Furthermore, the limited duration of the experiment restricts the total number of participants, resulting in insufficient usable data from subjective simulations. Currently, participant eligibility requires manual verification, which may lead to ineligible participants and unusable data. The experimental process requires extensive supervision and operation by numerous staff, consuming significant human and material resources. Raw test data must be manually entered into computers, increasing the risk of data entry errors.

[0004] Furthermore, existing network simulation software such as NIST Net and NS2 focus solely on the characteristics of the network itself, lacking settings for voice transmission and encoding. Therefore, relying solely on these simulations is insufficient for audio research. Experiments using real networks to study the impact of audio frame parameters on audio quality face challenges in precisely setting these parameters, resulting in a lack of repeatability and requiring expensive equipment such as routers and switches, along with related hardware and software support. Traditional experimental environments fail to consider precise control over the waveform, volume, and recording methods of terminal devices, nor do they account for the impact of latency on the quality of the interactive audio experience, rendering the experimental platform unstable and impractical. Summary of the Invention

[0005] This invention provides an intelligent simulation method and system based on multi-source audio features using neural algorithms. It solves the problem of existing audio source instrument detection and recognition methods erroneously detecting non-user simulated line audio source instruments, thereby effectively solving the problem of audio source instrument simulation robots erroneously executing to simulate non-user simulated line audio source instruments, resulting in simulation failure, as well as the distortion problem between the end-effector and the user simulated line audio source instruments.

[0006] According to a first aspect of the present invention, the present invention claims protection for an intelligent simulation method based on multi-source audio features of a neural algorithm, characterized by comprising the following steps:

[0007] Audio source instrument data acquisition: During the operation of the audio source instrument, the user metadata information and MPEG source encoded audio frames of the audio source data at the current position of the audio acquisition device are simulated using an MPEG audio acquisition device. The user identification of the current frame is extracted based on the Gaussian neural network method, and the quality level audio frames of the audio source data are acquired.

[0008] User identification involves calculating the real-time average user metadata information within the pre-selected pitch center and interrupted segments of the quality-level audio frames, and using the real-time average user metadata information as the basis for determining the recall information of the simulation model in subsequent steps.

[0009] The simulation operation combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic, and combined with the adaptive adjustment strategy of the real-time recall of the simulation model, the corresponding positions of the simulation model and the audio source data are matched.

[0010] The audio source instrument simulation determines the result based on the recall of the simulation model. The audio source instrument's cycle manager feeds back the audio source instrument through timing signals and stores the simulated audio source instrument in the audio source instrument whitelist library at the audio receiver end through a trusted database.

[0011] Specifically, the audio source instrument data acquisition involves using an MPEG audio acquisition device to simulate user metadata information and MPEG source-encoded audio frames of the audio source data at the current location of the audio acquisition device during the instrument's operation. Based on a Gaussian neural network method, the user of the current frame is identified, and the quality level audio frames of the audio source data are acquired, including:

[0012] The sample sound detection method based on cepstral coefficients and pyramidal algorithm collects the audio frame coordinates of sample sounds from audio frames in audio source data;

[0013] Among them, by installing a pitch audio acquisition device on the simulation model, it is possible to collect audio frames of different quality levels of audio source instruments and background in real time from the audio source data.

[0014] When the pitch audio acquisition device is vertically receiving the audio source instrument within the preset range, the audio source instrument is arranged at multiple angles. The abnormal pitch values ​​returned by the blank areas between the audio source instruments and the excessively small pitch values ​​returned by the sample sounds with large differences in volume peak and valley values ​​are filtered out to obtain the average user metadata information of the sample sounds on the audio source data.

[0015] The sample sound detection algorithm using cepstral coefficients and pyramidal algorithms achieves noise reduction and separation between the simulated sound and the sample sound in the background, and optimizes the average user metadata information of the obtained sample sound.

[0016] Specifically, user identification involves calculating real-time average user metadata information within pre-selected pitch centers and interrupted segments of audio frames of different quality levels. This real-time average user metadata information is used as the basis for determining the recall information of the simulation model in subsequent steps. This includes:

[0017] The Gaussian neural network method based on user metadata information extracts sample sound regions according to the pitch weights of each region cluster;

[0018] Among them, the input quality level audio frames are divided into K region clusters based on the spectrum clustering algorithm;

[0019] Calculate the initial significance value of the region cluster k in the pitch audio frame;

[0020] The mono prior is adjusted to the new user metadata information weight, the initial saliency value and the stereo mapping are corrected, and the corrected saliency value is obtained.

[0021] After obtaining the corrected saliency value, the average pitch of the output salient user region is combined with the sample sound separated and collected in the MPEG space and the corresponding position coordinates to obtain the position-pitch integrated information of the sample sound.

[0022] In the subsequent feedback recall calculation step, the position-pitch integrated information will be used as the initial input information to perform feature adaptive simulation of the audio source data by the simulation model.

[0023] Specifically, the simulation operation combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic and an adaptive adjustment strategy for the real-time recall of the simulation model, the corresponding positions of the simulation model and the audio source data are matched. This includes:

[0024] Fixed pitch audio acquisition device location: The relative position between the fixed pitch audio acquisition device and the user's audio output, and the relative position data between the pitch audio acquisition device and the simulation model are also acquired.

[0025] The performance value of the audio source instrument is obtained by acquiring video frames of audio source data at the simulation model using a pitch audio acquisition device. The video frame data is three-dimensional spatial data including user metadata information. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is calculated based on the video frame data. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is the performance value of the audio source instrument.

[0026] Adjusting the recall state of the simulation model: Compare the performance value of the source instrument with the expected performance value of the source instrument, and adjust the position recall state of the simulation model so that the performance value of the source instrument meets the requirements of the expected performance value of the source instrument. The adjustment of the position recall state of the simulation model includes adjusting the simulation model to increase or decrease.

[0027] Specifically, the audio source instrument simulation determines the result based on the recall of the simulation model. The audio source instrument's cycle manager feeds back the audio source instrument via timing signals and stores the simulated audio source instrument in the audio receiver's whitelist database using a trusted database. This includes:

[0028] The whitelist of musical instruments for audio sources includes at least a first whitelist, a second whitelist, and a third whitelist.

[0029] The grades should include at least a low quality grade, a satisfactory quality grade, and a high quality grade.

[0030] The audio source instruments are all turned off when the simulation model is not performing feedback operations;

[0031] Based on the recall determination results of the simulation model, when the simulation model simulates low-quality audio source instruments and the recall determination results of the simulation model are adjusted to indicate a decrease in user pitch, the first whitelist library is closed and the second whitelist library is opened.

[0032] When the simulation model simulates an audio source instrument of acceptable quality level and the recall result of the simulation model is adjusted to indicate a decrease in user pitch, the second whitelist library is closed and the third whitelist library is opened.

[0033] When the simulation model simulates low-quality audio source instruments and the recall determination result of the simulation model is adjusted to the user's pitch increase, the first whitelist library is closed.

[0034] When the simulation model simulates an audio source instrument of acceptable quality level and the recall result of the simulation model is adjusted to the user's pitch increase, the second whitelist library is closed and the first whitelist library is opened.

[0035] When the simulation model simulates a high-quality audio source instrument and the recall result of the simulation model is adjusted to indicate a decrease in user pitch, the third whitelist library is turned off.

[0036] When the simulation model simulates a high-quality audio source instrument and the recall result of the simulation model is adjusted to the user's pitch increase, the third whitelist library is closed and the second whitelist library is opened.

[0037] According to a second aspect of the present invention, the present invention claims protection for an intelligent simulation system based on multi-source audio features of a neural algorithm, comprising:

[0038] The audio source instrument data acquisition module uses an MPEG audio acquisition device to simulate user metadata information and MPEG source encoded audio frames of the audio source data at the current position of the audio acquisition device during the operation of the audio source instrument. It extracts the user identification of the current frame based on the Gaussian neural network method and acquires the quality level audio frames of the audio source data.

[0039] The user identification module calculates the real-time average user metadata information within the pre-selected pitch center and interrupted segments of the quality-level audio frames, and uses the real-time average user metadata information as the basis for determining the recall information of the simulation model in subsequent steps.

[0040] The simulation operation module combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic and combined with the adaptive adjustment strategy of the real-time recall of the simulation model, it matches the corresponding positions of the simulation model and the audio source data.

[0041] The audio source instrument simulation module determines the result based on the recall of the simulation model. The audio source instrument period manager feeds back the audio source instrument through timing signals and stores the simulated audio source instrument in the audio source instrument whitelist library at the audio receiver end through a trusted database.

[0042] Specifically, the audio source instrument data acquisition module includes:

[0043] The sample sound detection method based on cepstral coefficients and pyramidal algorithm collects the audio frame coordinates of sample sounds from audio frames in audio source data;

[0044] Among them, by installing a pitch audio acquisition device on the simulation model, it is possible to collect audio frames of different quality levels of audio source instruments and background in real time from the audio source data.

[0045] When the pitch audio acquisition device is vertically receiving the audio source instrument within the preset range, the audio source instrument is arranged at multiple angles. The abnormal pitch values ​​returned by the blank areas between the audio source instruments and the excessively small pitch values ​​returned by the sample sounds with large differences in volume peak and valley values ​​are filtered out to obtain the average user metadata information of the sample sounds on the audio source data.

[0046] The sample sound detection algorithm using cepstral coefficients and pyramidal algorithms achieves noise reduction and separation between the simulated sound and the sample sound in the background, and optimizes the average user metadata information of the obtained sample sound.

[0047] Specifically, the user identification module includes:

[0048] The Gaussian neural network method based on user metadata information extracts sample sound regions according to the pitch weights of each region cluster;

[0049] Among them, the input quality level audio frames are divided into K region clusters based on the spectrum clustering algorithm;

[0050] Calculate the initial significance value of the region cluster k in the pitch audio frame;

[0051] The mono prior is adjusted to the new user metadata information weight, the initial saliency value and the stereo mapping are corrected, and the corrected saliency value is obtained.

[0052] After obtaining the corrected saliency value, the average pitch of the output salient user region is combined with the sample sound separated and collected in the MPEG space and the corresponding position coordinates to obtain the position-pitch integrated information of the sample sound.

[0053] In the subsequent feedback recall calculation step, the position-pitch integrated information will be used as the initial input information to perform feature adaptive simulation of the audio source data by the simulation model.

[0054] Specifically, the simulation operation module includes:

[0055] Fixed pitch audio acquisition device location: The relative position between the fixed pitch audio acquisition device and the user's audio output, and the relative position data between the pitch audio acquisition device and the simulation model are also acquired.

[0056] The performance value of the audio source instrument is obtained by acquiring video frames of audio source data at the simulation model using a pitch audio acquisition device. The video frame data is three-dimensional spatial data including user metadata information. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is calculated based on the video frame data. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is the performance value of the audio source instrument.

[0057] Adjusting the recall state of the simulation model: Compare the performance value of the source instrument with the expected performance value of the source instrument, and adjust the position recall state of the simulation model so that the performance value of the source instrument meets the requirements of the expected performance value of the source instrument. The adjustment of the position recall state of the simulation model includes adjusting the simulation model to increase or decrease.

[0058] Specifically, the audio source instrument simulation module includes:

[0059] The whitelist of musical instruments for audio sources includes at least a first whitelist, a second whitelist, and a third whitelist.

[0060] The grades should include at least a low quality grade, a satisfactory quality grade, and a high quality grade.

[0061] The audio source instruments are all turned off when the simulation model is not performing feedback operations;

[0062] Based on the recall determination results of the simulation model, when the simulation model simulates low-quality audio source instruments and the recall determination results of the simulation model are adjusted to indicate a decrease in user pitch, the first whitelist library is closed and the second whitelist library is opened.

[0063] When the simulation model simulates an audio source instrument of acceptable quality level and the recall result of the simulation model is adjusted to indicate a decrease in user pitch, the second whitelist library is closed and the third whitelist library is opened.

[0064] When the simulation model simulates low-quality audio source instruments and the recall determination result of the simulation model is adjusted to the user's pitch increase, the first whitelist library is closed.

[0065] When the simulation model simulates an audio source instrument of acceptable quality level and the recall result of the simulation model is adjusted to the user's pitch increase, the second whitelist library is closed and the first whitelist library is opened.

[0066] When the simulation model simulates a high-quality audio source instrument and the recall result of the simulation model is adjusted to indicate a decrease in user pitch, the third whitelist library is turned off.

[0067] When the simulation model simulates a high-quality audio source instrument and the recall result of the simulation model is adjusted to the user's pitch increase, the third whitelist library is closed and the second whitelist library is opened.

[0068] This invention claims protection for an intelligent simulation method and system based on multi-source audio features using neural algorithms. During the operation of an audio source instrument, the method simulates user metadata information and MPEG source-encoded audio frames from the audio source data. It extracts the identified user for the current frame using a Gaussian neural network, collects quality-level audio frames, and calculates the real-time average user metadata information within pre-selected pitch centers and interrupted segments as the basis for determining the recall information of the simulation model. Then, it determines the real-time recall of the simulation model and combines it with an adaptive adjustment strategy to match the corresponding positions of the simulation model and the audio source data. Finally, the audio source instrument's cycle manager feeds back the audio source instrument via timing signals and stores the simulated audio source instrument in a whitelist database at the audio receiver end through a trusted database. This scheme accurately identifies audio source instrument content of different levels by combining accurate audio frame recognition with coarse-grained quality performance values, achieving the effect of accurately collecting audio source instruments based on adaptive audio source instrument levels. Attached Figure Description

[0069] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0070] Figure 1 This is a flowchart illustrating the workflow of an intelligent simulation method based on multi-source audio features using neural algorithms, as described in this invention.

[0071] Figure 2 This is the second workflow diagram of an intelligent simulation method based on multi-source audio features using neural algorithms, which is involved in this invention.

[0072] Figure 3 This is the third workflow diagram of an intelligent simulation method based on multi-source audio features using neural algorithms, which is involved in this invention.

[0073] Figure 4 This is the fourth workflow diagram of an intelligent simulation method based on multi-source audio features using neural algorithms, which is involved in this invention.

[0074] Figures 5a-5e This is a schematic diagram illustrating the operation of an intelligent simulation method based on multi-source audio features using neural algorithms, as described in this invention.

[0075] Figure 6 This is a structural block diagram of an intelligent simulation system based on multi-source audio features using neural algorithms, which is involved in this invention. Detailed Implementation

[0076] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0077] According to a first embodiment of the present invention, referring to the appendix Figure 1 This invention claims protection for an intelligent simulation method based on multi-source audio features using neural algorithms, characterized by comprising the following steps:

[0078] Audio source instrument data acquisition: During the operation of the audio source instrument, the user metadata information and MPEG source encoded audio frames of the audio source data at the current position of the audio acquisition device are simulated using an MPEG audio acquisition device. The user identification of the current frame is extracted based on the Gaussian neural network method, and the quality level audio frames of the audio source data are acquired.

[0079] User identification involves calculating the real-time average user metadata information within the pre-selected pitch center and interrupted segments of the quality-level audio frames, and using the real-time average user metadata information as the basis for determining the recall information of the simulation model in subsequent steps.

[0080] The simulation operation combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic, and combined with the adaptive adjustment strategy of the real-time recall of the simulation model, the corresponding positions of the simulation model and the audio source data are matched.

[0081] The audio source instrument simulation determines the result based on the recall of the simulation model. The audio source instrument's cycle manager feeds back the audio source instrument through timing signals and stores the simulated audio source instrument in the audio source instrument whitelist library at the audio receiver end through a trusted database.

[0082] In this embodiment, for a large number of musical instruments, some or all of the instruments can be selected as user instruments based on factors such as the technology, business, and playback effect in conjunction with animation for detecting average pitch simulation. The simulation recognition model is used in advance to detect the average pitch simulation of these user instruments during audio data operation, and the attribute information of these user instruments in terms of time, type, intensity, frequency, energy, etc. is obtained. This attribute information is recorded in the simulation file, and the simulation file is associated with the audio data. That is, the average pitch simulation has attribute information obtained by identifying the audio feature information of the audio data based on the simulation recognition model.

[0083] For scenarios where audio data is played for the first time, a simulation file associated with the audio data can be requested from the audio receiver, thereby reading the average pitch simulation of the user's instrument during operation from the simulation file.

[0084] If the audio data is online audio data provided by the audio receiver, the audio player can send the ID of the audio data to the audio receiver. The audio receiver can then use this ID to retrieve the simulation file associated with the audio data and send it to the audio player.

[0085] If the audio data is local audio data provided by a computer device, the audio player can send the name, audio fingerprint (such as a hash value), and other identifiers of the audio data to the audio receiver. The audio receiver uses these identifiers to check if a simulation file for the audio data exists. If it does, the receiver sends the simulation file to the audio player. If it does not, the receiver can request the audio player to upload the audio data, and then simulate the average pitch of the user's instrument during operation. The receiver uses the data's attribute information to create a corresponding simulation file and sends the simulation file to the audio player.

[0086] Experiments show that, following human auditory perception, the performance (such as recall and precision) of the simulation recognition model trained with highly simulated songs (i.e., the second audio data) is basically the same as that of the simulation recognition model trained with randomly generated songs (i.e., the second audio data). In order to save resources, songs (i.e., the second audio data) can be generated by random interleaving.

[0087] For details, please refer to the appendix. Figure 2 The audio source instrument data acquisition process involves using an MPEG audio acquisition device to simulate user metadata information and MPEG source-encoded audio frames at the current location of the audio source data during the instrument's operation. Based on a Gaussian neural network method, the user identity of the current frame is extracted, and the quality level audio frames of the audio source data are acquired. Specifically, this includes:

[0088] The sample sound detection method based on cepstral coefficients and pyramidal algorithm collects the audio frame coordinates of sample sounds from audio frames in audio source data;

[0089] Among them, by installing a pitch audio acquisition device on the simulation model, it is possible to collect audio frames of different quality levels of audio source instruments and background in real time from the audio source data.

[0090] When the pitch audio acquisition device is vertically receiving the audio source instrument within the preset range, the audio source instrument is arranged at multiple angles. The abnormal pitch values ​​returned by the blank areas between the audio source instruments and the excessively small pitch values ​​returned by the sample sounds with large differences in volume peak and valley values ​​are filtered out to obtain the average user metadata information of the sample sounds on the audio source data.

[0091] The sample sound detection algorithm using cepstral coefficients and pyramidal algorithms achieves noise reduction and separation between the simulated sound and the sample sound in the background, and optimizes the average user metadata information of the obtained sample sound.

[0092] In this embodiment, a cepstral coefficient sequence of the audio signal to be detected is obtained. The cepstral coefficient sequence is an N-dimensional vector, where N is the window length of the window function used to window the audio signal to be detected. Each element in the cepstral coefficient sequence is used to characterize the cepstral coefficients of each sampling point.

[0093] Based on the cepstral coefficient sequence, the low-energy spectral band in the audio signal to be detected is determined;

[0094] Based on the low-energy spectrum band, it is determined whether the audio signal to be detected has a frequency band loss. If it is determined that the audio signal to be detected has a frequency band loss, then it is determined that the audio signal to be detected is distorted.

[0095] A threshold-based denoising method can be used to separate the sample sound from the simulated sound. The denoising result depends primarily on the threshold value. Considering real-time performance and simplification, a pyramid algorithm is chosen to automatically determine the denoising threshold for each grayscale audio frame of the audio source data. After the threshold is determined, the grayscale audio frames are binarized, and the sample sound denoising is essentially complete. However, some outliers, influenced by illumination, are still incorrectly retained in the denoising results. These outliers can be removed using morphological erosion, and the coordinates of the retained audio frame can be collected.

[0096] At this point, most of the sample sounds have been successfully separated from the MPEG audio frames of the audio source data through noise reduction. To ensure the accuracy of recognition, secondary processing is required based on the pitch audio frames of the audio source data, and this processing is combined with the noise reduction results to collect more accurate sample sound feedback pitch values.

[0097] For details, please refer to the appendix. Figure 3 User identification involves calculating real-time average user metadata information within pre-selected pitch centers and interrupted segments of audio frames of different quality levels. This real-time average user metadata information is used as the basis for determining the recall information of the simulation model in subsequent steps. Specifically, this includes:

[0098] The Gaussian neural network method based on user metadata information extracts sample sound regions according to the pitch weights of each region cluster;

[0099] Among them, the input quality level audio frames are divided into K region clusters based on the spectrum clustering algorithm;

[0100] Calculate the initial significance value of the region cluster k in the pitch audio frame;

[0101] The mono prior is adjusted to the new user metadata information weight, the initial saliency value and the stereo mapping are corrected, and the corrected saliency value is obtained.

[0102] After obtaining the corrected saliency value, the average pitch of the output salient user region is combined with the sample sound separated and collected in the MPEG space and the corresponding position coordinates to obtain the position-pitch integrated information of the sample sound.

[0103] In the subsequent feedback recall calculation step, the position-pitch integrated information will be used as the initial input information to perform feature adaptive simulation of the audio source data by the simulation model.

[0104] In this embodiment, a Gaussian neural network method based on user metadata information is introduced to achieve high efficiency and robustness in the simulation of the audio source instrument. Since the user metadata information of the audio source instrument directly affects the feature adaptive adjustment of the machine during operation, the sample sound detection results are made more accurate by adding a new feature suppression end for optimization.

[0105] The input MPEG audio frames are clustered based on the spectrum clustering algorithm. The audio is divided into K region clusters. Combining the average pitch of each region's audio frames within the corresponding pitch audio frames, the region's pitch significance value for that audio frame is calculated as follows:

[0106]

[0107] in, For audio frames The significance value of pitch k in the middle region, It is the average Euclidean difference between region k and region i within the pitch space. This represents the ratio of the average pitch value of region k to the total pitch value of the entire audio frame. To further highlight the pitch differences between different region clusters, this project assigns pitch weights to each region within the corresponding pitch audio frame:

[0108]

[0109] in, This is the pitch weight assigned to region k. Indicates Gaussian normalization, It represents the maximum pitch value among all audio frames within the entire pitch audio frame. It is the average pitch value of the audio frames in region k. It is a fixed pitch value, set to

[0110]

[0111] in, This represents the minimum pitch value among all audio frames within the entire pitch audio frame. Based on the above parameters, the initial significance value of region clustering k in the pitch audio frame is calculated:

[0112]

[0113] This algorithm optimizes the center-to-binocular prior theory by adjusting the monocular prior to a new weighted user metadata information. This scheme applies the modification to the Gaussian neural network method and represents the improved binocular mapping as follows: Based on equations (2) and (3), the initial saliency value and the dual-channel mapping are further modified as follows:

[0114]

[0115] in This represents the significant value obtained after correction.

[0116] After obtaining the detection results, the scheme combines the average pitch of the output significant user region with the sample sound and its corresponding position coordinates, which were denoised and separated in MPEG space. This results in the position-pitch integrated information of the sample sound. In the subsequent feedback recall calculation step, this integrated information serves as the initial input, and based on this, the user pitch is adaptively simulated to the characteristics of the audio source data.

[0117] For details, please refer to the appendix. Figure 4 The simulation operation combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic and an adaptive adjustment strategy for the real-time recall of the simulation model, the corresponding positions of the simulation model and the audio source data are matched. Specifically, this includes:

[0118] Fixed pitch audio acquisition device location: The relative position between the fixed pitch audio acquisition device and the user's audio output, and the relative position data between the pitch audio acquisition device and the simulation model are also acquired.

[0119] The performance value of the audio source instrument is obtained by acquiring video frames of audio source data at the simulation model using a pitch audio acquisition device. The video frame data is three-dimensional spatial data including user metadata information. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is calculated based on the video frame data. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is the performance value of the audio source instrument.

[0120] The audio source data MPEG audio frames are captured in real time by a pitch audio acquisition device. After audio frame processing to reduce the influence of the simulated sound and background, the average pitch of the sample sound in two fixed regions of the audio frame is obtained. and Used for calculating feedback recall to ensure the accuracy of feedback.

[0121] In this embodiment, refer to the appendix Figures 5a-5e The audio source instruments can include violins, recorders, guitars, pianos, and drum kits. The audio data from the performance of these instruments is sent to the audio acquisition device. During operation, the audio acquisition device collects audio frames of the audio source data at a frame rate of 10 frames per second. For each audio frame, the device judges and calculates the motion of the features. With real-time adjustments to the feedback recall, the features adaptively adjust their height and angle.

[0122] In this embodiment, when the percentage range of qualified audio frames in the simulation model is below the first percentage range of qualified audio frames in the audio source data and the performance value of the audio source instrument is greater than the first performance value, the audio source instrument is identified as a high-quality level.

[0123] When the percentage range of qualified audio frames in the simulation model is below the second percentage and above the first percentage of qualified audio frames in the audio source data, and the performance value of the audio source instrument is greater than the second performance value and less than the first performance value, the audio source instrument is considered to be of qualified quality level.

[0124] Specifically, the first ratio, the second ratio, the first performance value, and the second performance value are set according to the different types of instruments and data of different audio sources;

[0125] Specifically, the audio source instrument simulation determines the result based on the recall of the simulation model. The audio source instrument's cycle manager feeds back the audio source instrument via timing signals and stores the simulated audio source instrument in the audio receiver's whitelist database using a trusted database. This includes:

[0126] The whitelist of musical instruments for audio sources includes at least a first whitelist, a second whitelist, and a third whitelist.

[0127] The grades should include at least a low quality grade, a satisfactory quality grade, and a high quality grade.

[0128] The audio source instruments are all turned off when the simulation model is not performing feedback operations;

[0129] Based on the recall determination results of the simulation model, when the simulation model simulates low-quality audio source instruments and the recall determination results of the simulation model are adjusted to indicate a decrease in user pitch, the first whitelist library is closed and the second whitelist library is opened.

[0130] When the simulation model simulates an audio source instrument of acceptable quality level and the recall result of the simulation model is adjusted to indicate a decrease in user pitch, the second whitelist library is closed and the third whitelist library is opened.

[0131] When the simulation model simulates low-quality audio source instruments and the recall determination result of the simulation model is adjusted to the user's pitch increase, the first whitelist library is closed.

[0132] When the simulation model simulates an audio source instrument of acceptable quality level and the recall result of the simulation model is adjusted to the user's pitch increase, the second whitelist library is closed and the first whitelist library is opened.

[0133] When the simulation model simulates a high-quality audio source instrument and the recall result of the simulation model is adjusted to indicate a decrease in user pitch, the third whitelist library is turned off.

[0134] When the simulation model simulates a high-quality audio source instrument and the recall result of the simulation model is adjusted to the user's pitch increase, the third whitelist library is closed and the second whitelist library is opened.

[0135] Specifically, based on the height of the audio source data, the height of the simulation model of the audio source data, and the performance values ​​of the audio source instruments, the acceptable quality level and high quality level of the audio source instruments are determined, as well as the low quality level determined by audio frame recognition. After determining the recall result of the simulation model, the audio source data is generally classified from top to bottom as low quality level, acceptable quality level, and high quality level. This scheme performs a coarse classification by determining whether the recall result of the simulation model is upward or downward. Then, it performs a fine classification by determining the acceptable quality level and high quality level of the audio source instruments based on the height of the simulation model of the audio source data and the performance values ​​of the audio source instruments. Audio source instruments that should belong to the acceptable quality level or high quality level but do not belong to this type of level are removed, achieving a more accurate collection of audio source instruments.

[0136] According to a second embodiment of the present invention, referring to the appendix Figure 6 This invention claims protection for an intelligent simulation system based on multi-source audio features using neural algorithms, characterized by comprising:

[0137] The audio source instrument data acquisition module uses an MPEG audio acquisition device to simulate user metadata information and MPEG source encoded audio frames of the audio source data at the current position of the audio acquisition device during the operation of the audio source instrument. It extracts the user identification of the current frame based on the Gaussian neural network method and acquires the quality level audio frames of the audio source data.

[0138] The user identification module calculates the real-time average user metadata information within the pre-selected pitch center and interrupted segments of the quality-level audio frames, and uses the real-time average user metadata information as the basis for determining the recall information of the simulation model in subsequent steps.

[0139] The simulation operation module combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic and combined with the adaptive adjustment strategy of the real-time recall of the simulation model, it matches the corresponding positions of the simulation model and the audio source data.

[0140] The audio source instrument simulation module determines the result based on the recall of the simulation model. The audio source instrument period manager feeds back the audio source instrument through timing signals and stores the simulated audio source instrument in the audio source instrument whitelist library at the audio receiver end through a trusted database.

[0141] Specifically, the audio source instrument data acquisition module includes:

[0142] The sample sound detection method based on cepstral coefficients and pyramidal algorithm collects the audio frame coordinates of sample sounds from audio frames in audio source data;

[0143] Among them, by installing a pitch audio acquisition device on the simulation model, it is possible to collect audio frames of different quality levels of audio source instruments and background in real time from the audio source data.

[0144] When the pitch audio acquisition device is vertically receiving the audio source instrument within the preset range, the audio source instrument is arranged at multiple angles. The abnormal pitch values ​​returned by the blank areas between the audio source instruments and the excessively small pitch values ​​returned by the sample sounds with large differences in volume peak and valley values ​​are filtered out to obtain the average user metadata information of the sample sounds on the audio source data.

[0145] The sample sound detection algorithm using cepstral coefficients and pyramidal algorithms achieves noise reduction and separation between the simulated sound and the sample sound in the background, and optimizes the average user metadata information of the obtained sample sound.

[0146] Specifically, the user identification module includes:

[0147] The Gaussian neural network method based on user metadata information extracts sample sound regions according to the pitch weights of each region cluster;

[0148] Among them, the input quality level audio frames are divided into K region clusters based on the spectrum clustering algorithm;

[0149] Calculate the initial significance value of the region cluster k in the pitch audio frame;

[0150] The mono prior is adjusted to the new user metadata information weight, the initial saliency value and the stereo mapping are corrected, and the corrected saliency value is obtained.

[0151] After obtaining the corrected saliency value, the average pitch of the output salient user region is combined with the sample sound separated and collected in the MPEG space and the corresponding position coordinates to obtain the position-pitch integrated information of the sample sound.

[0152] In the subsequent feedback recall calculation step, the position-pitch integrated information will be used as the initial input information to perform feature adaptive simulation of the audio source data by the simulation model.

[0153] Specifically, the simulation operation module includes:

[0154] Fixed pitch audio acquisition device location: The relative position between the fixed pitch audio acquisition device and the user's audio output, and the relative position data between the pitch audio acquisition device and the simulation model are also acquired.

[0155] The performance value of the audio source instrument is obtained by acquiring video frames of audio source data at the simulation model using a pitch audio acquisition device. The video frame data is three-dimensional spatial data including user metadata information. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is calculated based on the video frame data. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is the performance value of the audio source instrument.

[0156] Adjusting the recall state of the simulation model: Compare the performance value of the source instrument with the expected performance value of the source instrument, and adjust the position recall state of the simulation model so that the performance value of the source instrument meets the requirements of the expected performance value of the source instrument. The adjustment of the position recall state of the simulation model includes adjusting the simulation model to increase or decrease.

[0157] Specifically, the audio source instrument simulation module includes:

[0158] The whitelist of musical instruments for audio sources includes at least a first whitelist, a second whitelist, and a third whitelist.

[0159] The grades should include at least a low quality grade, a satisfactory quality grade, and a high quality grade.

[0160] The audio source instruments are all turned off when the simulation model is not performing feedback operations;

[0161] Based on the recall determination results of the simulation model, when the simulation model simulates low-quality audio source instruments and the recall determination results of the simulation model are adjusted to indicate a decrease in user pitch, the first whitelist library is closed and the second whitelist library is opened.

[0162] When the simulation model simulates an audio source instrument of acceptable quality level and the recall result of the simulation model is adjusted to indicate a decrease in user pitch, the second whitelist library is closed and the third whitelist library is opened.

[0163] When the simulation model simulates low-quality audio source instruments and the recall determination result of the simulation model is adjusted to the user's pitch increase, the first whitelist library is closed.

[0164] When the simulation model simulates an audio source instrument of acceptable quality level and the recall result of the simulation model is adjusted to the user's pitch increase, the second whitelist library is closed and the first whitelist library is opened.

[0165] When the simulation model simulates a high-quality audio source instrument and the recall result of the simulation model is adjusted to indicate a decrease in user pitch, the third whitelist library is turned off.

[0166] When the simulation model simulates a high-quality audio source instrument and the recall result of the simulation model is adjusted to the user's pitch increase, the third whitelist library is closed and the second whitelist library is opened.

[0167] Those skilled in the art will understand that the contents disclosed herein can be varied and modified in many ways. For example, the various devices or components described above can be implemented in hardware, or in software, firmware, or a combination of some or all of the three.

[0168] This disclosure uses flowcharts to illustrate the steps of a method according to embodiments of this disclosure. It should be understood that the preceding or following steps are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes.

[0169] Those skilled in the art will understand that all or part of the steps in the above methods can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk. Optionally, all or part of the steps in the above embodiments can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiments can be implemented in hardware or as a software functional module. This disclosure is not limited to any particular combination of hardware and software.

[0170] Unless otherwise defined, all terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that terms such as those defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and not as having an idealized or highly formalized meaning, unless expressly defined herein.

[0171] The foregoing description is intended to illustrate the present disclosure and should not be construed as limiting it. While several exemplary embodiments of the present disclosure have been described, those skilled in the art will readily understand that many modifications may be made to the exemplary embodiments without departing from the novel teachings and advantages of the present disclosure. Therefore, all such modifications are intended to be included within the scope of the present disclosure as defined by the claims. It should be understood that the foregoing description is intended to illustrate the present disclosure and should not be construed as limiting it to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The present disclosure is defined by the claims and their equivalents.

[0172] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0173] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, adjustments and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. An intelligent simulation method based on multi-source audio features using neural algorithms, characterized in that, Including the following steps: Audio source instrument data acquisition: During the operation of the audio source instrument, the user metadata information and MPEG source encoded audio frames of the audio source data at the current position of the audio acquisition device are simulated using an MPEG audio acquisition device. The user identification of the current frame is extracted based on the Gaussian neural network method, and the quality level audio frames of the audio source data are acquired. User identification involves calculating the real-time average user metadata information within the pre-selected pitch center and interrupted segments of the audio frames of the quality level, and using the real-time average user metadata information as the basis for determining the recall information of the simulation model in subsequent steps. The simulation operation combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic and combined with the adaptive adjustment strategy of the real-time recall of the simulation model, the simulation model is matched with the corresponding position of the audio source data. The audio source instrument is simulated. Based on the recall of the simulation model, the period manager of the audio source instrument feeds back the audio source instrument through a timing signal, and stores the simulated audio source instrument in the audio source instrument whitelist library at the audio receiver through a trusted database. The audio source instrument simulation, based on the recall of the simulation model, determines the result. The audio source instrument's cycle manager feeds back the audio source instrument via a timing signal, and stores the simulated audio source instrument in a trusted database in the audio source instrument whitelist at the audio receiver. Specifically, this includes: The audio source instrument whitelist library includes at least a first whitelist library, a second whitelist library, and a third whitelist library; The grades include at least a low quality grade, a qualified quality grade, and a high quality grade; The audio source instruments are all turned off when the simulation model is not performing feedback operations; Based on the recall determination result of the simulation model, when the simulation model simulates a low-quality audio source instrument and the recall determination result of the simulation model is adjusted to a decrease in user pitch, the first whitelist library is closed and the second whitelist library is opened. When the simulation model simulates an audio source instrument of acceptable quality level and the recall determination result of the simulation model is adjusted to a decrease in user pitch, the second whitelist library is closed and the third whitelist library is opened. When the simulation model simulates a low-quality audio source instrument and the recall determination result of the simulation model is adjusted to the user's pitch increase, the first whitelist database is closed. When the simulation model simulates an audio source instrument of acceptable quality level and the recall determination result of the simulation model is adjusted to the user's pitch increase, the second whitelist library is closed and the first whitelist library is opened. When the simulation model simulates a high-quality audio source instrument and the recall determination result of the simulation model is adjusted to indicate a decrease in user pitch, the third whitelist library is closed. When the simulation model simulates a high-quality audio source instrument and the recall determination result of the simulation model is adjusted to the user's pitch increase, the third whitelist library is closed and the second whitelist library is opened.

2. The intelligent simulation method based on multi-source audio features using neural algorithms as described in claim 1, characterized in that, The audio source instrument data acquisition, during the operation of the audio source instrument, uses an MPEG audio acquisition device to simulate user metadata information and MPEG source-encoded audio frames of the audio source data at the current position of the audio acquisition device. Based on a Gaussian neural network method, the user identification of the current frame is extracted, and the quality level audio frames of the audio source data are acquired, specifically including: The sample sound detection method based on cepstral coefficients and pyramidal algorithm collects the audio frame coordinates of sample sounds from audio frames in audio source data; Among them, by installing the pitch audio acquisition device on the simulation model, the quality level audio frames of different audio source instruments and backgrounds in the audio source data can be acquired in real time. When the pitch audio acquisition device is vertically receiving the audio source instrument within the preset range, the audio source instrument is arranged at multiple angles. The abnormal pitch values ​​returned by the blank areas between the audio source instruments and the excessively small pitch values ​​returned by the sample sounds with large differences in volume peak and valley values ​​are filtered out to obtain the average user metadata information of the sample sounds on the audio source data. The sample sound detection algorithm using cepstral coefficients and a pyramidal algorithm achieves noise reduction and separation between the simulated sound and the sample sound in the background, and optimizes the average user metadata information of the obtained sample sound.

3. The intelligent simulation method based on multi-source audio features using neural algorithms as described in claim 2, characterized in that, The user identification process involves calculating real-time average user metadata information within a pre-selected pitch center and interrupted segments of the audio frames at the specified quality levels. This real-time average user metadata information is then used as the basis for determining the recall information of the simulation model in subsequent steps. Specifically, this includes: The Gaussian neural network method based on user metadata information extracts sample sound regions according to the pitch weights of each region cluster; Among them, the input quality level audio frames are divided into K region clusters based on the spectrum clustering algorithm; Calculate the initial significance value of the region cluster k in the pitch audio frame; The mono prior is adjusted to the new user metadata information weight, the initial saliency value and the stereo mapping are corrected, and the corrected saliency value is obtained. After obtaining the corrected saliency value, the average pitch of the output saliency user region is combined with the sample sound that has been denoised and separated in the MPEG space and the corresponding position coordinates to obtain the position-pitch integrated information of the sample sound. In the subsequent feedback recall calculation step, the position-pitch integrated information will be used as the initial input information to perform the simulation model to adaptively simulate the features of the audio source data.

4. The intelligent simulation method based on multi-source audio features using neural algorithms as described in claim 3, characterized in that, The simulation operation combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic and an adaptive adjustment strategy for the real-time recall of the simulation model, the corresponding positions of the simulation model and the audio source data are matched. Specifically, this includes: Fixed pitch audio acquisition device position: The relative position between the fixed pitch audio acquisition device and the user's audio output, and the relative position data between the pitch audio acquisition device and the simulation model are also acquired. The performance value of the audio source instrument is obtained by acquiring video frames of audio source data at the simulation model using a pitch audio acquisition device. The video frame data is three-dimensional spatial data including user metadata information. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is calculated based on the video frame data. This difference is the performance value of the audio source instrument. Adjusting the recall state of the simulation model: Comparing the performance value of the audio source instrument with the expected performance value of the audio source instrument, the positional recall state of the simulation model is adjusted so that the performance value of the audio source instrument meets the requirements of the expected performance value of the audio source instrument. The adjustment of the positional recall state of the simulation model includes adjusting the simulation model to increase or decrease.

5. An intelligent simulation system based on multi-source audio features using neural algorithms, characterized in that, include: The audio source instrument data acquisition module, during the operation of the audio source instrument, uses an MPEG audio acquisition device to simulate user metadata information and MPEG source-encoded audio frames of the audio source data at the current position of the audio acquisition device, extracts the user identification of the current frame based on the Gaussian neural network method, and acquires the quality level audio frames of the audio source data. The user identification module calculates the real-time average user metadata information within the pre-selected pitch center and interrupted segments of the audio frames of the quality level, and uses the real-time average user metadata information as the basis for determining the recall information of the simulation model in subsequent steps. The simulation operation module combines the calculated real-time average user metadata information with the inherent parameters of the audio source instrument to determine the real-time recall of the simulation model. Based on the simulation logic and combined with the adaptive adjustment strategy of the real-time recall of the simulation model, the module matches the corresponding position of the simulation model with the audio source data. The audio source instrument simulation module determines the result based on the recall of the simulation model. The period manager of the audio source instrument feeds back the audio source instrument through a timing signal and stores the simulated audio source instrument in the audio source instrument whitelist library at the audio receiver end through a trusted database. The audio source instrument simulation module specifically includes: The audio source instrument whitelist library includes at least a first whitelist library, a second whitelist library, and a third whitelist library; The grades include at least a low quality grade, a qualified quality grade, and a high quality grade; The audio source instruments are all turned off when the simulation model is not performing feedback operations; Based on the recall determination result of the simulation model, when the simulation model simulates a low-quality audio source instrument and the recall determination result of the simulation model is adjusted to a decrease in user pitch, the first whitelist library is closed and the second whitelist library is opened. When the simulation model simulates an audio source instrument of acceptable quality level and the recall determination result of the simulation model is adjusted to a decrease in user pitch, the second whitelist library is closed and the third whitelist library is opened. When the simulation model simulates a low-quality audio source instrument and the recall determination result of the simulation model is adjusted to the user's pitch increase, the first whitelist database is closed. When the simulation model simulates an audio source instrument of acceptable quality level and the recall determination result of the simulation model is adjusted to the user's pitch increase, the second whitelist library is closed and the first whitelist library is opened. When the simulation model simulates a high-quality audio source instrument and the recall determination result of the simulation model is adjusted to indicate a decrease in user pitch, the third whitelist library is closed. When the simulation model simulates a high-quality audio source instrument and the recall determination result of the simulation model is adjusted to the user's pitch increase, the third whitelist library is closed and the second whitelist library is opened.

6. The intelligent simulation system based on multi-source audio features using neural algorithms as described in claim 5, characterized in that, The audio source instrument data acquisition module specifically includes: The sample sound detection method based on cepstral coefficients and pyramidal algorithm collects the audio frame coordinates of sample sounds from audio frames in audio source data; Among them, by installing the pitch audio acquisition device on the simulation model, the quality level audio frames of different audio source instruments and backgrounds in the audio source data can be acquired in real time. When the pitch audio acquisition device is vertically receiving the audio source instrument within the preset range, the audio source instrument is arranged at multiple angles. The abnormal pitch values ​​returned by the blank areas between the audio source instruments and the excessively small pitch values ​​returned by the sample sounds with large differences in volume peak and valley values ​​are filtered out to obtain the average user metadata information of the sample sounds on the audio source data. The sample sound detection algorithm using cepstral coefficients and a pyramidal algorithm achieves noise reduction and separation between the simulated sound and the sample sound in the background, and optimizes the average user metadata information of the obtained sample sound.

7. The intelligent simulation system based on multi-source audio features using neural algorithms as described in claim 6, characterized in that, The user identification module specifically includes: The Gaussian neural network method based on user metadata information extracts sample sound regions according to the pitch weights of each region cluster; Among them, the input quality level audio frames are divided into K region clusters based on the spectrum clustering algorithm; Calculate the initial significance value of the region cluster k in the pitch audio frame; The mono prior is adjusted to the new user metadata information weight, the initial saliency value and the stereo mapping are corrected, and the corrected saliency value is obtained. After obtaining the corrected saliency value, the average pitch of the output saliency user region is combined with the sample sound that has been denoised and separated in the MPEG space and the corresponding position coordinates to obtain the position-pitch integrated information of the sample sound. In the subsequent feedback recall calculation step, the position-pitch integrated information will be used as the initial input information to perform the simulation model to adaptively simulate the features of the audio source data.

8. The intelligent simulation system based on multi-source audio features using neural algorithms as described in claim 7, characterized in that, The simulation operation module specifically includes: Fixed pitch audio acquisition device position: The relative position between the fixed pitch audio acquisition device and the user's audio output, and the relative position data between the pitch audio acquisition device and the simulation model are also acquired. Acquiring the performance value of the audio source instrument: Video frames of audio source data at the simulation model are acquired using a pitch audio acquisition device. The data in the video frames is three-dimensional spatial data including user metadata information. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is calculated based on the data in the video frames. The difference between the simulation result score of the simulation model and the evaluation score of the audio source data is the performance value of the audio source instrument. Adjusting the recall status of the simulation model: The performance value of the audio source instrument is compared with the expected performance value of the audio source instrument. The position recall status of the simulation model is adjusted so that the performance value of the audio source instrument meets the requirements of the expected performance value of the audio source instrument. The adjustment of the position recall status of the simulation model includes adjustments to increase and decrease the simulation model.