Multi-channel audio file generation method and device, computer device and storage medium

By performing vocal separation and speaker channel gain calculation on stereo audio files, audio files adapted to multi-channel audio systems are generated, solving the problem of insufficient audio sources for multi-channel audio systems, improving user experience and reducing production costs.

CN119207338BActive Publication Date: 2026-03-31GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, multi-channel audio systems have a limited number of sound sources, resulting in a poor user experience. Furthermore, traditional methods are costly to manufacture and difficult to adapt to audio systems with different speaker configurations.

Method used

By separating human voices from stereo audio files, determining speaker placement, and calculating channel gain, a multi-channel audio file adapted to the target audio system is generated. Signal separation and gain adjustment are then performed using a pre-trained stereo separation model and signal filters.

Benefits of technology

It enables the flexible conversion of stereo audio files into audio files that are compatible with different multi-channel audio systems, improving the user experience, reducing production costs, and adapting to audio systems with different speaker configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207338B_ABST
    Figure CN119207338B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, computer device, and storage medium for generating multi-channel audio files. The method includes: separating the human voice from a stereo audio file to obtain a human voice signal and an accompaniment signal; determining the arrangement positions of multiple speakers in a target audio system and determining the channel gain of each speaker based on their arrangement positions; generating audio signals corresponding to the speakers based on the human voice signal, accompaniment signal, and speaker channel gains, and merging them into a multi-channel audio file adapted to the target audio system. This invention achieves the goal of converting stereo audio files into different multi-channel formats to adapt to different multi-channel audio systems. Users can determine the channel gain of each speaker according to the number and arrangement positions of the speakers in the target audio system, thereby converting the stereo audio file into a multi-channel audio file adapted to the target audio system, which can improve the user experience of multi-channel audio systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a method, apparatus, computer device, and storage medium for generating multi-channel audio files. Background Technology

[0002] As society develops, people's auditory requirements for audio also increase. Traditional stereo sound systems with two speakers on the left and right sides can no longer meet the needs of some users. As a result, multi-channel sound systems with a more immersive and realistic feel have emerged, such as sound systems with 5.1.2 speakers, 7.1 speakers, and 7.1.4 speakers.

[0003] Multi-channel audio systems with different speaker configurations require different audio formats. Users often need to find suitable audio signals to experience the effects of multi-channel audio systems. However, most music and other audio files on the market are stereo audio files, and the number of audio sources that are compatible with multi-channel audio systems is relatively small. For example, many music songs do not have multi-speaker formats. Users have fewer audio signals to choose from when using multi-channel audio systems, resulting in a poor user experience. Summary of the Invention

[0004] This invention provides a method, apparatus, computer device, and storage medium for generating multi-channel audio files. By converting stereo audio files into different multi-channel audio formats, it adapts to different multi-channel audio systems, thereby solving the problem of a limited number of compatible audio sources for multi-channel audio systems, which leads to a poor user experience.

[0005] To address the above problems, a method for generating multi-channel audio files is provided, including:

[0006] The input stereo audio file is processed to separate the human voice, resulting in the human voice signal and the accompaniment signal;

[0007] Determine the placement of multiple speakers in the target audio system, and determine the channel gain of each speaker based on the placement of the multiple speakers;

[0008] The audio signal corresponding to the speaker is generated based on the human voice signal, the accompaniment sound signal and the channel gain of the speaker, and then merged into a multi-channel audio file adapted to the target audio system.

[0009] A multi-channel audio file generation device is provided, comprising:

[0010] The separation module is used to separate the human voice from the input stereo audio file, obtaining the human voice signal and the accompaniment signal;

[0011] The determination module is used to determine the arrangement positions of multiple speakers in the target audio system, and to determine the channel gain of each speaker based on the arrangement positions of the multiple speakers.

[0012] The generation module is used to generate the corresponding audio signal for the speaker based on the human voice signal, the accompaniment sound signal and the channel gain of the speaker, and merge them into a multi-channel audio file adapted to the target audio system.

[0013] A computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the multi-channel audio file generation method described above.

[0014] A computer-readable storage medium is provided, which stores a computer program, characterized in that the computer program, when executed by a processor, implements the steps of the above-described multi-channel audio file generation method.

[0015] In one solution provided by the aforementioned multi-channel audio file generation method, apparatus, computer equipment, and storage medium, the input stereo audio file is subjected to voice separation to obtain voice signals and accompaniment signals. Then, the arrangement positions of multiple speakers in the target audio system are determined, and the channel gain of each speaker is determined based on the arrangement positions. The corresponding speaker audio signals are then generated based on the voice signals, accompaniment signals, and speaker channel gains, and merged into a multi-channel audio file adapted to the target audio system. This achieves the goal of converting stereo audio files into different channel formats to adapt to different multi-channel audio systems, solving the problem of limited audio sources suitable for multi-channel audio systems. Users can determine the gain of each speaker according to the number and arrangement positions of speakers in the target audio system, thereby flexibly converting stereo audio files into multi-channel audio files adapted to the target audio system, thus improving the user experience of multi-channel audio systems. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the structure of a multi-channel audio file generation system according to an embodiment of the present invention;

[0018] Figure 2 This is a flowchart illustrating a method for generating multi-channel audio files according to an embodiment of the present invention;

[0019] Figure 3 This is an audio spectrum of a stereo audio file in one embodiment of the present invention;

[0020] Figure 4 Yes Figure 3 Audio spectrogram of a 7.1 channel audio file obtained by converting a stereo audio file;

[0021] Figure 5 yes Figure 2 A schematic diagram of the implementation process of step S20;

[0022] Figure 6 , Figure 7 This is a schematic diagram showing the placement of each speaker in a 7.1.4 channel audio system;

[0023] Figure 8 This is a schematic diagram of a multi-channel audio file generation device according to an embodiment of the present invention;

[0024] Figure 9 This is another structural schematic diagram of a multi-channel audio file generation device in one embodiment of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] The multi-channel audio file generation method provided in this embodiment of the invention can be applied to, for example... Figure 1 The multi-channel audio file generation system includes a multi-channel audio file generation device, a target audio system, and a terminal device. The terminal device and the target audio system communicate with the multi-channel audio file generation device via a network. The target audio system is a multi-channel audio system, comprising multiple speakers.

[0027] When a stereo audio file needs to be converted to a multi-speaker format, the user inputs the stereo audio file to be converted into the multi-channel audio file generation device (which can be done via a terminal device). The multi-channel audio file generation device then separates the human voice from the input stereo audio file, obtaining human voice signals and accompaniment signals. It determines the placement of multiple speakers in the target audio system and determines the channel gain of each speaker based on their placement. Then, based on the human voice signal, accompaniment signal, and speaker channel gain, it generates the corresponding speaker audio signals and merges them into a multi-channel audio file adapted to the target audio system. Finally, the multi-channel audio file is directly sent to the target audio system for playback or exported to a terminal device for later transmission to the target audio system for playback. The method proposed in this embodiment achieves the goal of converting stereo audio files into different multi-channel formats to adapt to different multi-channel audio systems. It solves the problems of a limited number of audio sources that can be adapted to multi-channel audio systems and the difficulty for users to find suitable audio sources. Users can determine the gain of each speaker (i.e., determine the weight of each speaker) according to the number of speakers in the target audio system and their corresponding arrangement positions, thereby flexibly converting stereo audio files into multi-channel audio files that can be adapted to the target audio system, which can improve the user's experience of using multi-channel audio systems.

[0028] In this embodiment, the terminal device can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The multi-channel audio file generation device can be various personal computer devices, or it can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0029] In one embodiment, such as Figure 2 As shown, a method for generating multi-channel audio files is provided, which can be applied to... Figure 1 Taking a multi-channel audio file generation device as an example, the following steps are included:

[0030] S10: Separate the human voice from the stereo audio file to obtain the human voice signal and the accompaniment signal.

[0031] When a stereo audio file needs to be converted to a multi-speaker format, the user inputs the stereo audio file to be converted into the multi-channel audio file generating device. Then, the multi-channel audio file generating device separates the human voice from the input stereo audio file to obtain the human voice signal and the accompaniment signal.

[0032] In this embodiment, after acquiring the stereo audio file input by the user, the multi-channel audio file generation device can obtain a pre-trained stereo separation model, then input the stereo audio file into the stereo separation model, and use the stereo separation model to separate the human voice from the audio signal, obtaining the human voice signal and the accompaniment signal. In this embodiment, using a pre-trained stereo separation model for human voice separation is simple and convenient, and the separated data is more accurate.

[0033] S20: Determine the placement of multiple speakers in the target audio system, and determine the channel gain of each speaker based on the placement of the multiple speakers.

[0034] It's important to understand that different types of multi-channel audio systems have different speaker configurations, meaning the number and placement of speakers vary between systems (e.g., an ABC channel audio system). An ABC channel audio system consists of a surround speakers (front, rear, left, and right) + b subwoofers + c overhead speakers, totaling a+b+c speakers. Here, 'a' represents the a surround speakers; 'b' represents the b specially designed subwoofers, requiring a separate subwoofer to operate; and 'c' represents the c overhead speakers, mounted above the listener's position. If 'c' is not specified, it indicates no overhead speakers are used. For example, a 7.1.4 speaker system consists of 7 surround speakers (center, left and right front, left and right surround, left and right rear), plus 1 subwoofer and 4 overhead speakers, for a total of 12 speakers forming a 12-channel audio system; while 5.1, 5.1.2, 7.1, and 7.1.4 channel audio systems have 6, 8, 8, and 12 speakers respectively.

[0035] Therefore, after separating the vocal and accompaniment signals from the stereo audio file, the multi-channel audio file generation device needs to determine the speaker configuration of the target audio system, that is, to determine the number of speakers (number of channels) and the arrangement of the speakers (speaker positions of multiple channels) in the target audio system. After determining the arrangement of the speakers in the target audio system, the channel gain of each speaker needs to be determined based on the arrangement of the speakers.

[0036] S30: Generates the corresponding speaker audio signal based on the human voice signal, accompaniment sound signal and speaker channel gain, and merges them into a multi-channel audio file adapted to the target audio system.

[0037] After determining the channel gain of each speaker, the multi-channel audio file generation device also needs to generate the corresponding speaker audio signals based on the vocal signal, accompaniment signal, and speaker channel gain, and then merge them into a multi-channel audio file adapted to the target audio system. For example, the vocal signal, accompaniment signal, and speaker channel gain can be directly assigned to the corresponding speaker channels to generate the corresponding speaker audio signals. Then, the audio signals from each speaker can be merged to obtain a multi-channel audio file adapted to the target audio system.

[0038] When determining the channel gain of each speaker, the placement of each speaker in the target audio system is taken into account. On the one hand, the number and placement of the speakers can be flexibly changed to make the generated multi-channel audio files compatible with different multi-channel audio systems. On the other hand, the audio playback effect of the corresponding speaker can be adjusted according to the channel gain of each speaker, thereby improving the playback effect of the entire target audio system and enhancing the user's listening experience.

[0039] The method proposed in this embodiment solves the problems of a limited number of compatible audio sources for multi-channel audio systems and the difficulty for users to find suitable audio sources. Users can determine the gain of each speaker (i.e., determine the weight of each speaker) according to the number of speakers in the target audio system and their corresponding arrangement positions, thereby flexibly converting stereo audio files into multi-channel audio files that are compatible with the target audio system, which can improve the user's experience of using the multi-channel audio system.

[0040] Furthermore, most music tracks on the market are in stereo audio, while there are many types of multi-channel audio systems, such as 5.1, 5.1.2, 7.1, and 7.1.4 channels. If a music track is to be compatible with all types of multi-channel audio systems, the stereo audio needs to be converted into different multi-speaker formats for different users to download. This is costly and increases network data and storage requirements. In addition, to maximize the effect, multi-speaker audio sources on the market require the placement and number of speakers in the multi-channel audio system to be fixed. In reality, users may not be able to meet these requirements well due to various factors (such as limited space in the car), resulting in poor music playback quality.

[0041] The method provided in this embodiment effectively solves the above-mentioned problems by decomposing stereo music tracks into vocal and accompaniment signals, determining the gain of each speaker according to a certain algorithm, and then redistributing the vocal and accompaniment signals to the corresponding speakers according to the gain (weight) of each speaker. First, the algorithm can be flexibly changed according to the number and arrangement of speakers in different multi-channel audio systems, thereby adapting to different multi-channel audio systems. Users can convert multi-channel audio files that match the number of speakers in their own audio system without having to convert all types of audio files, reducing production costs and storage volume. Second, users can create multi-channel audio files suitable for their own audio system based on the number and arrangement of speakers in their own audio system, maximizing the auditory enjoyment.

[0042] Taking a stereo audio file as an example, the audio spectrum of the stereo audio file is as follows: Figure 3 As shown, the method provided in this embodiment generates a multi-channel audio file adapted to a 7.1-channel audio system, i.e., a 7.1-channel format audio file. The audio spectrum of this 7.1-channel format audio file is as follows: Figure 4 As shown. From Figure 3 and Figure 4 The comparison shows that, compared with traditional two-channel stereo audio files, the multi-channel audio files of this application have a more immersive and surround sound experience, which can greatly enhance the user's auditory enjoyment.

[0043] In one embodiment, step S10 involves separating the human voice from the input stereo audio file to obtain a human voice signal and an accompaniment signal, specifically including the following steps:

[0044] S11: Obtain the pre-trained stereo separation model, which includes a signal transformation layer and a signal filter.

[0045] After obtaining the stereo audio file input by the user, the multi-channel audio file generation device needs to acquire a pre-trained stereo separation model. This model is then used to separate the human voice from the stereo audio file, obtaining the human voice signal and the accompaniment signal. The stereo separation model is a neural network model, which includes a signal transformation layer and a signal filter.

[0046] S12: Input the stereo audio file into the stereo separation model, perform Fourier transform through the signal transformation layer to obtain the transformed audio signal, and perform human voice separation on the transformed audio signal through the signal filter to obtain the human voice signal and the accompaniment signal.

[0047] After obtaining the pre-trained stereo separation model, the stereo separation model is called and the stereo audio file is input into the stereo separation model to achieve the following steps: Fourier transform is performed through the signal transformation layer to obtain the transformed audio signal, and the human voice is separated from the transformed audio signal through the signal filter to obtain the human voice signal and the accompaniment signal.

[0048] In this embodiment, a pre-trained stereo separation model is obtained, and then the stereo audio file is input into the stereo separation model. A Fourier transform is performed through a signal transformation layer to obtain a transformed audio signal, and a signal filter is used to separate the vocals from the transformed audio signal, resulting in vocal and accompaniment signals. This clarifies the specific steps for separating vocals from the input stereo audio file to obtain vocal and accompaniment signals. Using a neural network model including a signal transformation layer and a signal filter for vocal separation of stereo audio files is simple, convenient, and highly accurate. Furthermore, during the training of the stereo separation model, the network parameters of the signal transformation layer and the signal filter are iterated simultaneously, improving the coordination between the signal transformation layer and the signal filter, thereby enhancing the accuracy of vocal separation.

[0049] In other embodiments, after obtaining the stereo audio file input by the user, the multi-channel audio file generation device first performs a Fourier transform on the stereo audio file based on the Fourier algorithm to obtain a transformed audio signal. Then, the transformed audio signal is input to a signal filter to separate the human voice from the transformed audio signal, resulting in a human voice signal and an accompaniment signal. The human voice separation can be performed using a conventional signal filter, which can meet a certain separation effect without the need for model training.

[0050] The signal filter can be obtained by iteratively training the network parameters of a filter based on stereo audio (including vocals and accompaniment) and standard stereo vocals. The training process is as follows: First, the stereo audio is subjected to a Fourier transform to obtain the transformed stereo audio. Then, the transformed stereo audio is input into the filter for vocal separation, resulting in vocals and accompaniment. The residuals of the vocals and accompaniment are compared, and the filter parameters are updated based on the residuals. This vocal separation process is repeated until the residuals of the vocals and accompaniment meet the requirements. When the filter has achieved satisfactory vocals and accompaniment separation, and the filter converges, the current filter parameters are output as the signal filter for subsequent use. Training the signal filter in this way can further improve the accuracy of the signal filter, thereby improving the accuracy of the signal data obtained from the subsequent vocal separation by the mechanism.

[0051] In one embodiment, before step S10, a stereo separation model needs to be pre-trained so that after acquiring the input stereo audio file, the stereo separation model can be used to separate the human voice from the audio signal to obtain the accompaniment signal. Specifically, the stereo separation model is trained in the following way:

[0052] S01: Acquire stereo audio samples including vocals and accompaniment, as well as standard vocals from the stereo audio samples.

[0053] When training the model, the multi-channel audio file generation device first needs to acquire stereo audio samples including human voice and accompaniment, as well as standard human voice samples. These stereo audio samples can be one or multiple samples. Using a single stereo audio sample for model training requires less iteration data, reducing the number of iterations and quickly obtaining a stereo separation model that meets the separation requirements. Using multiple stereo audio samples for model training, while requiring more iteration data and increasing the number of iterations, can improve the training accuracy of the stereo separation model. The stereo audio sample can be determined based on actual needs.

[0054] S02: Input the stereo audio samples into the pre-trained neural network model, perform Fourier transform through the Fourier transform network to obtain the Fourier signal, and then separate the human voice and the accompaniment through the Fourier signal using a filter.

[0055] In this embodiment, the pre-trained neural network model includes a Fourier transform network and a filter. The multi-channel audio file generation device inputs stereo audio samples into the pre-trained neural network model, performs Fourier transform through the Fourier transform network to obtain Fourier signals, and then uses the filter to separate the human voice from the Fourier signals, resulting in separated human voice and separated accompaniment.

[0056] S03: Determine the residuals between the separated human voice and the standard human voice;

[0057] S04: Iteratively update the filter parameters based on the residuals until the separated human voice obtained by the filter meets the effect requirements, and output the current Fourier transform network and filter as a stereo separation model.

[0058] After obtaining the separated vocals and accompaniment, the multi-channel audio file generation device needs to determine the residual between the separated vocals and the standard vocals. Then, based on the residual, the filter parameters are iteratively updated until the separated vocals meet the effect requirements. At this point, the current Fourier transform network and filter are output as a stereo separation model. Specifically, after determining the residual between the separated vocals and the standard vocals, it is necessary to determine if the residual is less than a preset residual value. If the residual is greater than the preset residual value, the separation effect of the filter is determined to be unsatisfactory. The filter network parameters are then adjusted, and steps S02-S03 are repeated until the residual is less than or equal to the preset residual value. When the separated vocals and accompaniment meet the effect requirements, the filter is considered to have converged and can separate the vocals and accompaniment to a satisfactory effect. The current Fourier transform network and filter are then output as a stereo separation model. Here, the output Fourier transform network is the signal transformation layer, and the output filter is the signal filter.

[0059] In this embodiment, stereo audio samples including human voice and accompaniment are acquired, along with a standard human voice from the stereo audio samples. The stereo audio samples are input into a pre-trained neural network model, and a Fourier transform is performed using a Fourier transform network to obtain Fourier signals. Human voice separation is then performed on the Fourier signals using a filter, resulting in separated human voice and separated accompaniment. The residuals between the separated human voice and the standard human voice are determined. The filter parameters are iteratively updated based on the residuals until the separated human voice obtained by the filter meets the effect requirements. The current Fourier transform network and filter are then output as the stereo separation model. This clarifies the training process of the stereo separation model. Through the above training method, a stereo separation model with good separation effect can be trained, providing a solid foundation for subsequent human voice separation of stereo audio files.

[0060] Furthermore, in this embodiment, the stereo separation model training process only iteratively updates the filter parameters, which reduces the number of network parameter iterations and thus accelerates model convergence. In other embodiments, the parameters of the Fourier transform network can be updated simultaneously during the stereo separation model training process, thereby further improving the accuracy of the stereo separation model.

[0061] In one embodiment, such as Figure 5 As shown, step S20, which involves determining the channel gain of each speaker based on the arrangement of the multiple speakers, specifically includes the following steps:

[0062] S21: Set an actual listening position in the target audio system, and determine the listening distance of each speaker based on the actual listening position and the speaker placement.

[0063] S22: Determine the channel gain of each speaker based on the listening distance of the speaker.

[0064] After determining the placement of the multiple speakers in the target audio system, an actual listening position (i.e., the listener's position) needs to be set in the target audio system. In this embodiment, the actual listening position in the target audio system may not coincide with the optimal listening position of the target audio system, or it may coincide with the optimal listening position of the target audio system.

[0065] Understandably, the optimal listening position for a target audio system is located at the center of a symmetrical shape formed by the surround speakers. For multi-channel audio systems, the placement of the surround speakers is generally based on the center speaker's position. Specifically, first establish an xyz coordinate system, with the xy plane parallel to the horizontal plane. The origin of the coordinate system points towards the center speaker's position as the positive y-axis, and the z-axis is perpendicular to the xy plane. Other speakers are arranged around the origin at different heights. In this multi-channel audio system, the surround speakers (excluding the center speaker) are symmetrical about each other relative to the yz plane, and the sky speakers are also symmetrical about each other relative to the yz plane, forming multiple symmetrical speaker groups. The origin of the coordinate system is then the center of the multi-channel audio system, i.e., the optimal listening position. Depending on the number of speakers in different multi-channel audio systems, these multiple symmetrical speaker groups include two or three symmetrical speaker groups corresponding to the surround speakers: a main symmetrical speaker group, a surround symmetrical speaker group, and a rear symmetrical speaker group (some multi-channel audio systems do not have this rear symmetrical speaker group), as well as one or two symmetrical speaker groups corresponding to the sky speakers. Taking a 7.1.4 channel audio system as an example, the speaker placement is as follows: Figure 6 , Figure 7 As shown, the 7.1.4-channel audio system includes seven surround speakers: right speaker R, center speaker C, left speaker L, right surround speaker Rs, left surround speaker Ls, left rear speaker Lb, and right rear speaker Rb; one woofer LEF; and four sky speakers: upper left front surround speaker Ltf, upper right front surround speaker Rtf, upper left rear surround speaker Ltb, and upper right rear surround speaker Rtb. O is the origin of the coordinate system (i.e., the optimal listening position), and the direction from the origin O to the placement of the center speaker C is the positive y-axis. The five speaker groups R and L, Rs and Ls, Lb and Rb, Ltf and Rtf, and Ltb and Rtb are symmetrical speaker groups, each pair being symmetrical with respect to the yz plane. R and L, Rs and Ls, and Lb and Rb are the three symmetrical speaker groups corresponding to the surround speakers; Ltf and Rtf, and Ltb and Rtb are the two symmetrical speaker groups corresponding to the sky speakers.

[0066] After setting an actual listening position in the target audio system, the listening distance of each speaker needs to be determined based on the actual listening position and the speaker placement. The listening distance of a speaker is the straight-line distance between the speaker and the actual listening position. Then, the channel gain of each speaker is determined based on the straight-line distance between the speaker and the actual listening position. For example, the farther the straight-line distance between the speaker and the actual listening position, the greater the channel gain; the closer the straight-line distance, the smaller the channel gain, ensuring that the same audio signal can be heard from different listening positions.

[0067] When using a multi-channel audio system, the listening experience varies depending on the listening position, with the optimal listening position providing the best auditory experience. Therefore, traditional multi-channel audio file generation methods typically use the optimal listening position as a fixed standard for audio conversion, assuming the actual listening position is the optimal location. However, in real life, due to space limitations and other factors, listeners may not be in the optimal listening position (e.g., in complex, small spaces like vehicles or homes), and the listener's position can change in different scenarios. Using the optimal listening position as a fixed standard for audio conversion can lead to suboptimal sound quality after playing the audio file. In this embodiment, the straight-line distance between the actual listening position and the speaker placement is determined, and the channel gain of each speaker is then determined. When calculating the speaker channel gain, the relative position of the actual listening position and each speaker is considered, allowing the channel gain of each speaker to change with the actual listening position (listener's position). This ensures that the optimal listening experience can be obtained at different positions within the speaker's enclosure, providing the same immersive auditory experience even if the listener is not in the center position.

[0068] In this embodiment, an actual listening position is set in the target audio system, and the straight-line distance between the speakers and the actual listening position is determined based on the actual listening position and the speaker placement. Then, the channel gain of each speaker is determined based on the listening distance of each speaker. This refines the steps of determining the channel gain of each speaker based on the placement of multiple speakers. When calculating the speaker channel gain, the relative positions of the actual listening position and each speaker are taken into account. The channel gain of each speaker can be adjusted according to the actual listening position, thereby achieving optimal listening performance at different positions within the speaker enclosure and improving the user experience.

[0069] In other embodiments, the actual listening position may be disregarded, and the optimal listening position of the target audio system may be used as a standard to determine the straight-line distance between the speaker and the optimal listening position, thereby determining the channel gain of each speaker. For the specific determination process, please refer to the calculation process of determining the channel gain of each speaker based on the actual listening position, which will not be repeated here.

[0070] In one embodiment, the channel gain includes vocal gain and accompaniment gain. Step S22, which determines the channel gain of each speaker based on the listening distance of each speaker, specifically includes the following steps:

[0071] S221: Determine the sound field width factor for each loudspeaker, which includes the vocal width factor and the accompaniment width factor.

[0072] S222: Determine the vocal gain of each speaker based on the vocal width factor and listening distance of each speaker.

[0073] S223: Determine the accompaniment gain of each speaker based on the accompaniment width coefficient and listening distance of each speaker.

[0074] After determining the listening distance of each speaker based on the actual listening position and speaker placement, it is also necessary to determine the sound field width factor of each speaker. In this embodiment, the channel gain includes vocal gain and accompaniment gain. Correspondingly, the sound field width factor includes vocal width factor and accompaniment width factor.

[0075] After determining the sound field width factor of each speaker, the vocal gain of each speaker is determined based on its vocal width factor and listening distance (the straight-line distance between each speaker and the actual listening position). Simultaneously, the accompaniment gain of each speaker is determined based on its accompaniment width factor and listening distance. This sound field width factor is a preset constant vector. The sound field width factor controls the energy contribution of each channel.

[0076] It's important to understand that different sound field widths create different subjective auditory experiences, and these effects can be achieved by controlling the sound field width. If the sound field width is too wide, the sound appears as a single, diffuse area without clear localization, making the sound sound blurry and unreal. Conversely, a narrower sound field width provides more precise localization, resulting in clearer vocals. However, the sound field width shouldn't be too narrow either. Because human voices are singular in audio, their sound field width is relatively small. For example, in a live concert, the singer's location is clear, allowing you to perceive the specific sound source but not pinpoint it precisely; the sound field width for human voices is relatively narrow. In contrast, accompaniment in a concert involves various instruments arranged in a dispersed manner. If the accompaniment's sound field width is narrow, different instruments may overlap in some location, resulting in a poor listening experience. Only with a sufficiently wide sound field can different instruments be effectively separated, leading to a better auditory effect.

[0077] In summary, the vocal width coefficient and the accompaniment width coefficient should be different in the sound field width coefficient. The vocal width coefficient is used to flexibly control the sound field width formed by the vocal signal, while the accompaniment width coefficient is used to flexibly control the sound field width formed by the accompaniment signal. This achieves the goal of avoiding widening the vocal signal in the music while widening the accompaniment signal, thereby improving the listening effect. While enhancing the sense of immersion and surround sound, it also restores the listening effect of a real stage or concert to a greater extent.

[0078] In this embodiment, the sound field width coefficient can be a set of fixed constant coefficients determined in advance based on audio playback effect tests for subsequent calculations. In other embodiments, multiple sound field width coefficients can be preset, each including a corresponding vocal width coefficient and accompaniment width coefficient, so that the subsequently generated multi-channel audio files have different listening effects, that is, different sound field width coefficients can correspond to different listening effects. Since different users have different listening preferences, when generating multi-channel audio files, users can select a sound field width coefficient that meets their preferences from multiple sound field width coefficients according to their actual needs. This sound field width coefficient includes a vocal width coefficient and accompaniment width coefficient that meet their listening preferences, so that the multi-channel audio file generation device generates the channel gain of each speaker according to the sound field width coefficient selected by the user, thereby obtaining a multi-channel audio file that meets the user's listening preferences and satisfies the listening needs of different users.

[0079] In this embodiment, the sound field width coefficient of each speaker is determined, including the vocal width coefficient and the accompaniment width coefficient. Then, based on the vocal width coefficient and listening distance of each speaker, the vocal gain of each speaker is determined, and based on the accompaniment width coefficient and listening distance of each speaker, the accompaniment gain of each speaker is determined. This refines the specific steps for determining the channel gain of each speaker based on the listening distance. When calculating the channel gain of each speaker, the vocal width coefficient is used to flexibly control the sound field width formed by the vocal signal, and the accompaniment width coefficient is used to flexibly control the sound field width formed by the accompaniment signal. This achieves the goal of avoiding widening the vocal signal in the music while widening the accompaniment signal, thereby improving the listening effect. While enhancing the immersive and surround sound, it also reproduces the listening effect of a real stage or concert to a greater extent.

[0080] In one embodiment, step S222, which involves determining the vocal gain of each speaker based on the vocal width coefficient and listening distance, specifically includes the following steps:

[0081] S2221: When the target audio system has sky speakers, the multiple speakers are divided into a first speaker group and a second speaker group.

[0082] After determining whether the target audio system has overhead speakers, if overhead speakers are confirmed, the multiple speakers need to be divided into a first speaker group and a second speaker group according to their placement. The second speaker group consists entirely of overhead speakers. The first speaker group may include all speakers except the overhead speakers; alternatively, the first speaker group may consist only of surround speakers, excluding woofers.

[0083] To ensure the accuracy of the channel gain of each speaker, it is necessary to consider the influence of the height position of the sky speaker on the channel gain, group the positions of different speakers, and use different algorithms to calculate the gain of the surround speakers and sky speakers.

[0084] S2222: For the first loudspeaker group, the voice gain of each loudspeaker is determined based on the first human voice constraint model, the human voice width coefficient of each loudspeaker, and the listening distance.

[0085] After dividing the multiple loudspeakers into a first loudspeaker group and a second loudspeaker group, for the first loudspeaker group, the voice gain of each loudspeaker is determined based on the first human voice constraint model, the human voice width coefficient of each loudspeaker, and the straight-line distance between each loudspeaker and the actual listening position (i.e., the listening distance).

[0086] The first human voice constraint model includes multiple constraint formulas. For example, it can include a first constraint formula and a second constraint formula. The first constraint formula can constrain each speaker based on the principle of vector superposition of sound signals; the second constraint formula can constrain any group of left and right symmetrical speakers based on the principle of sound symmetry on the left and right sides of the listener; the second constraint formula can constrain speakers located on the same side of the actual listening position based on the principle that the sound effect on the same side of the listener is the same.

[0087] S2223: For the second speaker group, determine the voice gain of each speaker based on the second human voice constraint model and the listening distance of each speaker.

[0088] Meanwhile, for the second speaker group, the voice gain of each speaker is determined based on the second voice constraint model and the straight-line distance between each speaker and the actual listening position (i.e., the listening distance). The constraint formulas for the second voice constraint model differ from those of the first voice constraint model; the second voice constraint model's formula considers the influence of the height of the sky speaker on the listening effect.

[0089] S2224: When the target audio system does not have sky speakers, determine the voice gain of each speaker based on the first human voice constraint model, the human voice width coefficient of the speaker and the listening distance.

[0090] After determining the sound field width coefficient of the loudspeaker, it is determined whether there is a sky loudspeaker in the target audio system. When it is determined that there is no sky loudspeaker in the target audio system, there is no need to consider the difference of sky loudspeakers. At this time, the voice gain of each loudspeaker can be determined directly based on the first human voice constraint model, the human voice width coefficient of each loudspeaker, and the straight-line distance between each loudspeaker and the actual listening position (i.e., the listening distance).

[0091] In other embodiments, it is not necessary to determine whether the target audio system has sky speakers. Instead, the voice gain of each speaker can be determined directly based on the first human voice constraint model, the human voice width coefficient of the speaker, and the straight-line distance between each speaker and the actual listening position, thus reducing the number of judgment steps.

[0092] In this embodiment, when the target audio system does not have overhead speakers, the voice gain of each speaker is determined based on the first voice constraint model, the speaker's voice width coefficient, and the listening distance. When the target audio system has overhead speakers, the multiple speakers are divided into a first speaker group and a second speaker group. The voice gain of each speaker in the first speaker group is determined based on the first voice constraint model, the speaker's voice width coefficient, and the listening distance. The voice gain of each speaker in the second speaker group is determined based on the second voice constraint model and the listening distance. The specific calculation process for determining the voice gain of each speaker based on the speaker's voice width coefficient and listening distance is clarified. Different channel gain calculation methods are adopted according to the speaker configuration of different target audio systems, improving the accuracy of channel gain technology. When the target audio system has overhead speakers, the influence of the overhead speaker's height on the channel gain is considered, thus enabling the calculation of more accurate channel gains.

[0093] In one embodiment, when it is determined that the target audio system does not have a sky speaker, for example, the target audio system is a 5.1 channel audio system or a 7.1 channel audio system, the subwoofer of the target audio system can be determined, and then the preset gain is used as the vocal gain of the subwoofer. The vocal gain of the remaining speakers only needs to be determined based on the vocal width coefficient of the speaker and the straight-line distance between the speaker and the actual listening position.

[0094] In this embodiment, the preset gain can be a pre-set fixed gain value, such as 0. The playback device of the subwoofer is a subwoofer. The purpose of the subwoofer is to reproduce the low-frequency accompaniment such as drums and bass in the music as much as possible, and it will not play vocals. In step S10, the music has already separated the vocals and accompaniment. The effect of the low-frequency accompaniment after gaining is not significant. Therefore, a fixed gain value can be directly set as the vocal gain or accompaniment gain of the subwoofer, reducing the amount of data processing in the channel gain calculation process.

[0095] In one embodiment, step S2222 or step S2224, namely determining the voice gain of each speaker based on the first voice constraint model, the voice width coefficient of each speaker, and the listening distance, specifically includes the following steps:

[0096] S201: Determine the first included angle of each speaker.

[0097] When calculating the channel gain of each speaker, the first step is to determine the first angle between each speaker and its actual listening position. This first angle is the angle between the speaker ray and the positive y-axis in the xyz coordinate system; the speaker ray is the ray formed by the actual listening position pointing towards the speaker's placement; and the positive y-axis is the direction from the origin of the xyz coordinate system towards the center speaker's placement. Taking a 7.1.4-channel audio system as an example, the positive y-axis direction can be found in... Figure 6 .

[0098] S202: Determine the second included angle of each speaker.

[0099] At the same time, it is also necessary to determine the second included angle of each speaker based on the actual listening position and the placement position of each speaker. The second included angle is the horizontal included angle of the speaker ray, that is, the angle formed by the ray formed by the actual listening position pointing to the placement position of the speaker and the horizontal plane.

[0100] S203: Input the human voice width coefficient, first included angle, second included angle and listening distance of each speaker into the first human voice constraint model for solution to obtain the human voice gain of each speaker.

[0101] After determining the first and second included angles of each speaker, the human voice width coefficient, the first included angle, the second included angle, and the listening distance of each speaker are input into the first human voice constraint model for solution to obtain the human voice gain of each speaker.

[0102] In this embodiment, a first angle and a second angle are determined for each speaker. The first angle is the angle between the speaker ray and the positive y-axis, where the speaker ray is the ray formed by the actual listening position pointing to the speaker, and the positive y-axis is the direction from the origin of the coordinate system to the center speaker. The second angle is the horizontal angle of the speaker ray. Then, the vocal width coefficient of each speaker, the first angle, the second angle, and the listening distance are input into a first vocal constraint model for solution, resulting in the vocal gain of each speaker. This clarifies the specific steps for determining the vocal gain of each speaker based on the first vocal constraint model, the vocal width coefficient of each speaker, and the listening distance. By using data such as the listening distance, the first angle, and the second angle, the relative position of the actual listening position and each speaker is defined from multiple dimensions. Thus, the vocal gain of the speaker is calculated based on the relative position of the actual listening position and each speaker, and the vocal width coefficient. This not only allows listeners to obtain the same immersive experience regardless of their position but also avoids widening the vocal signal in the music, improving the listening effect.

[0103] Specifically, the first voice constraint model includes a first constraint formula, a second constraint formula, and a third constraint formula. The voice width coefficient, first angle, second angle, and listening distance of each loudspeaker are input into the first voice constraint model for solution to obtain the voice gain of each loudspeaker. The steps include:

[0104] S2031: Adjust the symmetrical speaker groups in the target audio system according to the actual listening position to obtain multiple corresponding speaker groups.

[0105] At this point, it is necessary to first obtain multiple symmetrical speaker groups in the target audio system. As mentioned earlier, the positive direction of the y-axis is the direction from the origin of the coordinate system to the placement position of the center speaker. Each speaker in the symmetrical speaker group is symmetrical with respect to the yz plane. Figure 6 As shown, multiple symmetrical speaker groups include a main symmetrical speaker group, a surround symmetrical speaker group, a rear symmetrical speaker group, and one or two symmetrical speaker groups corresponding to the sky speakers. Then, based on the actual listening position, it is determined whether adjustments are needed to the symmetrical speaker groups in the target audio system: when the actual listening position is on the y-axis, no adjustment is needed, and the multiple symmetrical speaker groups are directly recorded as multiple corresponding speaker groups; when the actual listening position deviates from the y-axis, the center speaker is added to the main symmetrical speaker group as a new corresponding speaker group, thus combining with other symmetrical speaker groups to obtain multiple corresponding speaker groups. The new corresponding speaker group includes (center speaker + speaker located next to the center speaker and away from the actual listening position) and speaker located next to the center speaker and close to the actual listening position, that is, the center speaker + speaker located next to the center speaker and away from the actual listening position corresponds to the speaker located next to the center speaker and close to the actual listening position. For example... Figure 6 As shown, if the actual listening position deviates from the y-axis, and the actual listening position is located to the left of the y-axis ( Figure 6 If the left side is selected, the center speaker will be added to the main symmetrical speaker group (including the left speaker L and speaker R) as a new corresponding speaker group. This new corresponding speaker group includes the speakers corresponding to (C+R) and L.

[0106] S2032: Input the first included angle and listening distance of each speaker into the first constraint formula, input the second included angle and listening distance of each speaker in the corresponding speaker group into the second constraint formula, input the human voice width coefficient and listening distance of each speaker into the third constraint formula, and solve to obtain the human voice gain of each channel.

[0107] By inputting the first included angle and listening distance of each speaker into the first constraint formula, the second included angle and listening distance of each speaker in the corresponding speaker group into the second constraint formula, and the vocal width coefficient and listening distance of each speaker into the third constraint formula, the vocal gain of each channel can be obtained.

[0108] In order to obtain a better listening effect, based on the principle of vector superposition of sound signals, the first constraint formula can be obtained according to the placement position of the speaker and the actual listening position. The first constraint formula is expressed as follows:

[0109]

[0110]

[0111] Where n represents the nth speaker; θ n P represents the first included angle of loudspeaker n, which is the angle formed by the ray pointing from the actual listening position to the placement position of loudspeaker n, and the positive direction of the y-axis; n L represents the vocal gain of speaker n; n This represents the listening distance of speaker n, which is the straight-line distance between speaker n and the actual listening position. This represents the components of the audio signal from speaker n in the xy plane.

[0112] This represents the energy of the audio signal from speaker n.

[0113] To ensure consistent sound quality on both sides of the listener's position, based on the principle of left-right symmetry of the listener's position (actual listening position), a second constraint formula can be obtained for each corresponding speaker group. The second constraint formula is expressed as follows:

[0114]

[0115] Among them, P left With P rig h t To represent the channel gain of two speakers in a corresponding speaker group, it can be understood as the vocal gain of the speaker on the left side of the actual listening position and the vocal gain of the speaker on the right side of the actual listening position in a corresponding speaker group; θ left With θ rig h t Let L represent the second included angle between two speakers in a corresponding speaker group, that is, the horizontal included angle of the left speaker at the actual listening position and the horizontal included angle of the right speaker at the actual listening position; 1left With L rig h t These represent the listening distances of two speakers in a corresponding speaker group, specifically the straight-line distance between the left speaker and the actual listening position, and the straight-line distance between the right speaker and the actual listening position.

[0116] Since the sound field width factor can affect the sound field width and thus the sound effect, the sound field width formed by the human voice signal can be flexibly controlled through the human voice width factor. According to the listening preferences of different listeners, the human voice gain of each speaker can be adjusted using the human voice width factor, thereby obtaining the third constraint formula for constraining each speaker group. The third constraint formula is expressed as follows:

[0117]

[0118] Where n represents the nth speaker; θ n P represents the first included angle of speaker n; n L represents the vocal gain of speaker n; n The listening distance of speaker n; a n This represents the vocal width coefficient for each speaker n.

[0119] In this embodiment, the first human voice constraint model includes a first constraint formula, a second constraint formula, and a third constraint formula. The symmetrical speaker groups in the target audio system are adjusted according to the actual listening position to obtain multiple corresponding speaker groups. The first included angle and listening distance of each speaker are input into the first constraint formula, the second included angle and listening distance of each speaker in the corresponding speaker group are input into the second constraint formula, and the human voice width coefficient and listening distance of each speaker are input into the third constraint formula. The human voice gain of each channel is then calculated, clarifying the specific process of calculating the human voice gain of each speaker and providing a data foundation for the subsequent allocation of audio signals to each speaker.

[0120] In one embodiment, to ensure similar sound effects on the left and right sides of the listener, based on the principle of consistent sound effects on the same side and considering the different listening preferences of different listeners, a vocal width coefficient is used to constrain the speakers located on the same side of the actual listening position to obtain a better listening experience. That is, after adjusting the symmetrical speaker groups in the target audio system according to the actual listening position to obtain multiple corresponding speaker groups, it is necessary to group the multiple speakers on the left and right sides according to the actual listening position, and record the multiple speakers located on the same side of the actual listening position as the same-side speaker group. Then, the first included angle and listening distance of each speaker are input into the first constraint formula, the second included angle and listening distance of each speaker in the corresponding speaker group are input into the second constraint formula, and the vocal width coefficient and listening distance of each speaker in the same-side speaker group are input into the third constraint formula to solve for the vocal gain of each channel.

[0121] Among them, the human voice width coefficient is used to constrain each speaker in the same-side speaker group, thus obtaining the third constraint formula for constraining the same-side speaker group. The third constraint formula can also be expressed as follows:

[0122]

[0123] Where j represents the j-th speaker in the same-side speaker group, that is, the j-th speaker located on the same side as the actual listening position, and j is less than n; θ j P represents the first included angle of speaker j in the same-side speaker group; j L represents the vocal gain of speaker j in the same-side speaker group; j This indicates the listening distance of speaker j in the same-side speaker group; a j This represents the vocal width coefficient of speaker j in the same-side speaker group.

[0124] The vocal width coefficient is a preset constant vector. By coordinating the vocal width coefficient with the accompaniment width coefficient, the goal is to avoid widening the vocal signal in the music and only widen the accompaniment signal. The vocal width coefficients of speaker groups on different sides of the actual listening position can be the same or different, depending on the user's listening preferences. Similarly, the vocal width coefficients of individual speakers within the same side speaker group can be the same or different, also depending on the user's listening preferences. When the vocal width coefficients of individual speakers within the same side speaker group are different, the vocal width coefficient of each speaker decreases as the distance between that speaker and the center speaker increases, resulting in a better subjective listening experience for the subsequently generated multi-channel audio file.

[0125] In this approach, multiple speakers in the first speaker group can be grouped into left and right sides to obtain two same-side speaker groups. The second speaker group (i.e., the sky speaker group) does not apply the sound field width coefficient (voice width coefficient and accompaniment width coefficient), so it does not need to be grouped into left and right sides. That is, for the first speaker group, the symmetrical speaker groups in the target audio system are adjusted according to the actual listening position to obtain multiple corresponding speaker groups. The multiple speakers are then grouped into left and right sides according to the actual listening position, and the multiple speakers located on the same side of the actual listening position are recorded as same-side speaker groups. Then, the first included angle and listening distance of each speaker are input into the first constraint formula, the second included angle and listening distance of each speaker in the corresponding speaker group are input into the second constraint formula, and the voice width coefficient and listening distance of each speaker in the same-side speaker group are input into the third constraint formula to solve for the voice gain of each channel.

[0126] For example, a 7.1.4 channel audio system, such as Figure 6 and Figure 7As shown, L, Ls, Lb, Ltf, and Ltb are all on the same side, which can be a single-sided speaker group; R, Rs, Rb, Rtf, and Rtb are also on the same side, which can also be a single-sided speaker group. Since Ltf, Ltb, Rtf, and Rtb are sky channels, the two single-sided speaker groups included in the first speaker group are L, Ls, and Lb; and R, Rs, and Rb, respectively. In this case, the vocal width coefficients of L, Ls, and Lb are represented by a1, a2, and a3, respectively. a1, a2, and a3 can be the same or different. According to experimental results, when a1 = 2a2 = 3a3, the subjective listening experience of the generated multi-channel audio file is better.

[0127] In this embodiment, by adjusting the symmetrical speaker groups in the target audio system according to the actual listening position, multiple corresponding speaker groups are obtained. The multiple speakers are then grouped into left and right sides according to the actual listening position. The multiple speakers located on the same side of the actual listening position are recorded as the same-side speaker group. Then, the first included angle and listening distance of each speaker are input into the first constraint formula, the second included angle and listening distance of each speaker in the corresponding speaker group are input into the second constraint formula, and the human voice width coefficient and listening distance of each speaker in the same-side speaker group are input into the third constraint formula. The human voice gain of each channel is then obtained by solving the formula. This clarifies another specific process for calculating the human voice gain of each speaker. By using the human voice width coefficient to constrain the speakers located on the same side of the actual listening position, a better listening experience can be obtained.

[0128] In one embodiment, step S2223, namely for the second speaker group, determines the voice gain of each speaker based on the second voice constraint model and the listening distance of each speaker, specifically including the following steps:

[0129] S22231: Determine the first included angle of each speaker.

[0130] The process of determining the first included angle of the loudspeaker is described above and will not be repeated here.

[0131] S22232: Input the first included angle and listening distance of each speaker in the second speaker group into the second human voice constraint model for solution, and obtain the human voice gain of each speaker in the second speaker group.

[0132] Then, the first included angle and listening distance of each speaker in the second speaker group are input into the second human voice constraint model for solution, to obtain the human voice gain of each speaker in the second speaker group. That is, the first included angle and listening distance of each sky speaker are input into the second human voice constraint model for solution, to obtain the human voice gain of each sky speaker. Considering that the sky speakers affect the panoramic sound effect, sound field width adjustment is unnecessary. Furthermore, considering the influence of the sky speaker height on the sound effect, a fourth constraint formula is obtained for each sky speaker, which is the second human voice constraint model.

[0133] When the number of sky speakers (represented by m) is 4, that is, the second speaker group includes 4 speakers (sky speakers), m = 1, 2, 3, 4, the second human voice constraint model is as follows:

[0134]

[0135]

[0136] Where θ1, θ2, θ3, and θ4 represent the first included angles of the first to fourth sky speakers, respectively; P1, P2, P3, and P4 represent the human voice gain of the first to fourth sky speakers, respectively; and L1, L2, L3, and L4 represent the listening distances of the first to fourth sky speakers, respectively. The first and second sky speakers form one corresponding speaker group; the third and fourth sky speakers form another corresponding speaker group. The first and third sky speakers are located on the same side, forming a same-side speaker group; the second and third sky speakers are located on the same side, forming a same-side speaker group. (Reference) Figure 7 The first to fourth sky speakers can be the upper left front speaker group Ltf, the upper right front speaker group Rtf, the upper left rear speaker group Ltb, and the upper right rear speaker group Rtb, respectively.

[0137] When the number of sky speakers is 2, that is, when the second speaker group includes two speakers (sky speakers), the second human voice constraint model is as follows:

[0138]

[0139]

[0140] Where θ1 and θ2 represent the first included angle of the first sky speaker and the first included angle of the fourth sky speaker, respectively; P1 and P2 represent the voice gain of the first sky speaker and the channel gain of the fourth sky speaker, respectively; L1 and L2 represent the listening distance of the first sky speaker and the listening distance of the fourth sky speaker, respectively.

[0141] In other embodiments, the process of determining the voice gain of each speaker in the second speaker group can refer to the calculation process of the voice gain of each speaker in the first speaker group described above, and will not be repeated here.

[0142] In this embodiment, by determining the first included angle of each speaker, and then directly inputting the first included angle and listening distance of each speaker in the second speaker group into the second human voice constraint model for solution, the human voice gain of each speaker in the second speaker group is obtained. This clarifies the specific process of determining the human voice gain of each speaker in the second speaker group and provides an accurate data basis for subsequent calculations.

[0143] It is important to understand that the calculation process for the accompaniment gain of each speaker is similar to that for the vocal gain. Specifically, in step S223, the accompaniment gain of each speaker is determined based on its accompaniment width coefficient and listening distance, which includes the following steps:

[0144] S2221: When the target sound system does not have a sky speaker, determine the accompaniment gain of each speaker based on the first accompaniment constraint model, the speaker's vocal width coefficient, and the listening distance.

[0145] After determining the sound field width coefficient of the loudspeaker, it is determined whether there is a sky loudspeaker in the target sound system. When it is determined that there is no sky loudspeaker in the target sound system, there is no need to consider the difference of sky loudspeakers. At this time, the accompaniment gain of each loudspeaker can be determined directly based on the first human voice constraint model, the accompaniment width coefficient of each loudspeaker, and the straight-line distance (i.e., listening distance) between each loudspeaker and the actual listening position.

[0146] S2232: When the target audio system has sky speakers, the multiple speakers are divided into a first speaker group and a second speaker group.

[0147] S2233: For the first loudspeaker group, the accompaniment gain of each loudspeaker is determined based on the first accompaniment constraint model, the vocal width coefficient of each loudspeaker, and the listening distance.

[0148] S2234: For the second speaker group, determine the accompaniment gain of each speaker based on the second accompaniment constraint model and the listening distance of each speaker.

[0149] Specifically, for the second speaker group, based on the second accompaniment constraint model and the listening distance of each speaker, the accompaniment gain of each speaker is determined, including the following steps:

[0150] S22231: Determine the first included angle of each speaker.

[0151] The process of determining the first included angle of the loudspeaker is described above and will not be repeated here.

[0152] S22232: Input the first included angle and listening distance of each speaker in the second speaker group into the second accompaniment constraint model for solution to obtain the accompaniment gain of each speaker in the second speaker group.

[0153] Then, the first included angle and listening distance of each speaker in the second speaker group are input into the second accompaniment constraint model for solution, to obtain the accompaniment gain of each speaker in the second speaker group. That is, the first included angle and listening distance of each sky speaker are input into the second voice constraint model for solution, to obtain the voice gain of each sky speaker. Considering that the sky speakers affect the panoramic sound effect, sound field width adjustment is unnecessary, and considering the influence of the sky speaker height on the sound effect, a fourth constraint formula is obtained for each sky speaker, namely the second accompaniment constraint model. In this embodiment, the second accompaniment constraint model is the same as the second identical constraint model.

[0154] When the number of sky speakers is 4, that is, the second speaker group includes 4 speakers (sky speakers), the second accompaniment constraint model is as follows:

[0155]

[0156]

[0157] Where θ1, θ2, θ3, and θ4 represent the first included angles of the first to fourth sky speakers, respectively; P1, P2, P3, and P4 represent the accompaniment gain of the first to fourth sky speakers, respectively; and L1, L2, L3, and L4 represent the listening distances of the first to fourth sky speakers, respectively. The first and second sky speakers form one corresponding speaker group; the third and fourth sky speakers form another corresponding speaker group. The first and third sky speakers are located on the same side, forming a same-side speaker group; the second and third sky speakers are located on the same side, forming a same-side speaker group. (Reference) Figure 7 The first to fourth sky speakers can be the upper left front speaker group Ltf, the upper right front speaker group Rtf, the upper left rear speaker group Ltb, and the upper right rear speaker group Rtb, respectively.

[0158] When the number of sky speakers is 2, that is, when the second speaker group includes two speakers (sky speakers), the second accompaniment constraint model is as follows:

[0159]

[0160]

[0161] Where θ1 and θ2 represent the first included angle of the first sky speaker and the first included angle of the fourth sky speaker, respectively; P1 and P2 represent the accompaniment gain of the first sky speaker and the channel gain of the fourth sky speaker, respectively; L1 and L2 represent the listening distance of the first sky speaker and the listening distance of the fourth sky speaker, respectively.

[0162] In other embodiments, the process of determining the accompaniment gain of each speaker in the second speaker group can refer to the calculation process of the accompaniment gain of each speaker in the first speaker group described above, and will not be repeated here.

[0163] In this embodiment, by determining the first included angle of each speaker, and then directly inputting the first included angle and listening distance of each speaker in the second speaker group into the second accompaniment constraint model for solution, the accompaniment gain of each speaker in the second speaker group is obtained. This clarifies the specific process of determining the accompaniment gain of each speaker in the second speaker group, and provides an accurate data basis for subsequent calculations.

[0164] In one embodiment, step S2221 or S2223, which involves determining the vocal gain of each speaker based on the first accompaniment constraint model, the vocal width coefficient of each speaker, and the listening distance, specifically includes the following steps:

[0165] S211: Determine the first included angle and the second included angle for each speaker.

[0166] When calculating the channel gain of each speaker, it is first necessary to determine the first and second included angles of each speaker based on the actual listening position and the placement position of each speaker.

[0167] S212: Input the vocal width coefficient, first included angle, second included angle and listening distance of each speaker into the first accompaniment constraint model for solution to obtain the accompaniment gain of each speaker.

[0168] In this embodiment, by determining the first and second included angles of each speaker, the vocal width coefficient, the first included angle, the second included angle, and the listening distance of each speaker are input into the first accompaniment constraint model for solution, thereby obtaining the accompaniment gain of each speaker. This clarifies the specific steps for determining the vocal gain of each speaker based on the first accompaniment constraint model, the vocal width coefficient of each speaker, and the listening distance. By using data such as the listening distance, the first included angle, and the second included angle, the relative position of the actual listening position and each speaker is defined from multiple dimensions. Thus, the accompaniment gain of the speaker is calculated based on the relative position of the actual listening position and each speaker, and the vocal width coefficient. This not only allows for the same immersive experience regardless of the listening position, but also avoids widening the vocal signal in the music, thus improving the listening effect.

[0169] Specifically, the first accompaniment constraint model includes the fifth constraint formula, the second constraint formula, and the sixth constraint formula. The vocal width coefficient, first angle, second angle, and listening distance of each speaker are input into the first accompaniment constraint model for solution, yielding the accompaniment gain of each speaker. The calculation process for the accompaniment gain of each speaker is as follows:

[0170] S2121: Obtain multiple corresponding speaker groups in the target audio system.

[0171] The process for determining multiple corresponding speaker groups has been described above and will not be repeated here.

[0172] S2122: Input the first included angle and listening distance of each speaker into the first constraint formula, input the second included angle and listening distance of each speaker in the corresponding speaker group into the second constraint formula, input the vocal width coefficient and listening distance of each speaker into the third constraint formula, and solve to obtain the accompaniment gain of each channel.

[0173] As mentioned above, the second constraint formula is expressed as follows:

[0174]

[0175] Among them, P left With P rig h t Let θ represent the accompaniment gain of each of the two speakers in a corresponding speaker group; left With θ rig h t Let L be the second included angle between two speakers in a corresponding speaker group. 1left With L rig h t These represent the listening distances of two speakers in a corresponding speaker group.

[0176] To achieve better listening results, based on the principle of vector superposition of sound signals, the fifth constraint formula can be obtained. The first constraint formula is expressed as follows:

[0177]

[0178]

[0179] Where n represents the nth speaker; θ n P represents the first included angle of speaker n; n L represents the accompaniment gain of speaker n; n This indicates the listening distance of speaker n. This represents the components of the audio signal from speaker n in the xy plane.

[0180] Since the sound field width factor can affect the sound field width and thus the sound effect, the accompaniment width factor can be used to adjust the accompaniment gain of each speaker according to the listening preferences of different listeners. This leads to the sixth constraint formula for constraining each speaker group, which is expressed as follows:

[0181]

[0182] Where n represents the nth speaker; θ n P represents the first included angle of speaker n; n L represents the accompaniment gain of speaker n; n b represents the listening distance of speaker n; n This represents the accompaniment width coefficient for each speaker n.

[0183] In one embodiment, to ensure that the sound effects on the left and right sides of the listener are similar and to obtain a better listening experience, after obtaining multiple corresponding speaker groups, it is also necessary to determine the speaker groups on the same side of the target audio system (the determination process is as described above). Then, the first included angle and listening distance of each speaker are input into the fifth constraint formula, the second included angle and listening distance of each speaker in the corresponding speaker group are input into the second constraint formula, and the accompaniment width coefficient and listening distance of each speaker in the same side speaker group are input into the sixth constraint formula to solve for the accompaniment gain of each channel.

[0184] At this point, only the accompaniment width coefficient is used to constrain each speaker in the same-side speaker group, thus obtaining the sixth constraint formula for constraining the same-side speaker group. This third constraint formula is expressed as follows:

[0185]

[0186] Where j represents the j-th speaker; θ j P represents the first included angle of speaker j; j L represents the accompaniment gain of speaker j; j Indicates the listening distance of speaker j; b j This represents the accompaniment width coefficient for each speaker j.

[0187] The accompaniment width coefficient is a preset constant vector. By coordinating the vocal width coefficient and the accompaniment width coefficient, the goal is to avoid widening the vocal signal in the music and only widen the accompaniment signal. The accompaniment width coefficients for speaker groups on the same side but located on different sides of the actual listening position can be the same or different, depending on the user's listening preference. Furthermore, the accompaniment width coefficients for each speaker in the same side speaker group can be the same or different, depending on the user's listening preference.

[0188] In one embodiment, step S30, which generates the audio signal for the corresponding speaker based on the human voice signal, the accompaniment sound signal, and the channel gain of each speaker, specifically includes the following steps:

[0189] S31: Determine the subwoofer and subwoofer channel gain in the target audio system.

[0190] After determining the channel gain of each speaker, it is necessary to determine the subwoofer in the target audio system and the channel gain of the subwoofer (subwoofer).

[0191] S32: Perform low-pass filtering on the accompaniment signal to obtain the filtered accompaniment signal, and distribute the channel gain of the subwoofer and the filtered accompaniment signal to the subwoofer to obtain the audio signal of the subwoofer.

[0192] Then, the accompaniment signal is input into a low-pass filter for low-pass filtering to obtain a filtered accompaniment signal. The channel gain of the subwoofer and the filtered accompaniment signal are then assigned to the subwoofer to obtain the subwoofer's audio signal. For the bass channel, the goal is to reproduce the low-frequency accompaniment (drums, bass, etc.) in the music as much as possible. Since step S10 has already separated the vocals and accompaniment in the stereo audio, it is only necessary to perform low-pass filtering on the separated accompaniment and assign it to the subwoofer along with the channel gain of the subwoofer; there is no need to assign a vocal signal. In this embodiment, the channel gain of the subwoofer may only include the accompaniment gain. Directly assigning the accompaniment gain and the filtered accompaniment signal to the subwoofer can reduce the computational load of the subwoofer's vocal gain. In other embodiments, the channel gain of the subwoofer can be a fixed value (e.g., 0). This fixed value is then used as the channel gain of the subwoofer and assigned to the subwoofer along with the filtered accompaniment signal, further reducing the computational load required for sound gain, reducing the load, and increasing the generation speed of multi-channel audio files.

[0193] S33: Distribute the human voice signal, accompaniment signal and the channel gain of the remaining speakers to the corresponding speakers to obtain the audio signal of the corresponding speakers.

[0194] Finally, the human voice signal, accompaniment signal, and channel gain of the remaining speakers are distributed to the corresponding speakers to obtain the audio signals of the corresponding speakers. Then, the audio signals of all speakers are merged to obtain a multi-channel audio file.

[0195] In this embodiment, the subwoofer and its channel gain in the target audio system are determined. Then, the accompaniment signal is low-pass filtered to obtain a filtered accompaniment signal. The subwoofer's channel gain and the filtered accompaniment signal are then assigned to the subwoofer to obtain its audio signal. Simultaneously, the vocal signal, accompaniment signal, and the channel gains of the remaining speakers are assigned to their respective speakers to obtain their audio signals. This clarifies the specific steps for generating the corresponding speaker audio signals based on the vocal signal, accompaniment signal, and speaker channel gains. Directly assigning the subwoofer's accompaniment gain and the filtered accompaniment signal to the subwoofer ensures good bass channel sound quality while reducing the computational load on the subwoofer's channel gain. In other embodiments, low-pass filtering may be omitted.

[0196] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0197] In one embodiment, a multi-channel audio file generation apparatus is provided, which corresponds one-to-one with the multi-channel audio file generation method described in the above embodiments. For example... Figure 8 As shown, the multi-channel audio file generation device includes a separation module 801, a determination module 802, and a generation module 803. Detailed descriptions of each functional module are as follows:

[0198] The separation module 801 is used to separate the human voice from the input stereo audio file to obtain the human voice signal and the accompaniment signal;

[0199] The determination module 802 is used to determine the arrangement positions of multiple speakers in the target audio system, and to determine the channel gain of each speaker based on the arrangement positions of the multiple speakers.

[0200] The generation module 803 is used to generate the corresponding audio signal of the speaker based on the human voice signal, the accompaniment sound signal and the channel gain of the speaker, and merge them into a multi-channel audio file adapted to the target audio system.

[0201] Optionally, module 802 is specifically used for:

[0202] Set an actual listening position in the target audio system, and determine the listening distance of each speaker based on the actual listening position and the speaker placement. The listening distance is the straight-line distance between the speaker and the actual listening position.

[0203] The channel gain of each speaker is determined based on the listening distance of each speaker.

[0204] Optionally, the channel gain includes vocal gain and accompaniment gain, and the determination module 802 is further used for:

[0205] Determine the sound field width factor for each loudspeaker, which includes the vocal width factor and the accompaniment width factor;

[0206] The vocal gain of each speaker is determined based on its vocal width factor and listening distance.

[0207] The accompaniment gain of each speaker is determined based on the accompaniment width factor and listening distance of each speaker.

[0208] Optionally, the determining module 802 is further used for:

[0209] When the target audio system has sky speakers, the multiple speakers are divided into a first speaker group and a second speaker group, with the second speaker group consisting of sky speakers;

[0210] For the first loudspeaker group, the voice gain of each loudspeaker is determined based on the first human voice constraint model, the human voice width coefficient of each loudspeaker, and the listening distance.

[0211] For the second speaker group, the voice gain of each speaker is determined based on the second voice constraint model and the listening distance of each speaker. The constraint formula of the second voice constraint model is different from that of the first voice constraint model.

[0212] Optionally, the determining module 802 is further used for:

[0213] Determine the first included angle of each speaker. The first included angle is the angle between the speaker ray and the positive direction of the y-axis. The speaker ray is the ray formed by the actual listening position pointing to the speaker. The positive direction of the y-axis is the direction from the origin of the coordinate system to the center speaker.

[0214] Determine the second included angle for each speaker, which is the horizontal included angle of the speaker ray;

[0215] The human voice width coefficient, first included angle, second included angle, and listening distance of each speaker are input into the first human voice constraint model for solution to obtain the human voice gain of each speaker.

[0216] Optionally, the generation module 803 is specifically used for:

[0217] Determine the subwoofer and subwoofer channel gain in the target audio system;

[0218] The accompaniment signal is low-pass filtered to obtain the filtered accompaniment signal. The channel gain of the subwoofer and the filtered accompaniment signal are then distributed to the subwoofer to obtain the subwoofer audio signal.

[0219] The human voice signal, accompaniment signal, and channel gain of the remaining speakers are distributed to the corresponding speakers to obtain the audio signals of the corresponding speakers.

[0220] Optionally, the separation module 801 is specifically used for:

[0221] Obtain a pre-trained stereo separation model, which includes a signal transformation layer and a signal filter;

[0222] The stereo audio file is input into the stereo separation model, and Fourier transform is performed through the signal transformation layer to obtain the transformed audio signal. The transformed audio signal is then separated into human voice and accompaniment signals through the signal filter.

[0223] Specific limitations regarding the multi-channel audio file generation device can be found in the limitations of the multi-channel audio file generation method described above, and will not be repeated here. Each module in the aforementioned multi-channel audio file generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in a computer device, or stored in software in the memory of a computer device, so that the processor can call and execute the corresponding operations of each module.

[0224] In one embodiment, a computer device is provided, which may be a server. The computer device includes a processor, memory, a network interface, and a database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data used and generated by a multi-channel audio file generation method, such as stereo separation models, stereo audio files, separated vocal and accompaniment signals, and multi-channel audio files. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a multi-channel audio file generation method.

[0225] In one embodiment, such as Figure 8 As shown, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-described multi-channel audio file generation method.

[0226] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-channel audio file generation method described above.

[0227] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Furthermore, any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory.

[0228] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0229] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method of generating a multi-channel audio file, characterized by, The application relates to a method for generating a multi-channel audio file adapted to a target sound system. The method comprises: separating an input stereo audio file to obtain a vocal signal and an accompaniment signal; determining the arrangement positions of a plurality of loudspeakers in the target sound system, and determining the channel gains of the loudspeakers according to the arrangement positions of the loudspeakers; generating audio signals corresponding to the loudspeakers based on the vocal signal, the accompaniment signal and the channel gains of the loudspeakers, and merging the audio signals into a multi-channel audio file adapted to the target sound system; the determination of the channel gains of the loudspeakers according to the arrangement positions of the loudspeakers comprises: setting an actual listening position in the target sound system, and determining the listening distances of the loudspeakers according to the actual listening position and the arrangement positions of the loudspeakers, wherein the listening distance is the straight-line distance between the loudspeaker and the actual listening position; determining the channel gains of the loudspeakers according to the listening distances of the loudspeakers; 2. The multi-channel audio file generating method of claim 1, wherein, the channel gains comprise vocal gains and accompaniment gains, and the determination of the channel gains of the loudspeakers according to the listening distances of the loudspeakers comprises: determining the sound field width coefficients of the loudspeakers, wherein the sound field width coefficients comprise vocal width coefficients and accompaniment width coefficients; determining the vocal gains of the loudspeakers according to the vocal width coefficients and the listening distances of the loudspeakers; and determining the accompaniment gains of the loudspeakers according to the accompaniment width coefficients and the listening distances of the loudspeakers. the determination of the vocal gains of the loudspeakers according to the vocal width coefficients and the listening distances of the loudspeakers comprises: when the target sound system has a sky loudspeaker, the loudspeakers are divided into a first loudspeaker group and a second loudspeaker group, and the second loudspeaker group is composed of the sky loudspeaker; for the first loudspeaker group, the vocal gains of the loudspeakers are determined based on a first vocal constraint model, the vocal width coefficients and the listening distances of the loudspeakers; 3. The multi-channel audio file generating method of claim 2, wherein, for the second loudspeaker group, the vocal gains of the loudspeakers are determined based on a second vocal constraint model and the listening distances of the loudspeakers, and the second vocal constraint model has a constraint formula different from that of the first vocal constraint model. the determination of the vocal gains of the loudspeakers based on the first vocal constraint model, the vocal width coefficients and the listening distances of the loudspeakers comprises: determining the first included angles of the loudspeakers, wherein the first included angle is the included angle between a loudspeaker ray and the positive direction of a y-axis, the loudspeaker ray is a ray formed by the actual listening position pointing to the loudspeaker, and the positive direction of the y-axis is the direction of the coordinate origin pointing to a center loudspeaker; determining the second included angles of the loudspeakers, wherein the second included angle is the horizontal included angle of the loudspeaker ray; 4. The multi-channel audio file generation method according to any one of claims 1 to 3, characterized in that, inputting the vocal width coefficients, the first included angles, the second included angles and the listening distances of the loudspeakers into the first vocal constraint model to obtain the vocal gains of the loudspeakers. the generation of the audio signals corresponding to the loudspeakers based on the vocal signal, the accompaniment signal and the channel gains of the loudspeakers comprises: determining a bass speaker in the target sound system and a channel gain of the bass speaker; low-pass filtering the accompaniment sound signal to obtain a filtered accompaniment sound signal, and distributing the channel gain of the bass speaker and the filtered accompaniment sound signal to the bass speaker to obtain an audio signal of the bass speaker; distributing the vocal signal, the accompaniment sound signal and the channel gain of the remaining speakers to the corresponding speakers to obtain audio signals of the corresponding speakers.

5. The multi-channel audio file generation method according to any one of claims 1 to 3, characterized in that, The method comprises the following steps: obtaining a pre-trained stereo separation model, wherein the stereo separation model comprises a signal transformation layer and a signal filter; inputting the stereo audio file into the stereo separation model, performing Fourier transform on the stereo audio file through the signal transformation layer to obtain a transformed audio signal, and performing vocal separation on the transformed audio signal through the signal filter to obtain the vocal signal and the accompaniment sound signal.

6. A multi-channel audio file generating apparatus characterized by comprising: The method comprises the following steps: a separation module configured to separate a vocal signal and an accompaniment sound signal from an input stereo audio file; a determination module configured to determine arrangement positions of a plurality of speakers in a target sound system, and determine channel gains of the speakers according to the arrangement positions of the speakers; a generation module configured to generate audio signals of the speakers based on the vocal signal, the accompaniment sound signal and the channel gains of the speakers, and combine the audio signals into a multi-channel audio file adapted to the target sound system; The method comprises the following steps: setting an actual listening position in the target sound system, and determining listening distances of the speakers according to the actual listening position and the arrangement positions of the speakers, wherein the listening distance is a straight-line distance between the speaker and the actual listening position; determining the channel gains of the speakers according to the listening distances of the speakers; The channel gains comprise vocal gains and accompaniment gains, and the determination of the channel gains of the speakers according to the listening distances of the speakers comprises: determining sound field width coefficients of the speakers, wherein the sound field width coefficients comprise vocal width coefficients and accompaniment width coefficients; determining the vocal gains of the speakers according to the vocal width coefficients and the listening distances of the speakers; and determining the accompaniment gains of the speakers according to the accompaniment width coefficients and the listening distances of the speakers.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the multi-channel audio file generation method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising: The computer program is executed by the processor to implement the steps of the multi-channel audio file generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-channel audio system capable of tracking user and equipment with multi-channel audio system

    CN106686520A

  • Audio signal processing method and device and computer readable storage medium

    CN113347552A