Multi-view audio generation method and device, equipment, storage medium and program product

By arranging the radio unit around the camera unit and giving gain weight, the problem of audio data not being synchronously switched in multi-view video is solved, and the synchronous switching between audio and viewing angle is realized, improving the user experience.

CN120416618APending Publication Date: 2025-08-01MIGU VIDEO TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510694525.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, the audio data cannot be switched synchronously when switching viewing angles, which affects the user's on-site experience.

Method used

By arranging the radio unit around the camera unit corresponding to each viewing angle, different gain weights are assigned according to the position distance between the radio unit and the viewing angle, audio data of each viewing angle is obtained, and sound field synchronous switching is achieved by assigning the gain weight to the sliding position during viewing angle switching.

Benefits of technology

It realizes synchronous switching between audio data and perspective, enhancing the user's immersive on-site experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416618A_ABST
    Figure CN120416618A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view audio generation method and device, equipment, a storage medium and a program product, and relates to the technical field of 3D audio and video processing. The method comprises the following steps: acquiring a first playing request of video data, wherein the first playing request is a playing request of the video data at a first view angle; and obtaining first audio data corresponding to the video data of the first view angle according to audio signals collected by at least two sound receiving units related to the first view angle and the gain weight corresponding to each sound receiving unit. According to the scheme provided by the invention, the sound field can be associated with the visual angles, and the video data of each visual angle corresponds to different audio data, so that a better immersive on-site experience feeling can be brought to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 3D audio - video processing, and particularly to a multi - perspective audio generation method, device, equipment, storage medium and program product. Background Art

[0002] In the prior art, video data supporting multiple perspectives is available, but the video data of each perspective only supports using unified audio data. Thus, when switching perspectives, the sound field is not associated with the perspective, and the audio data does not synchronously switch with the perspective switch, thereby affecting the user's sense of presence experience. Summary of the Invention

[0003] Embodiments of the present invention provide a multi - perspective audio generation method, device, equipment, storage medium and program product to solve the problem that audio data in the prior art does not synchronously switch with the perspective switch.

[0004] In a first aspect, embodiments of the present invention provide a multi - perspective audio generation method, including:

[0005] Obtain a first playback request for video data, where the first playback request is a playback request for the video data of a first perspective;

[0006] According to the audio signals collected by at least two sound - collecting units related to the first perspective and the gain weight corresponding to each sound - collecting unit, obtain first audio data corresponding to the video data of the first perspective.

[0007] Optionally, the method further includes:

[0008] For each sound - collecting unit related to the first perspective, according to the position indication information of the first perspective and the position indication information of the sound - collecting unit, obtain the gain weight corresponding to the sound - collecting unit;

[0009] Among them, the farther the sound - collecting unit is from the position of the first perspective, the smaller the corresponding gain weight.

[0010] Optionally, each perspective corresponds to a camera unit for obtaining video data, and one sound - collecting unit is respectively arranged on both sides of the camera unit;

[0011] The first audio data includes left - channel audio data and right - channel audio data;

[0012] In the case of obtaining the left-channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: a sound collection unit on the left side of the camera unit corresponding to the first perspective and a sound collection unit on the left side of the camera unit corresponding to at least one left perspective, where at least one of the left perspectives is a perspective set on the left side of the first perspective;

[0013] In the case of obtaining the right-channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: a sound collection unit on the right side of the camera unit corresponding to the first perspective and a sound collection unit on the right side of the camera unit corresponding to at least one right perspective, where at least one of the right perspectives is a perspective set on the right side of the first perspective.

[0014] Optionally, the method further includes:

[0015] Obtaining a perspective switching request for the video data, where the perspective switching request is a playback request for the video data to switch from the first perspective to the second perspective, and the second perspective is a perspective adjacent to the first perspective;

[0016] Obtaining the audio data corresponding to the video data after the perspective switch according to the first audio data and the second audio data corresponding to the video data of the second perspective.

[0017] Optionally, obtaining the audio data corresponding to the video data after the perspective switch according to the first audio data and the second audio data corresponding to the video data of the second perspective includes:

[0018] Obtaining the gain weight corresponding to the first audio data and the gain weight corresponding to the second audio data according to the switching progress information indicated by the perspective switching request;

[0019] Obtaining the audio data corresponding to the video data after the perspective switch according to the first audio data, the second audio data, and the gain weights corresponding to the first audio data and the second audio data respectively.

[0020] Optionally, the switching progress information includes the sliding position input by the user;

[0021] The gain weight corresponding to the first audio data is related to the distance between the sliding position and the sliding start point, and the sliding start point corresponds to the first perspective;

[0022] The gain weight corresponding to the second audio data is related to the distance between the sliding position and the sliding end point, and the sliding end point corresponds to the second perspective.

[0023] In a second aspect, an embodiment of the present invention further provides a multi-perspective audio generation device, including:

[0024] A first acquisition module, configured to acquire a first playback request for video data, where the first playback request is a playback request for the video data from a first perspective;

[0025] A second acquisition module, configured to acquire first audio data corresponding to the video data from the first perspective according to audio signals collected by at least two sound collection units related to the first perspective and gain weights corresponding to each sound collection unit.

[0026] In a third aspect, an embodiment of the present invention further provides a multi - perspective audio generation device, including: a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor; the processor is configured to read the program in the memory to implement the multi - perspective audio generation method as described in the first aspect.

[0027] In a fourth aspect, an embodiment of the present invention further provides a computer - readable storage medium, configured to store a computer program, where the computer program, when executed by a processor, implements the multi - perspective audio generation method as described in the first aspect.

[0028] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including computer instructions, where the computer instructions, when executed by a processor, implement the multi - perspective audio generation method as described in the first aspect.

[0029] In an embodiment of the present invention, a first playback request for video data is acquired, where the first playback request is a playback request for the video data from a first perspective; first audio data corresponding to the video data from the first perspective is acquired according to audio signals collected by at least two sound collection units related to the first perspective and gain weights corresponding to each sound collection unit. The audio data corresponding to the video data of each perspective is different, and the sound field is associated with the perspective, realizing synchronous switching of audio data and perspective, thus greatly enhancing the user's immersive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for description in the embodiments of the present invention. Obviously, the following - described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0031] Figure 1 It is a flowchart of the multi - perspective audio generation method provided by an embodiment of the present invention;

[0032] Figure 2 It is a schematic diagram of an application scenario of the multi - perspective audio generation method provided by an embodiment of the present invention;

[0033] Figure 3 Schematic diagram of the relationship between the sliding position and the gain weight during the perspective switching process provided by the embodiment of the present invention;

[0034] Figure 4 Block diagram of the multi-perspective audio generation device provided by the embodiment of the present invention;

[0035] Figure 5 Schematic diagram of the hardware structure of the multi-perspective audio generation device provided by the embodiment of the present invention. Detailed implementation manners

[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] See Figure 1 , Figure 1 is the flowchart of the multi-perspective audio generation method provided by the embodiment of the present invention. As shown in Figure 1 below, it includes the following steps:

[0038] Step 101, obtain a first playback request for video data, where the first playback request is a playback request for the video data from a first perspective, and the first perspective is one of multiple perspectives of the video data;

[0039] Step 102, obtain first audio data corresponding to the video data from the first perspective according to the audio signals collected by at least two sound collection units related to the first perspective and the gain weight corresponding to each sound collection unit.

[0040] Among them, the at least two sound collection units related to the first perspective include at least two sound collection units within a preset distance range from the camera unit corresponding to the first perspective. Optionally, the number of sound collection units related to the first perspective can be set according to experience. The preset distance range can also be set according to experience.

[0041] The first audio data is related to the number of channels and satisfies one of the following:

[0042] In a mono scenario, the first audio data includes mono audio data;

[0043] In a stereo (two-channel) scenario, the first audio data includes left-channel audio data and right-channel audio data;

[0044] In a surround sound scenario, the first audio data includes multi-channel audio data.

[0045] Further, the video data of the first perspective is mixed and encoded with the corresponding first audio data to output the audio-visual data of the first perspective.

[0046] In a specific embodiment, the above method further includes:

[0047] For each sound collection unit related to the first perspective, according to the position indication information of the first perspective and the position indication information of the sound collection unit, obtain the gain weight corresponding to the sound collection unit;

[0048] Among them, the farther the sound collection unit is from the position of the first perspective, the smaller the corresponding gain weight.

[0049] As Figure 2 shown, a schematic diagram of the application scenario of the multi-perspective audio generation method provided by the embodiment of the present invention. This application scenario belongs to a stereo scenario. In Figure 2 , a triangular pattern represents a perspective. One perspective corresponds to one camera unit, which collects the video signal of this perspective for obtaining video data. There are a total of twelve perspectives, corresponding to twelve camera units. These twelve camera units are equally angularly arranged, and the positions of these twelve camera units are numbered according to the clock position, that is, the 1 o'clock camera unit, the 2 o'clock camera unit, the 3 o'clock camera unit,..., the 11 o'clock camera unit, the 12 o'clock camera unit. For example Figure 2 the 6 o'clock camera unit and the 12 o'clock camera unit shown in Figure 2 are shown, so as to obtain the position indication information of each camera unit. For example

[0050] It can be understood that the position indication information of each camera unit is also the position indication information of each perspective.

[0051] One sound collection unit is arranged equidistantly on both sides of each camera unit to collect audio signals. It can be understood that one sound collection unit is arranged equidistantly on both sides of each camera unit, that is, one sound collection unit is arranged between two adjacent camera units. This sound collection unit can be a MIC (Microphone), and there are a total of twelve sound collection units. These twelve sound collection units are also equally angularly arranged, and the positions of these twelve sound collection units are also numbered according to the clock position, that is Figure 2 the 0.5 sound collection unit, the 1.5 sound collection unit, the 2.5 sound collection unit,..., the 10.5 sound collection unit, the 11.5 sound collection unit shown in Figure 2The position indication information of the 5.5 sound collection unit in is "5.5", and the position indication information of the 1.5 sound collection unit is "1.5".

[0052] In the embodiment of the present invention, the gain weight corresponding to each sound collection unit related to the first viewing angle is negatively correlated with the position of each sound collection unit from the first viewing angle, that is, the farther the sound collection unit is from the first viewing angle, the smaller the corresponding gain weight.

[0053] To ensure the rapid convergence of the gain weight, an inverse function algorithm is adopted, as shown in the following formula (1), to obtain the gain weight.

[0054]

[0055] Where n represents the position indication information of the first viewing angle; x represents the position indication information of the sound collection unit; y represents the gain weight corresponding to the sound collection unit.

[0056] Here, taking Figure 2 the viewing angle corresponding to the 6 o'clock camera unit in as the first viewing angle n = 6, and the sound collection units related to the first viewing angle may include the 5.5 sound collection unit, 4.5 sound collection unit, 3.5 sound collection unit, 6.5 sound collection unit, 7.5 sound collection unit, 8.5 sound collection unit as an example, to illustrate the gain weight corresponding to each sound collection unit:

[0057] The position indication information of the 5.5 sound collection unit is x = 5.5, and the gain weight corresponding to the 5.5 sound collection unit is y = 1;

[0058] The position indication information of the 4.5 sound collection unit is x = 4.5, and the gain weight corresponding to the 4.5 sound collection unit is y = 1 / 3;

[0059] The position indication information of the 3.5 sound collection unit is x = 3.5, and the gain weight corresponding to the 3.5 sound collection unit is y = 1 / 5;

[0060] The position indication information of the 6.5 sound collection unit is x = 6.5, and the gain weight corresponding to the 6.5 sound collection unit is y = 1;

[0061] The position indication information of the 7.5 sound collection unit is x = 7.5, and the gain weight corresponding to the 7.5 sound collection unit is y = 1 / 3;

[0062] The position indication information of the 8.5 sound collection unit is x = 8.5, and the gain weight corresponding to the 8.5 sound collection unit is y = 1 / 5.

[0063] Further, according to the audio signals collected by each sound collection unit related to the first viewing angle and the corresponding gain weights, weighted summation is performed to obtain the audio data corresponding to the video data of the first viewing angle.

[0064] In one embodiment, each of the perspectives corresponds to a camera unit, the camera unit is used to acquire video data, and a sound collection unit is disposed on each side of the camera unit;

[0065] The first audio data includes: left-channel audio data and right-channel audio data;

[0066] When acquiring the left-channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: the sound collection unit on the left side of the camera unit corresponding to the first perspective and the sound collection units on the left side of the camera units corresponding to at least one left-side perspective, and at least one of the left-side perspectives is a perspective disposed on the left side of the first perspective;

[0067] When acquiring the right-channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: the sound collection unit on the right side of the camera unit corresponding to the first perspective and the sound collection units on the right side of the camera units corresponding to at least one right-side perspective, and at least one of the right-side perspectives is a perspective disposed on the right side of the first perspective.

[0068] As Figure 2 shown, taking the perspective corresponding to the 6 o'clock camera unit in Figure 2 as the first perspective as an example, the right-side perspectives may sequentially include the perspectives corresponding to the 5 o'clock camera unit and the 4 o'clock camera unit, and the left-side perspectives may sequentially include the perspectives corresponding to the 7 o'clock camera unit and the 8 o'clock camera unit.

[0069] When acquiring the right-channel audio data, the sound collection units related to the first perspective may include the 5.5 sound collection unit on the right side of the 6 o'clock camera unit, the 4.5 sound collection unit on the right side of the 5 o'clock camera unit, and the 3.5 sound collection unit on the right side of the 4 o'clock camera unit.

[0070] When acquiring the left-channel audio data, the sound collection units related to the first perspective may include the 6.5 sound collection unit on the left side of the 6 o'clock camera unit, the 7.5 sound collection unit on the left side of the 7 o'clock camera unit, and the 8.5 sound collection unit on the left side of the 8 o'clock camera unit.

[0071] From the above embodiments, the gain weights corresponding to these sound collection units can be obtained, and thus the left-channel audio data can be obtained, which is expressed as follows:

[0072] y6.1 = (1 * the audio signal collected by the 6.5 sound collection unit + 1 / 3 * the audio signal collected by the 7.5 sound collection unit + 1 / 5 * the audio signal collected by the 8.5 sound collection unit) / (1 + 1 / 3 + 1 / 5);

[0073] And, the right-channel audio data is obtained, which is expressed as follows:

[0074] y6.2 = (1 * the audio signal collected by the 5.5 microphone unit + 1 / 3 * the audio signal collected by the 4.5 microphone unit + 1 / 5 * the audio signal collected by the 3.5 microphone unit) / (1 + 1 / 3 + 1 / 5).

[0075] For ease of understanding, Table 1 below is used to illustrate the first audio data.

[0076]

[0077]

[0078] Table 1

[0079] In one embodiment, each of the perspectives corresponds to a camera unit for acquiring video data, and one of the microphone units is provided at the camera unit.

[0080] The first audio data includes: mono audio data;

[0081] When acquiring the mono audio data corresponding to the video data of the first perspective, at least two microphone units related to the first perspective include: the microphone unit provided at the camera unit corresponding to the first perspective and the microphone units provided at the camera units corresponding to at least one adjacent perspective, where at least one of the adjacent perspectives is adjacent to the first perspective.

[0082] In one embodiment, each of the perspectives corresponds to a camera unit for acquiring video data, and at least five of the microphone units are provided around the camera unit. Taking the five microphone units as an example, specifically, it includes the first microphone unit on the left front side of the camera unit, the second microphone unit on the right front side of the camera unit, the third microphone unit at the camera unit, the fourth microphone unit on the left rear side of the camera unit, and the fifth microphone unit on the right rear side of the camera unit;

[0083] The first audio data includes: left front channel audio data, right front channel audio data, center channel audio data, left rear channel audio data, and right rear channel audio data;

[0084] When acquiring the left front channel audio data corresponding to the video data of the first perspective, at least two microphone units related to the first perspective include: the microphone unit on the left front side of the camera unit corresponding to the first perspective and the microphone units on the left front side of the camera units corresponding to at least one left - hand perspective, where at least one of the left - hand perspectives is a perspective to the left of the first perspective.

[0085] When obtaining the left rear channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: a sound collection unit on the left rear side of the camera unit corresponding to the first perspective and at least one sound collection unit on the left rear side of the camera unit corresponding to at least one left perspective, where at least one of the left perspectives is a perspective on the left side of the first perspective;

[0086] When obtaining the right front channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: a sound collection unit on the right front side of the camera unit corresponding to the first perspective and at least one sound collection unit on the right front side of the camera unit corresponding to at least one right perspective, where at least one of the right perspectives is a perspective on the right side of the first perspective;

[0087] When obtaining the right rear channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: a sound collection unit on the right rear side of the camera unit corresponding to the first perspective and at least one sound collection unit on the right rear side of the camera unit corresponding to at least one right perspective, where at least one of the right perspectives is a perspective on the right side of the first perspective;

[0088] When obtaining the center channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: a sound collection unit arranged at the camera unit corresponding to the first perspective and at least one sound collection unit arranged at the camera unit corresponding to at least one adjacent perspective, where at least one of the adjacent perspectives is adjacent to the first perspective.

[0089] In an embodiment, the above method further includes:

[0090] Obtaining a perspective switching request for the video data, where the perspective switching request is a playback request for the video data to switch from the first perspective to the second perspective, and the second perspective is a perspective adjacent to the first perspective;

[0091] Obtaining the audio data corresponding to the video data after switching perspectives according to the first audio data and the second audio data corresponding to the video data of the second perspective.

[0092] In the embodiments of the present invention, the perspective switching is synchronized with the sound field switching. The user can trigger the perspective switching request through a sliding screen operation on the terminal player. The perspective switching direction corresponds to the sliding direction. For example, sliding the screen to the left means switching to a perspective on the left side of the first perspective, and the second perspective is the perspective on the left side of the first perspective; sliding the screen to the right means switching to a perspective on the right side of the first perspective, and the second perspective is the perspective on the right side of the first perspective.

[0093] It should be noted that the perspective switching is a gradual process, and the audio data is gradually switched synchronously during the perspective switching process.

[0094] In one embodiment, obtaining the audio data corresponding to the video data after switching perspectives according to the first audio data and the second audio data corresponding to the video data of the second perspective includes:

[0095] Obtaining the gain weight corresponding to the first audio data and the gain weight corresponding to the second audio data according to the switching progress information indicated by the perspective switching request;

[0096] Obtaining the audio data corresponding to the video data after switching perspectives according to the first audio data, the second audio data, and the gain weights corresponding to the first audio data and the second audio data respectively.

[0097] During the perspective switching process, obtaining the gain weight corresponding to the real-time first audio data and the gain weight corresponding to the second audio data according to the real-time switching progress information, so as to obtain the audio data corresponding to the video data after the real-time perspective switching. The progress of the perspective switching is linearly synchronized with the progress of the sound field switching.

[0098] In a specific embodiment, the switching progress information includes the sliding position input by the user;

[0099] The gain weight corresponding to the first audio data is related to the distance between the sliding position and the sliding start point, and the sliding start point corresponds to the first perspective;

[0100] The gain weight corresponding to the second audio data is related to the distance between the sliding position and the sliding end point, and the sliding end point corresponds to the second perspective.

[0101] Assume that the first perspective is the perspective corresponding to the 6 o'clock camera unit, the left-channel audio data corresponding to the video data of the first perspective is y6.1, the right-channel audio data corresponding to the video data of the first perspective is y6.2, the second perspective is the perspective corresponding to the 5 o'clock camera unit, the left-channel audio data corresponding to the video data of the second perspective is y5.1, the right-channel audio data corresponding to the video data of the second perspective is y5.2, the distance between the sliding start point and the sliding end point is L centimeters, the distance between the sliding position and the sliding start point is N centimeters, and the distance between the sliding position and the sliding end point is L - N centimeters.

[0102] As Figure 3 shown, Figure 3 is a schematic diagram of the relationship between the sliding position and the gain weight during the perspective switching process provided by the embodiment of the present invention. The left-channel audio data corresponding to the video data after switching perspectives is expressed as follows:

[0103] y6.1-5.1 = N / L * y6.1 + (L - N) / L * y5.1;

[0104] In addition, the right-channel audio data corresponding to the video data after the perspective is switched is expressed as follows:

[0105] y6.2 - 5.2 = N / L * y6.2 + (L - N) / L * y5.2.

[0106] Therefore, in the process of perspective switching in the embodiments of the present invention, the balanced switching of the sound field can be synchronously achieved.

[0107] The embodiments of the present invention provide a multi-perspective audio generation system, including a sound collection system, a source station, a CDN (Content Delivery Network), and a terminal player. Among them, the sound collection system is used to generate video data for each perspective according to the video signals collected by the camera units corresponding to each perspective, perform a primary gain weighting on the audio signals collected by each sound collection unit according to the gain weight corresponding to each sound collection unit in at least two sound collection units related to each perspective, generate the audio data corresponding to the video data for each perspective, and perform mixed-stream encoding on the video data and the corresponding audio data for each perspective through an online encoder to generate the audio-visual data for each perspective, and then distribute it to the CDN through the source station.

[0108] Next, the terminal player sends a multi-perspective video play request to the CDN, obtains the play addresses of all perspectives, and then sends a first play request to the CDN to request the audio-visual data of the first perspective. The audio-visual data of the first perspective is obtained through the CDN, and the video data and the first audio data are respectively decoded from the audio-visual data, and the first audio data may include left-channel audio data and right-channel audio data.

[0109] It should be noted that while the terminal player is playing the audio-visual data of the first perspective, it preloads the audio-visual data of the two perspectives on both sides of the first perspective. When the user triggers a perspective switching request through a screen sliding operation, the terminal player monitors the sliding start point and the sliding direction. According to whether the user slides the screen to the left or to the right, it determines the second perspective that the user wants to switch to (that is, when the user slides the screen to the left, it represents switching to the video data of the second perspective on the left side of the first perspective). At this time, the terminal decodes the left-channel audio data and the right-channel audio data in the second audio data from the preloaded audio-visual data of the second perspective, and performs a secondary weighted gain on the left-channel audio data in the first audio data corresponding to the video data of the first perspective and the left-channel data in the second audio data corresponding to the video data of the second perspective, and performs a secondary gain weighting on the right-channel audio data in the first audio data corresponding to the video data of the first perspective and the right-channel data in the second audio data corresponding to the video data of the second perspective.

[0110] In summary, by adopting the multi-view audio generation method according to the embodiments of the present invention, a sound collection unit is installed and arranged around the camera unit corresponding to each view to collect audio signals, and different gain weights are assigned to the sound collection units according to the distance between the sound collection units and the position of the view, that is, primary gain weighting, so as to obtain the audio data corresponding to the video data of each view, and realize the change of the sound field with the change of the view. Moreover, when the view is switched, the sound field is switched synchronously and balancedly. The gain weight is assigned according to the distances between the sliding position of the user on the screen and the sliding start point and the sliding end point respectively, that is, secondary gain weighting, so as to obtain the audio data corresponding to the video data after the view is switched in real time, and realize the synchronous and balanced switching and transfer of the sound field during the process of the user sliding on the screen, which can give the user a better immersive on-site experience.

[0111] See Figure 4 , the embodiments of the present invention further provide a multi-view audio generation device, and the device 400 includes:

[0112] A first acquisition module 401, configured to acquire a first playback request of video data, where the first playback request is a playback request of the video data of a first view;

[0113] A second acquisition module 402, configured to acquire first audio data corresponding to the video data of the first view according to the audio signals collected by at least two sound collection units related to the first view and the gain weight corresponding to each sound collection unit.

[0114] Optionally, the device further includes:

[0115] A third acquisition module, configured to obtain, for each sound collection unit related to the first view, the gain weight corresponding to the sound collection unit according to the position indication information of the first view and the position indication information of the sound collection unit;

[0116] Among them, the farther the sound collection unit is from the position of the first view, the smaller the corresponding gain weight.

[0117] Optionally, each view corresponds to a camera unit, the camera unit is configured to acquire video data, and one sound collection unit is respectively arranged on both sides of the camera unit;

[0118] The first audio data includes left-channel audio data and right-channel audio data;

[0119] When acquiring the left-channel audio data corresponding to the video data of the first view, at least two sound collection units related to the first view include: the sound collection unit on the left side of the camera unit corresponding to the first view and the sound collection units on the left side of the camera units corresponding to at least one left-side view, and at least one of the left-side views is a view set on the left side of the first view;

[0120] In the case of obtaining the right-channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: a sound collection unit on the right side of the camera unit corresponding to the first perspective and a sound collection unit on the right side of the camera unit corresponding to at least one right-side perspective, and at least one of the right-side perspectives is a perspective set on the right side of the first perspective.

[0121] Optionally, the apparatus further includes:

[0122] A fourth acquisition module, configured to acquire a perspective switching request for the video data, where the perspective switching request is a playback request for the video data to switch from the first perspective to a second perspective, and the second perspective is a perspective adjacent to the first perspective;

[0123] A fifth acquisition module, configured to acquire the audio data corresponding to the video data after the perspective is switched according to the first audio data and the second audio data corresponding to the video data of the second perspective.

[0124] Optionally, the fifth acquisition module is specifically configured to:

[0125] Acquire a gain weight corresponding to the first audio data and a gain weight corresponding to the second audio data according to the switching progress information indicated by the perspective switching request;

[0126] Acquire the audio data corresponding to the video data after the perspective is switched according to the first audio data, the second audio data, and the gain weights corresponding to the first audio data and the second audio data respectively.

[0127] Optionally, the switching progress information includes a sliding position input by a user;

[0128] The gain weight corresponding to the first audio data is related to the distance between the sliding position and the sliding start point, and the sliding start point corresponds to the first perspective;

[0129] The gain weight corresponding to the second audio data is related to the distance between the sliding position and the sliding end point, and the sliding end point corresponds to the second perspective.

[0130] The apparatus provided in the embodiments of the present invention can execute the above method embodiments, and the implementation principles and technical effects are similar, which will not be elaborated here in this embodiment.

[0131] Such as Figure 5As shown in the figure, the multi - perspective audio generation device according to an embodiment of the present invention includes: a processor 500; and a memory 520 connected to the processor 500 through a bus interface. The memory 520 is used to store programs and data used by the processor 500 when performing operations, and the processor 500 calls and executes the programs and data stored in the memory 520.

[0132] A transceiver 510, configured to perform the following processes under the control of the processor 500:

[0133] Obtain a first playback request for video data, where the first playback request is a playback request for the video data from a first perspective;

[0134] A processor 500, configured to read the program in the memory 520 and perform the following processes:

[0135] Obtain first audio data corresponding to the video data from the first perspective according to the audio signals collected by at least two sound - collecting units related to the first perspective and the gain weight corresponding to each sound - collecting unit.

[0136] Among them, in Figure 5 The bus architecture may include any number of interconnected buses and bridges. Specifically, various circuits represented by one or more processors represented by the processor 500 and the memory represented by the memory 520 are linked together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well - known in the art. Therefore, they will not be further described herein. The bus interface provides an interface. The transceiver 510 may be multiple components, that is, including a transmitter and a transceiver, and provides a unit for communicating with various other devices on the transmission medium. For different user devices, the user interface 530 may also be an interface capable of externally or internally connecting required devices, and the connected devices include but are not limited to a keypad, a display, a speaker, a microphone, a joystick, etc.

[0137] The processor 500 is responsible for managing the bus architecture and general processing, and the memory 520 can store data used by the processor 500 when performing operations.

[0138] Optionally, the processor 500 is further configured to read the computer program and perform the following steps:

[0139] For each sound - collecting unit related to the first perspective, obtain the gain weight corresponding to the sound - collecting unit according to the position indication information of the first perspective and the position indication information of the sound - collecting unit;

[0140] Among them, the farther the sound - collecting unit is from the position of the first perspective, the smaller the corresponding gain weight.

[0141] Optionally, each of the perspectives corresponds to a camera unit for acquiring video data, and a sound collection unit is respectively arranged on both sides of the camera unit;

[0142] The first audio data includes left-channel audio data and right-channel audio data;

[0143] When acquiring the left-channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: the sound collection unit on the left side of the camera unit corresponding to the first perspective and the sound collection units on the left side of the camera units corresponding to at least one left-side perspective, where at least one of the left-side perspectives is a perspective arranged on the left side of the first perspective;

[0144] When acquiring the right-channel audio data corresponding to the video data of the first perspective, at least two sound collection units related to the first perspective include: the sound collection unit on the right side of the camera unit corresponding to the first perspective and the sound collection units on the right side of the camera units corresponding to at least one right-side perspective, where at least one of the right-side perspectives is a perspective arranged on the right side of the first perspective.

[0145] Optionally, the transceiver 510 is further configured to perform the following steps under the control of the processor 500:

[0146] Obtain a perspective switching request for the video data, where the perspective switching request is a playback request for the video data to be switched from the first perspective to the second perspective, and the second perspective is a perspective adjacent to the first perspective;

[0147] The processor 500 is further configured to read the computer program and perform the following steps:

[0148] Obtain the audio data corresponding to the video data after switching the perspective according to the first audio data and the second audio data corresponding to the video data of the second perspective.

[0149] Optionally, the processor 500 is specifically configured to read the computer program and perform the following steps:

[0150] Obtain the gain weight corresponding to the first audio data and the gain weight corresponding to the second audio data according to the switching progress information indicated by the perspective switching request;

[0151] Obtain the audio data corresponding to the video data after switching the perspective according to the first audio data, the second audio data, and the gain weights corresponding to the first audio data and the second audio data respectively.

[0152] Optionally, the switching progress information includes the sliding position input by the user;

[0153] The gain weight corresponding to the first audio data is related to the distance between the sliding position and the sliding start position, and the sliding start point corresponds to the first view angle;

[0154] The gain weight corresponding to the second audio data is related to the distance between the sliding position and the sliding end position, and the sliding end point corresponds to the second view angle.

[0155] The multi-view audio generation device provided by the embodiments of the present invention can execute the above method embodiments, and the implementation principles and technical effects are similar, so they will not be elaborated here.

[0156] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the above multi-view audio generation method are implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0157] In addition, the specific embodiments of the present invention also provide a computer program product, including computer instructions. When the computer instructions are executed by a processor, each process of the above multi-view audio generation method embodiment is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0158] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0159] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can be physically included separately, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.

[0160] The integrated unit implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute some steps of the transceiver method described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0161] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A multi-view audio generation method, characterized in that, Including: A first playback request for obtaining video data, where the first playback request is a playback request for the video data from a first perspective; Obtaining first audio data corresponding to the video data from the first perspective according to audio signals collected by at least two sound collection units related to the first perspective and gain weights corresponding to each sound collection unit.

2. The multi-view audio generation method according to claim 1, wherein The method further includes: For each sound collection unit related to the first perspective, obtaining the gain weight corresponding to the sound collection unit according to the position indication information of the first perspective and the position indication information of the sound collection unit; Wherein, the farther the sound collection unit is from the position of the first perspective, the smaller the corresponding gain weight.

3. The multi-view audio generation method according to claim 1, wherein Each perspective corresponds to a camera unit for obtaining video data, and one sound collection unit is respectively arranged on both sides of the camera unit; The first audio data includes left-channel audio data and right-channel audio data; When obtaining the left-channel audio data corresponding to the video data from the first perspective, at least two sound collection units related to the first perspective include: the sound collection unit on the left side of the camera unit corresponding to the first perspective and the sound collection units on the left side of the camera units corresponding to at least one left-side perspective, and at least one of the left-side perspectives is a perspective arranged on the left side of the first perspective; When obtaining the right-channel audio data corresponding to the video data from the first perspective, at least two sound collection units related to the first perspective include: the sound collection unit on the right side of the camera unit corresponding to the first perspective and the sound collection units on the right side of the camera units corresponding to at least one right-side perspective, and at least one of the right-side perspectives is a perspective arranged on the right side of the first perspective.

4. The multi-view audio generation method according to claim 1, characterized in that, The method further includes: Obtaining a perspective switching request for the video data, where the perspective switching request is a playback request for the video data to switch from the first perspective to a second perspective, and the second perspective is a perspective adjacent to the first perspective; Obtaining audio data corresponding to the video data after perspective switching according to the first audio data and second audio data corresponding to the video data from the second perspective.

5. The multi-view audio generation method according to claim 4, wherein, Obtaining audio data corresponding to the video data after perspective switching according to the first audio data and second audio data corresponding to the video data from the second perspective includes: Obtaining the gain weight corresponding to the first audio data and the gain weight corresponding to the second audio data according to the switching progress information indicated by the perspective switching request; Obtaining audio data corresponding to the video data after perspective switching according to the first audio data, the second audio data, and the gain weights corresponding to the first audio data and the second audio data respectively.

6. The multi-viewpoint audio generation method according to claim 5, characterized in that The switching progress information includes the sliding position input by the user; The gain weight corresponding to the first audio data is related to the distance between the sliding position and the sliding start point, and the sliding start point corresponds to the first perspective; The gain weight corresponding to the second audio data is related to the distance between the sliding position and the sliding end point, and the sliding end point corresponds to the second perspective.

7. A multi-view audio generation device, characterized in that, Including: A first acquisition module, configured to acquire a first playback request for video data, where the first playback request is a playback request for the video data from a first perspective; A second acquisition module, configured to obtain first audio data corresponding to the video data from the first perspective according to audio signals collected by at least two sound collection units related to the first perspective and gain weights corresponding to each of the sound collection units.

8. A multi-view audio generation device, comprising: A transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor; characterized in that the processor is configured to read the program in the memory to implement the multi-perspective audio generation method according to any one of claims 1 to 6.

9. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-perspective audio generation method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by the processor, it implements the multi-perspective audio generation method according to any one of claims 1 to 6.