Video playing method, device and equipment
Through the combination of binocular cameras and microphone arrays, efficient stereo image and sound field reproduction is achieved, which solves the problem of high algorithm complexity in existing technologies and enhances the immersive experience of naked-eye 3D display.
Patent Information
- Application Number
- CN202510966798.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-24
AI Technical Summary
When generating immersive 3D images and sound fields, existing technologies have complex algorithm processing and poor results, making it difficult to achieve efficient naked-eye 3D display and spatial sound field reproduction.
A binocular camera is used to capture stereo images and display them through a naked-eye 3D display. A microphone array is used to record the direction of the sound source and play it back through an external speaker. Mixed encoding of images and audio is performed to ensure the accuracy of the sound source direction and the three-dimensional display.
It simplifies the complexity of multi-camera acquisition, improves the stereoscopic visual effect of naked-eye 3D display and the immersive feeling of spatial sound field, and avoids the limitation of the single position of the display's built-in audio.
Smart Images

Figure CN120835134A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital image processing, in particular to a video playing method, device and equipment. BACKGROUND
[0002] In order to enhance the immersive experience of live broadcast, video conference and other scenes, emerging display methods represented by naked-eye 3D display have become popular. At present, in order to obtain immersive audio and video experience, a method of generating new perspective 3D images by using 4 or more cameras to collect images and generating corresponding spatial sound field by tracking the binaural orientation of the audience is generally adopted. However, the algorithm processing process is complex and the restoration effect is poor. SUMMARY
[0003] Therefore, it is necessary to provide a video playing method, device and equipment aiming at the above technical problems.
[0004] In a first aspect, the present application provides a video playing method. The method comprises:
[0005] obtaining two original images and multiple original audio data;
[0006] splicing the two original images based on the input format of a to-be-played device to obtain a single image, wherein the single image comprises scene information in the two original images;
[0007] obtaining the sound source direction of the multiple original audio data based on a beamforming algorithm, performing noise reduction processing on the multiple original audio data based on the sound source direction to obtain single audio data;
[0008] mixing and encoding the single image and the single audio data to obtain mixed and encoded data;
[0009] decoding and transmitting the mixed and encoded data to the to-be-played device, so that the to-be-played device plays the single audio according to the sound source direction and displays the single image.
[0010] In one of the embodiments, the splicing the two original images based on the input format of the to-be-played device to obtain a single image comprises:
[0011] obtaining the matching area in one of the original images and the other original image, and determining the disparity of the matching area based on the displacement value between the matching areas;
[0012] obtaining a disparity image based on the disparity of multiple matching areas;
[0013] obtaining the relationship between the disparity and the depth based on the parameters of the to-be-played device;
[0014] convert the parallax image into a depth image based on the relationship between the parallax and the depth;
[0015] determine corresponding pixels of each pixel point under different viewing angles based on the parallax image and the depth image, perform pixel conversion based on a target viewing angle, and obtain a single-channel image.
[0016] In one of the embodiments, after obtaining the sound source direction of the multi-channel original audio data based on the beamforming algorithm, the method further comprises:
[0017] obtain a sound-emitting object in the sound source direction in the single-channel image, and obtain a target direction based on a sound-emitting position of the sound-emitting object.
[0018] In one of the embodiments, before obtaining the sound source direction of the multi-channel original audio data based on the beamforming algorithm, the method further comprises:
[0019] obtain a sound-emitting object in the single-channel image, and determine a sound-emitting direction;
[0020] adjust a parameter of the beamforming algorithm based on the sound-emitting direction, focus a beam direction to a direction indicated by the sound-emitting direction, and obtain the sound source direction of the multi-channel original audio data.
[0021] In one of the embodiments, the method further comprises:
[0022] in a case where the angle of the sound source direction exceeds a preset angle and in a case where a type of a microphone array in the playback device is a planar array, re-obtain multi-channel original audio data, and again perform judgment of the sound source direction;
[0023] or, in a case where the type of the microphone array in the playback device is a stereo array, convert the angle of the sound source direction exceeding the preset angle into a mirror angle relative to a camera connecting line in the playback device.
[0024] In one of the embodiments, the decoding and transmission of the mixed encoded data to the to-be-playback device for playing the single-channel audio according to the sound source direction and displaying the single-channel image comprises:
[0025] decode the mixed encoded data to obtain a single-channel image and a single-channel audio;
[0026] transmit the single-channel image to a display of the to-be-playback device, the display being configured to obtain position information of a target object, generate a vertical stripe image with odd and even columns alternated based on the position information, perform an anamorphic transformation on the vertical stripe image, and display the image after the anamorphic transformation;
[0027] The single-channel audio is transmitted to a loudspeaker of the to-be-played device, and the loudspeaker is used to calculate audio signal strength and phase played by each loudspeaker according to a sound source direction and a positional relationship of the loudspeaker, and play the single-channel audio.
[0028] In a second aspect, the present application further provides a device for video playing, the device comprising:
[0029] An acquisition module is configured to acquire two original images and multi-channel original audio data.
[0030] A processing module is configured to splice the two original images based on an input format of a to-be-played device to obtain a single-channel image, wherein the single-channel image comprises scene information in the two original images.
[0031] A sound source direction of the multi-channel original audio data is obtained based on a beamforming algorithm, and the multi-channel original audio data is denoised based on the sound source direction to obtain single-channel audio data.
[0032] A mixing module is configured to mix and encode the single-channel image and the single-channel audio data to obtain mixed and encoded data.
[0033] A playing module is configured to decode and transmit the mixed and encoded data to the to-be-played device, so that the to-be-played device plays the single-channel audio based on the sound source direction and displays the single-channel image.
[0034] In a third aspect, the present application further provides a device for video playing, the device comprising:
[0035] An audio-video acquisition device comprises a binocular camera, a microphone array, a processor and an audio output terminal module, wherein the binocular camera is configured to acquire images, and the microphone array is configured to acquire audio.
[0036] A display is configured to display images.
[0037] A loudspeaker is configured to play audio.
[0038] In a fourth aspect, the present application further provides a computer device, which comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements steps of a video playing method when executing the computer program.
[0039] In a fifth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program implements steps of a video playing method when executed by a processor.
[0040] In a sixth aspect, the disclosure also provides a computer program product. The computer program product comprises a computer program which, when executed by a processor, implements the steps of the video playing method.
[0041] The video playing method described above has at least the following beneficial effects:
[0042] The embodiment provided by the disclosure avoids the high algorithm complexity of the previous multi-camera collection and new view generation by collecting stereoscopic images through a binocular camera and making interlaced projection development through a naked-eye 3D display. The direction of the target sound source is recorded through a microphone array, and is transmitted to the other end together with the target sound source. After rendering, the sound field with direction is created through the playback of the external multi-speaker, avoiding the single sound emission position of the display built-in sound.
[0043] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the disclosure, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0045] Figure 1 The application environment diagram of the video playing method in an embodiment;
[0046] Figure 2 The device for video playing in an embodiment;
[0047] Figure 3 The device for video playing in an embodiment;
[0048] Figure 4 The flowchart of the video playing method in an embodiment;
[0049] Figure 5 The schematic diagram of the preset angle in an embodiment;
[0050] Figure 6 The structural block diagram of the device for video playing in an embodiment;
[0051] Figure 7 The internal structure diagram of the computer device in an embodiment;
[0052] Figure 8 The internal structure diagram of a server in an embodiment. DETAILED DESCRIPTION
[0053] In order for the ordinary person in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings.
[0054] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims. The terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or apparatus including a series of elements includes not only those elements, but also other elements not explicitly listed, or further includes elements inherent to such a process, method, product or apparatus. Without more limitations, it does not exclude the presence of other same or equivalent elements in the process, method, product or apparatus comprising the elements. For example, if the first, second, etc. terms are used to represent names, they do not represent any particular order.
[0055] The embodiments of the present disclosure provide a video playing method, which can be applied to an application environment as shown in Figure 1 The video playing device includes an audio and video collection device, which integrates a binocular camera, a microphone array, a processor, an audio output end and the like.
[0056] The binocular camera includes two RGB (Red green blue, three primary colors) cameras arranged horizontally with a small interval, and the distance is the normal human eye pupil distance, which can simulate the synchronous collection of images in front of the human eyes with width and depth information. In an embodiment of the present disclosure, the binocular camera refers to two cameras working simultaneously to provide images, but it is not limited that only two cameras are integrated on the audio and video module. A plurality of cameras can also be integrated on the audio and video collection module, and then two cameras in the most suitable position are dynamically switched to form a binocular camera according to the position of the collected target in front.
[0057] The microphone array uses a plurality of microphones to form an array to pick up surrounding sounds containing spatial information, including target voice and noise.
[0058] The processor performs audio and video signal processing and encoding work, including splicing two camera-captured images of the same frame according to the input format requirements (such as side-by-side left and right) of the naked-eye 3D display, denoising, de-echoing, and sound source angle information extraction of the microphone signal, mixing and encoding the image, and sending the mixed and encoded image to the computer. In addition, the processor receives audio signals and sound source angle information from the computer, renders the audio signals using the tangent law or other virtual sound image generation technology, and outputs multiple signals.
[0059] The audio output end maps the multiple signals output by the processor to the number and position of the loudspeakers one by one, and outputs the mapped signals after power amplification.
[0060] The computer serves as a control and application center, runs corresponding application software such as video live broadcast and video conference, and controls the normal work of the audio and video acquisition module and the peripheral devices such as the naked-eye 3D display.
[0061] In one embodiment of the present disclosure, Figure 2 For the video playing device in one embodiment, the audio and video acquisition module and the stereo loudspeaker can be arranged outside the naked-eye 3D display, and the existing naked-eye 3D display of the user is used to establish a structural connection with the naked-eye 3D display in a bracket or clip manner. In this embodiment, the horizontal distance between the binocular cameras is 35 mm, and the microphone array is a linear four-microphone array with a distance of 40 mm. A 3.5 mm stereo interface is used to connect the audio and video acquisition module and the stereo loudspeaker inside the peripheral device to facilitate separate production, packaging, and transportation. A USB interface is used to connect the peripheral device and the computer to realize plug and play.
[0062] In one embodiment of the present disclosure, Figure 3 For the video playing device in one embodiment, the audio and video acquisition module and the stereo loudspeaker are packaged into the frame of the naked-eye 3D display to form a spatial audio and video acquisition and playback integrated device. This device is suitable for users who have not purchased a naked-eye 3D display in advance. In this embodiment, the frame is a three-sided surrounding frame with a regular appearance. A locally protruding frame design can also be used, that is, the positions of the audio and video acquisition module and the stereo loudspeaker are protruding, and the frame design is not limited.
[0063] In some embodiments of the present disclosure, as Figure 4 shown, a video playing method is provided. The method is described by taking the processor in Figure 1 as an example. In one embodiment, the method can include the following steps:
[0064] S402: acquire two original images and multi-channel original audio data.
[0065] The two cameras with a distance close to the human binocular pupil distance form a binocular camera, and the two cameras work synchronously and simultaneously capture the same scene or object. Because the positions of the two cameras are different, just like the two eyes of a person observing things from different angles, the two original images obtained by shooting have left and right parallax. This left and right parallax is a key factor for producing stereoscopic vision, simulating the human eye to perceive the depth and spatial position of an object by observing the difference in imaging of the object on the left and right retinas, and the images with left and right parallax collected by the binocular camera provide basic data for subsequent applications requiring stereoscopic vision, such as naked-eye 3D display.
[0066] The microphone array can be composed of multiple microphones arranged in a certain manner, which can pick up sound signals containing target voice, and due to the cooperation of multiple microphones, the sound signals have spatial information. For example, the time and intensity of sound received by microphones at different positions will be different, and through the analysis of these differences, the spatial information such as the direction of the sound source can be determined, and the collected sound signals form multi-channel original audio signals.
[0067] S404: based on the input format of the to-be-played device, the two original images are spliced to obtain a single image, and the single image includes scene information in the two original images. The sound source direction of the multi-channel original audio data is obtained based on a beamforming algorithm, and the multi-channel original audio data is denoised based on the sound source direction to obtain single-channel audio data.
[0068] Different playback devices usually have specific requirements for the format of the image when receiving image data, such as resolution, color mode, encoding method, and arrangement form. Based on the input format of the to-be-played device, the two original images are spliced to ensure that the image can be normally played and displayed. The two original images synchronously captured by the binocular camera have left and right parallax because the positions of the cameras are different, so they capture the same scene but have different perspectives. The two images can be spliced and combined according to the format required by the to-be-played device to integrate them into one image. For example, in the side-by-side format, one of the images is placed on the left side and the other is placed on the right side to form a new wide image.
[0069] After the splicing operation, a complete single image is obtained, which contains all the scene information captured in the two original images. The arrangement of the image is adjusted to meet the input format requirements of the to-be-played device.
[0070] Due to the differences in the time and intensity of sound reaching different microphones, by analyzing these differences, the beamforming algorithm can calculate the direction of the sound source in space. For example, in an array composed of multiple microphones, when sound comes from a certain direction, the microphone closer to the sound source will receive the sound first, and the sound intensity is relatively large. Through processing and calculation of these signals, the beamforming algorithm can determine the specific direction of the sound source.
[0071] After obtaining the direction of the sound source, noise reduction processing is performed on the multiple original audio data based on the direction. The direction of the sound source can be used to distinguish target speech and noise. For sound signals from the direction of the sound source, reservation or enhancement is given, while for sound signals from other directions, they are considered as noise and are suppressed or attenuated. The interference of background noise on target speech can be reduced, and the quality of the audio can be improved. After noise reduction processing, the multiple audio data is merged or processed into single-channel audio data.
[0072] S406: Mix and encode the single-channel image and single-channel audio data to obtain mixed encoding data.
[0073] The single-channel image data and single-channel audio data that have been processed are integrated together using a specific encoding technology. The image and audio data are compressed and combined into a format that is convenient for storage, transmission and processing, and can contain both encoded image and audio information.
[0074] S408: Decode the mixed encoding data and transmit it to the to-be-played device, so that the to-be-played device plays the single-channel audio according to the direction of the sound source and displays the single-channel image.
[0075] After decoding, single-channel image data and single-channel audio data are obtained. The single-channel image data is transmitted to a display for playing, and the single-channel audio data is transmitted to a loudspeaker for playing.
[0076] In the above-mentioned video playing method, stereo images are collected by a binocular camera and are projected and developed by an naked-eye 3D display through interleaving, avoiding the high algorithm complexity of the previous generation of new view angle generation algorithm after multi-camera collection. The target sound source direction is recorded by a microphone array and is transmitted to the opposite end together with the target sound source, and after rendering, it is played back through a peripheral multi-loudspeaker to create a sound field with a sense of direction, avoiding the single sound position of the display built-in sound.
[0077] In some embodiments of the present disclosure, the input format of the to-be-played device is used to splice the two original images to obtain a single-channel image.
[0078] The matching area between the two original images is obtained based on the displacement value between the matching areas.
[0079] obtain a disparity image based on the disparities of the multiple matching regions;
[0080] obtain a relationship between the disparity and the depth based on parameters of the device to be played;
[0081] convert the disparity image into a depth image based on the relationship between the disparity and the depth;
[0082] determine corresponding pixels of each pixel point under different viewing angles based on the disparity image and the depth image, and perform pixel conversion based on a target viewing angle to obtain a single-channel image.
[0083] From the two original images, find a matching region in one of the images with the other image, which can be achieved by an image matching algorithm, such as a feature point matching algorithm, and find similar regions by comparing features in the images. After finding the matching region, calculate the displacement value between the matching regions, which represents the disparity of the matching region. The disparity is due to the difference in the positions of the same object in the two images caused by the different viewing angles of the two cameras.
[0084] Based on the disparity information of multiple matching regions, a disparity image is constructed, and each pixel value in the disparity image reflects the disparity size of the corresponding matching region. The disparity information of each part of the image is presented in the form of an image.
[0085] According to the related parameters of the device to be played, such as the camera's internal and external parameters, a mathematical relationship between the disparity and the depth is established, and the depth of the object is calculated from the disparity.
[0086] Based on the relationship between the disparity and the depth, the disparity value of each pixel in the disparity image is converted into a corresponding depth value, thereby obtaining a depth image. Each pixel value in the depth image represents the distance from the object at that position to the camera, so that the image contains the depth information of the scene.
[0087] Based on the disparity image and the depth image, determine the corresponding pixels of each pixel point under different viewing angles. Based on the disparity and depth information, find the pixels representing the same object point in different viewing angle images through a stereo matching algorithm. Then, convert the pixels of the two original images according to the target viewing angle. The conversion process can include geometric transformation operations such as affine transformation and projection transformation, which adjust the pixels to the appropriate position under the target viewing angle, and finally splice to obtain a single-channel image that meets the input format of the device to be played.
[0088] In some embodiments of the present disclosure, after obtaining the sound source direction of the multi-channel original audio data based on the beamforming algorithm, the method further comprises:
[0089] An object that makes sound in the single-channel image is obtained, and a target direction is obtained based on a sound-making position of the object.
[0090] After the multi-channel original audio data is processed using the beamforming algorithm, the approximate direction of the sound source in space, i.e., the sound source direction θ 11, has been obtained. Image information corresponding to the audio data is obtained, and an object that makes sound in the direction of the sound source is found in the image. Based on the specific sound-making position of the object in the image, such as human shape recognition, lip movement recognition, etc., a more accurate and specific direction, i.e., the target direction θ 12, is further determined.
[0091] In some embodiments of the present disclosure, before the sound source direction of the multi-channel original audio data is obtained based on the beamforming algorithm, the method further comprises:
[0092] An object that makes sound in the single-channel image is obtained, and a sound-making direction is determined;
[0093] Parameters of the beamforming algorithm are adjusted based on the sound-making direction, and the beam direction is focused on the direction indicated by the sound-making direction, so as to obtain the sound source direction of the multi-channel original audio data.
[0094] First, image information corresponding to the audio data is obtained, and an object that makes sound is found in the image. After the sound-making object is determined, the state of the object in the image is further analyzed, such as the orientation of a person, the part of an object that makes sound, etc., so as to determine the sound-making direction θ 21. For example, if a person in the image speaks to the left side, it can be preliminarily determined that the sound-making direction is the left side.
[0095] According to the sound-making direction θ 21, parameters of the beamforming algorithm are adjusted. The parameters can include weight settings of a microphone array, algorithm parameters of signal processing, etc. By adjusting the parameters, the beam direction of the beamforming algorithm is focused on the direction indicated by the sound-making direction θ 21. For example, if the sound-making direction is the left side, the parameters are adjusted to make the microphone array more focused on receiving sound signals in the left direction.
[0096] After the parameters of the beamforming algorithm are adjusted and the beam direction is focused on the sound-making direction θ 21, the algorithm is used to process the multi-channel original audio data, so as to obtain the sound source direction θ 22 of the multi-channel original audio data.
[0097] In some embodiments of the present disclosure, the method further comprises:
[0098] In the case where the angle of the sound source direction exceeds a preset angle, and in the case where the type of the microphone array in the playback device is a planar array, the multi-channel original audio data is re-obtained, and the judgment of the sound source direction is performed again;
[0099] Or, in the case of the microphone array type in the playback device being a stereo array, the angle of the sound source direction exceeding the preset angle is converted into a mirror angle relative to the camera line in the playback device.
[0100] Figure 5 For a schematic diagram of the preset angle in an embodiment, the preset angle θ can take a value of 0-180°, with the right side of the camera being 0° counterclockwise. When the angle of the sound source direction exceeds the preset angle, and the microphone array type in the playback device is a planar array, the multi-channel original audio data can be reacquired.
[0101] When the angle of the sound source direction exceeds the preset angle, and the microphone array type in the playback device is a stereo array, the angle of the sound source direction exceeding the preset angle can be converted into a mirror angle relative to the camera line in the playback device.
[0102] Through the above different processing methods, for different microphone array types, the abnormal sound source direction angle is reasonably disposed to ensure the accuracy of the sound source direction judgment, improve the effect and reliability of audio processing, and adapt to different application scenario requirements.
[0103] In some embodiments of the present disclosure, the mixed encoded data is decoded and transmitted to a to-be-played device, for the to-be-played device to play the single-channel audio according to the sound source direction and display the single-channel image.
[0104] The mixed encoded data is decoded to obtain a single-channel image and a single-channel audio.
[0105] The single-channel image is transmitted to a display of the to-be-played device, and the display is used to acquire position information of a target object, generate a vertical stripe image with odd and even phases based on the position information, perform a skew cut transformation on the vertical stripe image, and display the image after the skew cut transformation.
[0106] The single-channel audio is transmitted to a loudspeaker of the to-be-played device, and the loudspeaker is used to calculate the audio signal intensity and phase played by each loudspeaker according to the positional relationship between the sound source direction and the loudspeaker, and play the single-channel audio.
[0107] The mixed encoded data is decoded to obtain a single-channel image and a single-channel audio.
[0108] The single-channel image data obtained by decoding is transmitted to a display of a to-be-played device, and the display obtains position information of the target object. Based on the position information, the single-channel image is processed to generate vertical stripe images in odd and even columns, so as to prepare for a stereoscopic display effect. The image is divided into odd and even columns to form the vertical stripe images, and subsequent processing is used to make the left and right eyes see different image information, thereby generating a stereoscopic effect. The generated vertical stripe images are subjected to anamorphic transformation. The anamorphic transformation is a geometric transformation that causes the image to be deformed in a certain direction. The image after the anamorphic transformation is displayed by the display. Since the left and right eyes of the viewer see different anamorphic stripe images, the brain fuses the two images according to the binocular parallax principle, thereby enabling the viewer to perceive a stereoscopic image with depth.
[0109] The single-channel audio data obtained by decoding is transmitted to a loudspeaker of a to-be-played device. The loudspeaker calculates the audio signal intensity and phase of each loudspeaker according to the positional relationship between the sound source direction and the loudspeaker itself, and different loudspeakers at different positions play audio signals of different intensities and phases to simulate the effect of sound coming from different directions and enhance the spatial sense of the audio. The loudspeaker plays the single-channel audio according to the calculated audio signal intensity and phase, so that the listener can hear the sound with spatial sense, thereby improving the immersion and reality of the audio.
[0110] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0111] Based on the same inventive concept, the embodiments of the present disclosure also provide a video player for implementing the above-mentioned method for video playing. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in the following video player embodiment can be referred to the limitations of the video playing method in the above text, which will not be repeated here.
[0112] The apparatus can include a system (including a distributed system), software (application), module, component, server, client, etc. using the method described in the embodiments of the present specification in combination with necessary implementation hardware. Based on the same innovative concept, the apparatus in one or more embodiments provided by the embodiments of the present disclosure is described as follows. Since the implementation scheme of the apparatus to solve the problem is similar to the method, the implementation of the specific apparatus of the embodiments of the present specification can be referred to the implementation of the foregoing method, and the repeated parts will not be described herein. The term "unit" or "module" used below can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware or a combination of software and hardware is also possible and conceived.
[0113] In one embodiment, as shown in Figure 6 An apparatus 600 for video playing is provided, which can be the foregoing server, or a module, component, device, unit, etc. integrated in the server. The apparatus 600 can include:
[0114] The acquisition module 602 is configured to acquire two original images and multi-channel original audio data.
[0115] The processing module 604 is configured to splice the two original images based on an input format of a to-be-played device to obtain a single-channel image, the single-channel image including scene information in the two original images.
[0116] The processing module 604 is configured to obtain a sound source direction of the multi-channel original audio data based on the beamforming algorithm, and perform noise reduction processing on the multi-channel original audio data based on the sound source direction to obtain single-channel audio data.
[0117] The mixing module 606 is configured to perform hybrid coding on the single-channel image and the single-channel audio data to obtain hybrid coding data.
[0118] The playing module 608 is configured to decode the hybrid coding data and transmit the hybrid coding data to the to-be-played device, so that the to-be-played device plays the single-channel audio according to the sound source direction and displays the single-channel image.
[0119] As to the apparatus in the foregoing embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described herein.
[0120] The foregoing various modules in the video playing apparatus can be all or partially implemented by software, hardware, and a combination thereof. The foregoing various modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to the foregoing various modules.
[0121] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store video data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a video playback method is implemented.
[0122] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, memory, a communication interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal via wired or wireless communication. The wireless communication can be achieved via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements the video playback method. The display screen of the computer device can be a liquid crystal display or an electronic ink display. The input device of the computer device can be a touch layer covering the display screen, or keys, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse.
[0123] Those skilled in the art will understand that Figure 7 、 Figure 8 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present disclosure, and does not constitute a limitation on the computer device to which the solution of the present disclosure is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0124] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the present disclosure is implemented.
[0125] In an embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.
[0126] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. Any reference to memory, database or other medium used in the embodiments provided by the present disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided by the present disclosure can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided by the present disclosure can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0127] The technical features of the above embodiments can be combined in any manner. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present disclosure.
[0128] The above-described embodiments are merely illustrative of several embodiments of the present disclosure, which are described in a relatively specific and detailed manner, but should not be construed as limiting the scope of the patent of the present disclosure. It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present disclosure, and these all belong to the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the appended claims.
Claims
1. A method of video playback, characterized by, The method comprises: acquiring two original images and multi-channel original audio data; splicing the two original images based on an input format of a to-be-played device to obtain a single-channel image, the single-channel image comprising scene information in the two original images; obtaining a sound source direction of the multi-channel original audio data based on a beamforming algorithm, and performing noise reduction processing on the multi-channel original audio data based on the sound source direction to obtain single-channel audio data; mixing and encoding the single-channel image and the single-channel audio data to obtain mixed encoding data; decoding the mixed encoding data and transmitting the mixed encoding data to the to-be-played device, so that the to-be-played device plays the single-channel audio according to the sound source direction and displays the single-channel image.
2. The method of claim 1, wherein, The splicing of the two original images based on the input format of the to-be-played device to obtain a single-channel image comprises: acquiring a matching region in one of the original images and the other original image, determining a parallax of the matching region based on a displacement value between the matching regions; obtaining a parallax image based on the parallaxes of the matching regions; obtaining a relationship between the parallax and depth based on parameters of the to-be-played device; converting the parallax image into a depth image based on the relationship between the parallax and depth; determining corresponding pixels of each pixel point at different viewing angles based on the parallax image and the depth image, and performing pixel conversion based on a target viewing angle to obtain a single-channel image.
3. The method of claim 1, wherein, After obtaining the sound source direction of the multi-channel original audio data based on the beamforming algorithm, the method further comprises: acquiring a sound-emitting object in the sound source direction in the single-channel image, and obtaining a target direction based on a sound-emitting position of the sound-emitting object.
4. The method of claim 1, wherein, Before obtaining the sound source direction of the multi-channel original audio data based on the beamforming algorithm, the method further comprises: acquiring a sound-emitting object in the single-channel image, and determining a sound-emitting direction; adjusting parameters of the beamforming algorithm based on the sound-emitting direction, focusing a beam direction to a direction indicated by the sound-emitting direction, and obtaining the sound source direction of the multi-channel original audio data.
5. The method according to any of claims 3-4, characterized in that, The method further comprises: in a case where an angle of the sound source direction exceeds a preset angle and a case where a microphone array type in the to-be-played device is a planar array, re-acquiring multi-channel original audio data, and re-determining the sound source direction; or, in a case where the microphone array type in the to-be-played device is a stereo array, converting the angle of the sound source direction exceeding the preset angle into a mirror angle relative to a camera line in the to-be-played device.
6. The method of claim 1, wherein, The decoding of the mixed encoding data and the transmission of the mixed encoding data to the to-be-played device, so that the to-be-played device plays the single-channel audio according to the sound source direction and displays the single-channel image, comprises: decoding the mixed encoding data to obtain a single-channel image and single-channel audio; transmitting the single-channel image to a display of the to-be-played device, the display being configured to acquire position information of a target object, generate a vertical stripe image with odd and even columns alternated based on the position information, perform an anamorphic transformation on the vertical stripe image, and display the image after the anamorphic transformation. The single-channel audio is transmitted to a speaker of the to-be-played device, and the speaker is used to calculate audio signal intensity and phase played by each speaker according to a sound source direction and a positional relationship of the speaker, and play the single-channel audio.
7. An apparatus for video playback, the apparatus comprising: The device comprises: An acquisition module is configured to acquire two original images and multi-channel original audio data. A processing module is configured to splice the two original images based on an input format of a to-be-played device to obtain a single-channel image, wherein the single-channel image comprises scene information in the two original images. A sound source direction of the multi-channel original audio data is obtained based on a beamforming algorithm, and the multi-channel original audio data is denoised based on the sound source direction to obtain single-channel audio data. A mixing module is configured to mix and encode the single-channel image and the single-channel audio data to obtain mixed and encoded data. A playing module is configured to decode and transmit the mixed and encoded data to the to-be-played device, so that the to-be-played device plays the single-channel audio according to the sound source direction and displays the single-channel image.
8. An apparatus for video playback, the apparatus comprising: The device comprises: An audio and video acquisition device, the audio and video acquisition device comprising a binocular camera, a microphone array, a processor, and an audio output end module; the binocular camera is configured to acquire images, the microphone array is configured to acquire audio, and the audio output end module is configured to output the audio to a speaker; A display, the display being configured to display images; A speaker, the speaker being configured to play audio; A processor, the processor being configured to perform audio and video signal processing and encoding operations, and the processor being configured to implement the steps of the method according to any one of claims 1 to 6 when performing the audio and video signal processing and encoding operations. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor implements the steps of the method according to any one of claims 1 to 6 when executing the computer program.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-channel video hybrid coding method and device
CN102801979A
Video camera
CN106027919A
Binocular 720-degree panoramic acquisition system
CN106993177A
Binocular 720-degree panoramic acquisition system
WO2018068481A1
Video data processing method and system
WO2024061295A1