Video playing method, device, medium, equipment and vehicle for multi-sound field
By generating the correspondence between the perspective and the audio equipment in the vehicle, and controlling different audio equipment to play audio from different perspectives, the problem that the vehicle audio system cannot adapt to the viewing of movies from multiple perspectives is solved, and the auditory experience is improved.
Patent Information
- Application Number
- CN202210303546.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-03-24
AI Technical Summary
The existing vehicle audio system cannot adapt to multi-view movie viewing, resulting in poor auditory experience.
By generating the correspondence between the viewing angle and the audio equipment at different positions in the vehicle, different audio equipment are controlled to play audio from different viewing angles separately to create a multi-sound field effect.
It improves the auditory experience of passengers watching movies, presents a sense of space from different perspectives through the differences in vision and hearing, and provides a more on-site auditory experience.
Smart Images

Figure CN115442733B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of automotive technologies, and in particular, to a multi-acoustic field video playback method, apparatus, medium, device, and vehicle. Background Art
[0002] Multi-viewpoint viewing refers to providing video streams of multiple viewpoints to a user, and playing the video streams of different viewpoints through different screens. Utilizing the screens installed in a vehicle for multi-viewpoint viewing can meet different subjective viewing needs of passengers. However, currently, the audio systems in vehicles can often only create a fixed acoustic field effect and cannot adapt to different viewpoints, thus affecting the user's auditory experience. Summary of the Invention
[0003] To solve the above technical problems, the present disclosure provides a multi-acoustic field video playback method, apparatus, medium, device, and vehicle to create different acoustic field effects and improve the auditory experience of passengers during viewing.
[0004] The present disclosure provides a multi-acoustic field video playback method, including:
[0005] Obtaining video streams of multiple viewpoints; wherein, the video streams include audio;
[0006] Generating a first correspondence between the viewpoints and audio devices at different positions in the vehicle;
[0007] During the process of playing the video streams, controlling the different audio devices to play the audio of the video streams of different viewpoints respectively according to the first correspondence.
[0008] In some embodiments, the video streams further include pictures, and the method further includes:
[0009] When there are at least two display screens installed in the vehicle, generating a second correspondence between the viewpoints and the display screens;
[0010] During the process of playing the video streams, controlling the different display screens to play the pictures of the video streams of different viewpoints respectively according to the second correspondence.
[0011] In some embodiments, the generating the second correspondence between the viewpoints and the display screens includes:
[0012] Identifying the face positions of passengers in the vehicle;
[0013] Determining a first distance between each display screen and the face positions;
[0014] Obtaining a preset priority of the viewpoints;
[0015] Sort the perspectives in descending order of the said priority, and sort the display screens in ascending order of the said first distance;
[0016] Generate a second correspondence between the perspectives and the display screens with the same sorting order.
[0017] In some embodiments, the identifying the face positions of the passengers in the vehicle includes:
[0018] Obtain the target images of the passengers in the vehicle;
[0019] If multiple faces are detected in the target images, determine the proportion of the face area of each face in the target images;
[0020] Determine the face with the largest proportion of the face area as the target face;
[0021] Identify the face positions corresponding to the target face.
[0022] In some embodiments, the generating the second correspondence between the perspective and the display screen includes:
[0023] When there is a perspective with a recommended identifier, determine the perspective with the recommended identifier as the recommended perspective;
[0024] Determine the display screen corresponding to the shortest first distance as the main display screen, and / or select at least one main display screen from the display screens;
[0025] Generate the second correspondence between the recommended perspective and the main display screen.
[0026] In some embodiments, the video stream further includes pictures, and the method further includes:
[0027] When there is one display screen in the vehicle, obtain the timeline information of each video stream;
[0028] During the playback of the video stream, control the display screen to play the pictures of the video streams of different perspectives in the order of the timeline information.
[0029] In some embodiments, the controlling different audio devices to play the audio of the video streams of different perspectives according to the first correspondence includes:
[0030] In response to the switching of the perspective, obtain the target perspective after the switching;
[0031] Switch the first audio device to a second audio device having the first corresponding relationship with the target perspective; wherein, the first audio device is the audio device before the perspective switch.
[0032] Control the second audio device to play the audio of the video stream in the target perspective.
[0033] In some embodiments, the perspective switch includes: the current perspective corresponding to the main display screen is switched, and the current perspective includes the recommended perspective.
[0034] In some embodiments, generating the first corresponding relationship between the perspective and the audio devices at different positions in the vehicle includes:
[0035] According to the installation positions of the audio devices in the vehicle, allocate at least one audio device to each perspective, so that the relative direction of the at least one audio device in the vehicle matches the perspective.
[0036] Generate the first corresponding relationship between the perspective and the at least one audio device currently allocated to the perspective.
[0037] The present disclosure provides a multi-soundfield video playback device, including:
[0038] A video stream acquisition module, configured to acquire video streams of multiple perspectives; wherein, the video stream includes audio.
[0039] A first relationship generation module, configured to generate a first corresponding relationship between the perspective and the audio devices at different positions in the vehicle.
[0040] An audio playback module, configured to, during the playback of the video stream, control different audio devices to play the audio of the video streams of different perspectives respectively according to the first corresponding relationship.
[0041] The present disclosure further provides a computer-readable storage medium, which stores programs or instructions, and the programs or instructions cause a computer to execute the steps of any of the above methods.
[0042] The present disclosure further provides an electronic device, including:
[0043] One or more processors;
[0044] A memory, configured to store one or more programs or instructions;
[0045] The processor is configured to execute the steps of any of the above methods by calling the programs or instructions stored in the memory.
[0046] The present disclosure also provides a vehicle, which includes the above multi-soundfield video playback device.
[0047] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:
[0048] For the technical solutions provided by the embodiments of the present disclosure, first, video streams from multiple perspectives are acquired; then, a first correspondence between the perspectives and audio devices at different positions inside the vehicle is generated; during the playback of the video streams, according to the first correspondence, different audio devices can be controlled to play the audio of the video streams adapted to different perspectives respectively, so as to create different soundfield effects and greatly improve the auditory experience of passengers watching movies. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0051] Figure 1 It is a flowchart of a multi-soundfield video playback method provided by an embodiment of the present disclosure;
[0052] Figure 2 It is a schematic diagram of a scene of a perspective provided by an embodiment of the present disclosure;
[0053] Figure 3 It is a schematic diagram of audio devices inside a vehicle provided by an embodiment of the present disclosure;
[0054] Figure 4 It is a block diagram of the structure of a multi-soundfield video playback device provided by an embodiment of the present disclosure;
[0055] Figure 5 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] In order to be able to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0057] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.
[0058] Multi-perspective viewing is a newly added multi-screen linkage application scenario in vehicles. During actual use, the audio system in a vehicle often can only create a fixed sound field effect and cannot adapt to different perspectives, affecting the user's auditory experience. Based on this, embodiments of the present disclosure provide a multi-sound field video playback method, apparatus, medium, device, and vehicle. This technology is applicable to scenarios in vehicles where multi-perspective viewing is possible, such as variety shows, sports games, etc.; this technology plays the audio of the video streams of different said perspectives through different audio devices respectively, presents the spatial sense of the audio under different perspectives, and creates a sound field effect adapted to the perspectives, thereby improving the auditory experience of passengers.
[0059] Figure 1 It is a flowchart of a multi-sound field video playback method provided by an embodiment of the present disclosure. This method can be executed by a multi-sound field video playback device, and the multi-sound field video playback device can be implemented in a software and / or hardware manner. As Figure 1 shown, the method includes the following steps:
[0060] Step S102, obtain video streams of multiple perspectives; wherein, the video streams include audio.
[0061] In this embodiment, multiple acquisition devices arranged at different camera positions can acquire video streams from different perspectives at the scene of the program to obtain video streams of the program from at least one perspective. Taking a sports game program as an example, video streams can be acquired from perspectives such as the playing field, the audience seats, and the referee stand. The above acquisition devices at different camera positions include cameras for shooting pictures and sound collection devices for recording audio, and the cameras and sound collection devices can be of an integrated structure or independent structures respectively. When the cameras and sound collection devices are independent structures respectively, the shooting direction of the camera and the sound collection direction of the sound collection device at the same camera position are the same, so as to ensure the consistency of the spatial sense of the picture and the audio. The obtained video streams include audio and pictures, which can be used as the source data for the multi-sound field video playback solution provided by this embodiment.
[0062] In this embodiment, during the playback of the video stream, the audio and pictures of the video stream can be played through the audio devices and the display screen in the vehicle respectively. Based on this, different perspectives of auditory experiences are provided by the audio devices, and different perspectives of visual experiences are provided by the display screen, and the spatial sense of different perspectives is presented through the differences between hearing and vision.
[0063] Step S104, generate a first correspondence between the perspectives and the audio devices at different positions in the vehicle.
[0064] Among them, various methods such as random matching and user setting can be adopted to generate the first correspondence between the viewing angle and the audio devices at different positions inside the vehicle. For the convenience of understanding, a possible embodiment of generating the first correspondence is provided here. Refer to the following content.
[0065] According to the installation positions of the audio devices inside the vehicle, at least one audio device is assigned to each viewing angle, so that the relative directions of at least one audio device inside the vehicle match the viewing angle; a first correspondence is generated between the viewing angle and the at least one audio device assigned to the current viewing angle.
[0066] Relative spatial orientations can be formed between different viewing angles of the same program scene; for example Figure 2 as shown, the stage viewing angle, the tutor viewing angle, and the audience viewing angle form a relative spatial orientation arranged from top to bottom. Taking Figure 3 as an example, there are multiple audio devices inside the vehicle, corresponding to different installation positions of the vehicle.
[0067] In practical applications, among multiple viewing angles, there is usually a recommended viewing angle for watching movies, and this recommended viewing angle is, for example, a viewing angle with information such as a preset recommended identifier or the highest priority. When assigning audio devices to each viewing angle, the audio devices at the key installation positions that can present the sense of sound stereo can be assigned to the recommended viewing angle. Referring to Figure 3 , the key positions may include the audio devices at the four installation positions of the front, rear, left, and right inside the vehicle, that is Figure 3 the audio devices numbered 2, 4, 5, and 7 respectively in
[0068] Generally, the stage viewing angle is the recommended viewing angle, and the above-mentioned audio devices numbered 2, 4, 5, and 7 are assigned to the stage viewing angle. Figure 2 For other viewing angles except the recommended viewing angle, at least one audio device that matches the relative spatial orientation can be assigned to the current viewing angle according to the relative spatial orientation corresponding to the current viewing angle. For example, for the audience viewing angle in Figure 3 , its corresponding relative spatial orientation is the lower orientation. Therefore, at least one audio device (that is, at least one of the audio devices numbered 6, 7, and 8) that is also in the lower orientation in
[0069] Referring to the above method, at least one audio device that matches the relative spatial orientation is assigned to each viewing angle, and a first correspondence is generated between the viewing angle and the at least one audio device assigned to the current viewing angle. Through the first correspondence generated by this embodiment, the spatial consistency between the viewing angle and the audio device can be improved. Therefore, the use of the audio device with the first correspondence can improve the restoration degree of the sound field effect of the corresponding viewing angle.
[0070] Step S106, during the process of playing the video stream, according to the first correspondence relationship, control different audio devices to play the audio of the video stream from different perspectives respectively.
[0071] Among them, playing the video stream can be playing multiple perspective video streams simultaneously; during the process of playing multiple perspective video streams simultaneously, it can be controlled that only the audio devices having the first correspondence relationship with the recommended perspective play the audio of the video stream from the recommended perspective, or it can be controlled that the audio devices having the first correspondence relationship with two or more perspectives play the audio of the video streams from the corresponding perspectives respectively. Exemplarily, taking Figure 2 as an example, during the process of playing the video streams corresponding to the stage perspective and the auditorium perspective simultaneously, it can be controlled that only the audio devices numbered 2, 4, 5, and 7 having the first correspondence relationship with the stage perspective (i.e., the recommended perspective) play the audio of the video stream from the stage perspective. Or, it can also be controlled that the audio devices having the first correspondence relationship with the stage perspective play the audio of the video stream from the stage perspective, and, control the audio devices numbered 6 and 8 corresponding to the auditorium perspective to play the audio of the video stream from the auditorium perspective. In addition, in order to enhance the sound field effect of the audio corresponding to the recommended perspective, in this embodiment, the volume of the audio device corresponding to the recommended perspective can be set higher than that of other audio devices.
[0072] In this embodiment, different audio devices play the audio of the video stream from different perspectives respectively, restoring the audio collected from different perspectives. For example, the enthusiastic atmosphere on the stage, the screams in the auditorium, etc. can all be broadcast through the audio devices corresponding to different perspectives, enabling passengers to obtain a more immersive auditory experience.
[0073] In this embodiment, it can be combined with the embodiment of playing the video stream picture on the display screen to describe the multi - sound - field video playing method in more detail.
[0074] In one embodiment, the video streams from multiple perspectives can exist in the form of segments in a complete combined video stream. The video streams from each perspective existing in the form of segments in the combined video stream are usually applicable to the scenario where the vehicle has a relatively low configuration and only one display screen. By playing the combined video stream through this display screen, it is equivalent to playing the video streams from multiple perspectives.
[0075] Correspondingly, the method provided in this embodiment includes: when there is one display screen set in the vehicle, obtain the time - axis information of each video stream; during the process of playing the video stream, control the display screen to play the pictures of the video streams from different perspectives in the order of the time - axis information.
[0076] Among them, the video streams from each perspective obtained can be spliced to obtain a combined video stream, and timeline information of each video stream on the timeline of the combined video stream is generated; the timeline information can represent the start time and end time of the video stream. During the process of playing the video stream, the display screen is controlled to play the picture of the combined video stream, and as the playing progress of the combined video stream, the pictures of the video streams from different perspectives are played in the order of the timeline information.
[0077] In another embodiment, the video streams from multiple perspectives can exist in the form of independent individuals, and each video stream is associated with one perspective; this embodiment is applicable to the scenario of multiple display screens, and the pictures of the video streams from different perspectives are played on different display screens respectively. Correspondingly, the method provided in this embodiment includes the following steps (1) and (2).
[0078] (1) When there are at least two display screens in the vehicle, a second correspondence between the perspective and the display screen is generated.
[0079] During implementation, the face positions of the passengers in the vehicle can be recognized first. Specifically: obtain the target image of the passengers in the vehicle, for example, by using a camera inside or outside the vehicle facing the seat position to capture the target image of the passengers in the vehicle. If only one face is detected in the target image, the face position of this face is determined through image recognition. If multiple faces are detected in the target image, the face position can be recognized by referring to the following methods.
[0080] Method 1: Determine the proportion of the face area of each face in the target image; determine the target face as the face corresponding to the largest proportion of the face area; recognize the face position corresponding to the target face.
[0081] A relatively large proportion of the face area of the face indicates that the passenger is closer to the camera, the face angle is closer to the frontal angle rather than the side angle, or the face is more complete and there is no occlusion. Based on this, using the target face corresponding to the largest proportion of the face area for image recognition can improve the recognition accuracy of the face position.
[0082] Method 2: Generate face selection information; obtain the selection result for the face selection information, and determine the target face among multiple faces according to the selection result; recognize the face position corresponding to the target face in the target image.
[0083] Specifically, the generated face selection information can be simultaneously displayed on each display screen, which is conducive to passengers selecting any convenient display screen to perform operations. The passengers feedback the selection result to the face selection information through the display screen. When the selection result feedback by any display screen is obtained, the target face is determined among multiple faces according to the selection result, and the display of the face selection information on each display screen is cancelled. Recognize the face position corresponding to the target face in the target image.
[0084] It should be noted that the above two methods are only examples for identifying the position of the human face and should not be construed as limitations.
[0085] Then, a first distance between each display screen and the human face position is determined. Since the positions of the display screens in the vehicle are fixed and known, after determining the human face position, the first distance between each display screen and the human face position can be further determined.
[0086] Next, obtain the preset priority of the perspective; sort the perspectives from high to low according to the priority, and sort the display screens from near to far according to the first distance; generate a second correspondence between the perspectives and the display screens with the same sorting order.
[0087] The level of priority can represent the primary and secondary relationship of the perspectives. The perspective with the highest priority is generally the recommended perspective, such as the stage perspective in variety shows or the stadium perspective in sports competition programs; the priority of the perspective can be set during video acquisition or post-editing. The proximity of the first distance can represent the suitability of the viewing distance of the display screen. The display screen with the closest first distance is generally the display screen closest to the passenger and with the best viewing distance, such as the display screen in front of the passenger. Based on this, the perspectives and the display screens are sorted respectively. According to the sorting order, the perspective with a higher priority is matched with the display screen with a closer first distance, thereby generating a second correspondence between the perspective and the display screen.
[0088] In a possible embodiment of the second correspondence, when there is a perspective with a recommended logo, the perspective with the recommended logo is determined as the recommended perspective; the display screen corresponding to the closest first distance is determined as the main display screen, and / or, at least one main display screen is selected from the display screens; generate a second correspondence between the recommended perspective and the main display screen.
[0089] In this embodiment, the perspective with the recommended logo is set as the perspective with the highest priority, and then a second correspondence between the recommended perspective and the display screen is generated according to the level of priority and the proximity of the first distance. Correspondingly, the display screen corresponding to the closest first distance is generally the closest display screen that the passenger can view, and this display screen can be determined as the main display screen. Or, there may be at least one passenger in the vehicle who wants to view the video stream of the recommended perspective. Based on this, at least one display screen can be selected from multiple display screens as the main display screen based on the user's selection operation. Then, a second correspondence between the recommended perspective and at least one main display screen is generated.
[0090] In addition, for other perspectives and other display screens, a second correspondence between the perspective and the display screen can be generated according to the level of priority and the proximity of the first distance.
[0091] In this embodiment, a second correspondence is established between the perspective with a higher priority and the display screen that is closer to the first distance, so that the display screen and its corresponding perspective can provide a more friendly and comfortable visual experience for the user.
[0092] In some other embodiments, a second correspondence between the perspective and the display screen can also be generated by means of random matching or user settings.
[0093] (2) During the process of playing the video stream, according to the second correspondence, control different display screens to play the pictures of the video stream from different perspectives respectively.
[0094] Considering that the number of perspectives and the number of display screens may not be the same, several examples of playing pictures are provided as follows.
[0095] In one example, when the number of perspectives is more than the number of display screens, the first N perspectives with higher priorities can be determined first, where N is the number of display screens and N≥2; then a second correspondence in which the first N perspectives correspond to the display screens one by one is generated, and control different display screens to play the pictures of the video stream from the above N perspectives respectively.
[0096] In another example, when the number of perspectives is less than the number of display screens, the first M display screens with the closest first distance can be determined first, where M is the number of perspectives; then a second correspondence in which the perspectives correspond to the first M display screens one by one is generated, and control the first M display screens to play the pictures of the video stream from different perspectives respectively.
[0097] In this example, for the idle display screens that are not used to play pictures, the idle display screens can be controlled to display interactive messages such as bullet screens and comments of the video stream.
[0098] Based on the above embodiments in which one display screen is provided in the vehicle and the display screen is controlled to play the pictures of the video stream from different perspectives, or, based on the above embodiments in which at least two display screens are provided in the vehicle and different display screens are controlled to play the pictures of the video stream from different perspectives respectively, the following provides a method for controlling different audio devices to play the audio of the video stream from different perspectives, as shown below.
[0099] In response to a perspective switch, obtain the target perspective after the switch.
[0100] Among them, the perspective switch includes: the current perspective corresponding to the main display screen switches, and the current perspective includes the recommended perspective. Among them, when one display screen is provided in the vehicle, the main display screen is this display screen; when at least two display screens are provided in the vehicle, the main display screen is the display screen determined by means of the first distance or user selection.
[0101] Specifically, when the main display screen initially plays the video stream, the current viewing angle is generally the recommended viewing angle; during the process of the main display screen playing the video stream of the recommended viewing angle, if the viewing angle is switched, the picture played on the main display screen switches from the recommended viewing angle to the first viewing angle, and the first viewing angle is determined as the target viewing angle after the switch. Thereafter, the first viewing angle is the current viewing angle corresponding to the main display screen. During the process of the main display screen playing the video stream of the first viewing angle, if the viewing angle is switched, the picture played on the main display screen switches from the first viewing angle to the second viewing angle, and the second viewing angle is determined as the target viewing angle after the switch.
[0102] The switching of the viewing angle may also include: the current viewing angle corresponding to any display screen is switched, and any display screen may be the main display screen or other display screens that play pictures.
[0103] Exemplarily, the multiple display screens include a co-pilot display screen and a rear-row display screen, where the co-pilot display screen is the main display screen. According to the second correspondence, control the co-pilot display screen to play the video stream of the third viewing angle, and control the rear-row display screen to play the video stream of the fourth viewing angle. During this process, if the current viewing angle corresponding to the rear-row display screen is switched, that is, from the fourth viewing angle to the fifth viewing angle, then the fifth viewing angle is determined as the target viewing angle after the switch.
[0104] After obtaining the target viewing angle after the switch according to the above embodiments, switch the first audio device to the second audio device that has the first correspondence with the target viewing angle; where the first audio device is the audio device before the viewing angle switch; control the second audio device to play the audio of the video stream at the target viewing angle.
[0105] In some embodiments, in the scenario of only controlling the audio device that has the first correspondence with the current viewing angle corresponding to the main display screen to play audio, only when the viewing angle of the main display screen is switched will the audio device be correspondingly switched. Specifically, in response to the switching of the current viewing angle corresponding to the main display screen (assumed to be the stage viewing angle), obtain the target viewing angle after the switch (assumed to be the audience viewing angle). Switch the first audio devices numbered 2, 4, 5, and 7 that have the first correspondence with the stage viewing angle to the second audio devices numbered 6 and 8 that have the first correspondence with the audience viewing angle; control the above-mentioned second audio devices to play the audio of the video stream at the audience viewing angle.
[0106] In some embodiments, in a scenario where audio devices corresponding to at least two perspectives play audio respectively, the display screen for presenting the images of the at least two perspectives can be used as the target display screen. It can be understood that there are also at least two target display screens, which correspond to the perspectives one by one. When the perspective of the target display screen is switched, the audio devices will be switched accordingly. In one example, the at least two perspectives are the stage perspective and the auditorium perspective, and the corresponding target display screens are the co-pilot display screen and the rear-row display screen. During the process of playing the video streams of the above two perspectives, in response to the switching of the auditorium perspective corresponding to the rear-row display screen, the target perspective after switching is obtained as the tutor perspective. The first audio devices numbered 6 and 8 respectively, which have a first correspondence with the auditorium perspective, are switched to the second audio devices numbered 4 and 5 respectively, which have a first correspondence with the tutor perspective; the above-mentioned second audio devices are controlled to play the audio of the video stream from the tutor perspective.
[0107] In the above embodiments, when the perspective is switched, the images played on the display screen and the audio played on the audio device are synchronously switched, enabling the user to jointly perceive the spatial difference brought by the perspective change visually and auditorily.
[0108] In summary, the multi-soundfield video playback method provided by the embodiments of the present disclosure first obtains video streams of multiple perspectives, and then generates a first correspondence between the perspectives and the audio devices at different positions in the vehicle; during the process of playing the video streams, according to the first correspondence, different audio devices are controlled to play the audio of the video streams of different perspectives respectively. This technical solution utilizes the first correspondence between the perspectives and the audio devices, enabling different audio devices to play audio adapted to different perspectives, thereby creating different soundfield effects and greatly improving the auditory experience of passengers watching movies.
[0109] At the same time, combined with the multi-perspective movie viewing provided by the display screen in the vehicle, the images and audio are both matched with the perspectives, and the spatial sense of different perspectives is better presented through different display screens and different audio devices, meeting the dual experience of users visually and auditorily and further improving the movie viewing experience of passengers.
[0110] Corresponding to the multi-soundfield video playback method provided by the embodiments of the present disclosure, the embodiments of the present disclosure also provide a multi-soundfield video playback device. Figure 4 As shown in the structural block diagram of the multi-soundfield video playback device provided by the embodiments of the present disclosure, Figure 4 the multi-soundfield video playback device includes:
[0111] A video stream acquisition module 402, configured to acquire video streams of multiple perspectives; wherein, the video stream includes audio;
[0112] The first relationship generation module 404 is configured to generate a first correspondence between perspectives and audio devices at different positions inside the vehicle;
[0113] The audio playback module 406 is configured to, during the playback of the video stream, control different audio devices to respectively play the audio of the video stream from different perspectives according to the first correspondence.
[0114] In some embodiments, the video stream further includes a picture, and the apparatus further includes a picture playback module (not shown in the figure), and the picture playback module is configured to:
[0115] When at least two display screens are provided inside the vehicle, generate a second correspondence between perspectives and the display screens; during the playback of the video stream, control different display screens to respectively play the pictures of the video stream from different perspectives according to the second correspondence.
[0116] In some embodiments, the picture playback module is specifically configured to: identify the face positions of passengers inside the vehicle; determine the first distances between each display screen and the face positions; obtain the preset priorities of the perspectives; sort the perspectives from high to low according to the priorities, and sort the display screens from near to far according to the first distances; generate a second correspondence between the perspectives and the display screens with the same sorting order.
[0117] In some embodiments, the picture playback module is specifically configured to: obtain the target image of the passengers inside the vehicle; if multiple faces are detected in the target image, determine the proportion of the face areas of each face in the target image; determine the face corresponding to the largest proportion of the face area as the target face; identify the face position corresponding to the target face.
[0118] In some embodiments, the picture playback module is specifically configured to: when there is a perspective with a recommended identifier, determine the perspective with the recommended identifier as the recommended perspective; determine the display screen corresponding to the nearest first distance as the main display screen, and / or select at least one main display screen from the display screens; generate a second correspondence between the recommended perspective and the main display screen.
[0119] In some embodiments, the video stream further includes a picture, and the picture playback module is further configured to: when there is one display screen inside the vehicle, obtain the time axis information of each video stream; during the playback of the video stream, control the display screen to play the pictures of the video stream from different perspectives in the order of the time axis information.
[0120] In some embodiments, the audio playback module 406 is specifically configured to: in response to a perspective switch, obtain the switched target perspective; switch the first audio device to a second audio device having a first correspondence with the target perspective; wherein the first audio device is the audio device before the perspective switch; control the second audio device to play the audio of the video stream from the target perspective.
[0121] In some embodiments, the switching of the viewing angle includes: the switching of the current viewing angle corresponding to the main display screen, and the current viewing angle includes the recommended viewing angle.
[0122] In some embodiments, the first relationship generation module 404 is specifically configured to: allocate at least one audio device to each viewing angle according to the installation position of the in-vehicle audio device, so that the relative direction of at least one audio device in the vehicle matches the viewing angle; generate a first correspondence between the viewing angle and at least one audio device assigned to the current viewing angle.
[0123] The multi-soundfield video playback device disclosed in the above embodiments can execute the multi-soundfield video playback method disclosed in the above embodiments, and has the same or corresponding beneficial effects. To avoid repetition, it will not be elaborated here.
[0124] The embodiments of the present disclosure further provide a vehicle, which includes the above-mentioned multi-soundfield video playback device.
[0125] The embodiments of the present disclosure further provide a computer-readable storage medium, which stores a program or instructions, and the program or instructions cause a computer to execute the steps of any of the above methods.
[0126] Obtain video streams of multiple viewing angles; wherein, the video stream includes audio; generate a first correspondence between the viewing angle and audio devices at different positions in the vehicle; during the playback of the video stream, control different audio devices to play the audio of the video stream of different viewing angles according to the first correspondence.
[0127] Optionally, when executed by a computer processor, the computer-executable instructions can also be used to execute the technical solutions of any of the above multi-soundfield video playback methods provided by the embodiments of the present disclosure, and achieve the corresponding beneficial effects.
[0128] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solutions of the embodiments of the present disclosure, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk, or an optical disc of a computer, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present disclosure.
[0129] An embodiment of the present disclosure further provides an electronic device, including: one or more processors; a memory for storing one or more programs or instructions; the processor is configured to execute the steps of any of the above methods by invoking the programs or instructions stored in the memory, so as to achieve corresponding beneficial effects.
[0130] Figure 5 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present disclosure. As Figure 5 shown, the electronic device includes one or more processors 501 and a memory 502.
[0131] The processor 501 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0132] The memory 502 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 501 may run the program instructions to implement the multi-soundfield video playback method of the embodiment of the present disclosure described above, and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage media.
[0133] In one example, the electronic device may further include: an input device 503 and an output device 504, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0134] In addition, the input device 503 may further include, for example, a keyboard, a mouse, and the like.
[0135] The output device 504 may output various information to the outside, including determined distance information, direction information, etc. The output device 504 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0136] Of course, for simplicity, Figure 5 only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.
[0137] It should be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0138] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video playback method for multiple sound fields, characterized in that Including: Obtain video streams from multiple perspectives; wherein, the video streams include audio. Generate a first correspondence between the perspectives and audio devices at different positions inside the vehicle. During the playback of the video streams, according to the first correspondence, control different ones of the audio devices to respectively play the audio of the video streams from different perspectives. The video streams further include images, and the method further includes: When at least two display screens are provided inside the vehicle, determine a first distance between each of the display screens and the passengers inside the vehicle. Determine a second correspondence between the perspectives and the display screens according to the first distance. During the playback of the video streams, according to the second correspondence, control different ones of the display screens to respectively play the images of the video streams from different perspectives.
2. The method according to claim 1, wherein Generating the second correspondence between the perspectives and the display screens includes: Identify the face positions of the passengers inside the vehicle. Determine a first distance between each of the display screens and the face positions. Obtain a preset priority of the perspective. Sort the perspectives in descending order according to the priority, and sort the display screens in ascending order according to the first distance. Generate a second correspondence between the perspectives and the display screens with the same sorting order.
3. The method according to claim 2, wherein The identifying the face positions of the passengers inside the vehicle includes: Obtain a target image of the passengers inside the vehicle. If multiple faces are detected in the target image, determine the proportion of the face area of each of the faces in the target image. Determine the face corresponding to the largest proportion of the face area as the target face. Identify the face position corresponding to the target face.
4. The method according to claim 2, wherein Generating the second correspondence between the perspectives and the display screens includes: When there is a perspective with a recommended label, determine the perspective with the recommended label as the recommended perspective. Determine the display screen corresponding to the nearest first distance as the main display screen, and / or select at least one of the main display screens from the display screens. Generate the second correspondence between the recommended perspective and the main display screen.
5. The method according to claim 1, characterized in that, The video streams further include images, and the method further includes: When one display screen is provided inside the vehicle, obtain the time axis information of each of the video streams. During the playback of the video streams, control the display screen to play the images of the video streams from different perspectives in the order of the time axis information.
6. The method according to claim 4, characterized in that, The controlling different ones of the audio devices to respectively play the audio of the video streams from different perspectives according to the first correspondence includes: In response to a switch of the perspective, obtain a target perspective after the switch. Switch a first audio device to a second audio device having the first correspondence with the target perspective; wherein, the first audio device is the audio device before the perspective switch. Control the second audio device to play the audio of the video stream in the target perspective.
7. The method according to claim 6, wherein The switch of the perspective includes: a switch of the current perspective corresponding to the main display screen, and the current perspective includes the recommended perspective.
8. The method according to claim 1, characterized in that The generating the first correspondence between the perspectives and the audio devices at different positions inside the vehicle includes: Allocate at least one audio device for each of the perspectives according to the installation position of the audio device in the vehicle, so that the relative direction of the at least one audio device in the vehicle matches the perspective; Generate the first correspondence between the perspective and the at least one audio device allocated to the current perspective.
9. A video playback device with multiple sound fields, characterized in that, Comprising: A video stream acquisition module, configured to acquire video streams of multiple perspectives; wherein, the video stream includes audio; A first relationship generation module, configured to generate a first correspondence between the perspective and audio devices at different positions in the vehicle; An audio playback module, configured to, during the playback of the video stream, control different ones of the audio devices to respectively play the audio of the video streams of different perspectives according to the first correspondence; The video stream further includes a picture, and the apparatus further includes: When there are at least two display screens provided in the vehicle, determine a first distance between each of the display screens and a passenger in the vehicle; Determine a second correspondence between the perspective and the display screen according to the first distance; During the playback of the video stream, control different ones of the display screens to respectively play the pictures of the video streams of different perspectives according to the second correspondence.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instruction, and the program or instruction causes a computer to execute the steps of the method according to any one of claims 1 to 8.
11. An electronic device, characterized in that, Comprising: One or more processors; A memory, configured to store one or more programs or instructions; The processor is configured to execute the steps of the method according to any one of claims 1 to 8 by calling the program or instruction stored in the memory.
12. A vehicle, characterized in that, The vehicle includes the multi-sound-field video playback device according to claim 9.
Citation Information
Patent Citations
In-vehicle information reproducing apparatus
CN1980484A