Video processing device, terminal device, and video processing method
By setting up multiple acquisition sub-modules to generate stereoscopic video in remote communication scenarios and adjusting the display screen of terminal devices, the problem of poor on-site immersion for remote users is solved, and a three-dimensional stereoscopic user experience is achieved.
Patent Information
- Application Number
- CN202410813377.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-21
- Publication Date
- 2025-12-23
AI Technical Summary
In remote communication, the scene seen by the remote user is a fixed location, resulting in a poor sense of immersion and a poor user experience.
Multiple acquisition sub-modules are set up within the scene to generate single-point stereoscopic video, and the display screen of the terminal device is adjusted by the target stereoscopic video set to achieve a three-dimensional stereoscopic effect.
It enhances the user's sense of immersion and improves the user experience.
Smart Images

Figure CN121193879A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiments of the present application relate to video processing technology. More particularly, it relates to a video processing device, a terminal device and a video processing method. BACKGROUND
[0002] With the development of remote communication such as video conference and distance education, display devices such as conference large screen and electronic whiteboard have also been widely applied.
[0003] Taking the conference large screen as an example, the audio and video module is the main interactive medium of the conference large screen. At present, the audio and video module usually captures a scene at a fixed position, and at this time the picture seen by the user at the remote end is also the scene at the fixed position, which makes the user's on-site immersion poor, thereby resulting in poor user experience. SUMMARY
[0004] The embodiments of the present application provide a video processing device, a terminal device and a video processing method, which can be used to solve the problem in the related art that the picture seen by the remote user is the scene at a fixed position, which makes the user's on-site immersion poor, thereby resulting in poor user experience.
[0005] In a first aspect, the embodiments of the present application provide a video processing device, which comprises:
[0006] a processor configured to acquire a picture of a scene where the video processing device is located, and generate a three-dimensional space model of the scene according to the picture;
[0007] receive real-time pictures of the scene captured by a plurality of capturing sub-modules included in a capturing module, different capturing sub-modules being located at different positions of the scene;
[0008] generate a single-point three-dimensional video corresponding to each capturing sub-module based on the real-time pictures of the scene captured by the capturing sub-modules and the three-dimensional space model, wherein the single-point three-dimensional video is a video of the current scene observed from the position of the capturing sub-module;
[0009] send a target three-dimensional video set to a terminal device, so that the terminal device adjusts the picture form of the scene displayed by the terminal device based on the target three-dimensional video set and the operation of the user when the user performs the operation;
[0010] wherein the target three-dimensional video set is obtained based on the single-point three-dimensional videos corresponding to the plurality of capturing sub-modules, and the picture form includes at least one of the following: display position, display angle, display size.
[0011] In a second aspect, the embodiments of the present application provide a terminal device, which comprises:
[0012] A display screen for displaying a picture;
[0013] A processor connected with the display screen, the processor being configured to receive a target stereoscopic video set sent by the video processing device;
[0014] The processor is configured to respond to the operation of the user to adjust the picture form of the scene displayed by the display screen based on the target stereoscopic video set;
[0015] The target stereoscopic video set is obtained based on single-point stereoscopic videos corresponding to the plurality of acquisition sub-modules in the acquisition module; the single-point stereoscopic video is a video of the current scene observed from the position of the acquisition sub-module; the scene is the surrounding environment of the video processing device; and the picture form includes at least one of the following: display position, display angle, and display size.
[0016] In a third aspect, an embodiment of the present application provides a video processing method, and the method comprises:
[0017] Obtaining a picture of a scene in which the video processing device is located, and generating a stereoscopic space model of the scene according to the picture;
[0018] Obtaining real-time pictures of the scene collected by a plurality of acquisition sub-modules included in an acquisition module, different acquisition sub-modules being located at different positions of the scene;
[0019] Generating single-point stereoscopic videos corresponding to the acquisition sub-modules based on the real-time pictures of the scene collected by the acquisition sub-modules and the stereoscopic space model, wherein the single-point stereoscopic video is a video of the current scene observed from the position of the acquisition sub-module;
[0020] Sending a target stereoscopic video set to a terminal device, so that the terminal device adjusts the picture form of the scene displayed by the terminal device based on the target stereoscopic video set and the operation of the user when the user performs the operation;
[0021] The target stereoscopic video set is obtained based on single-point stereoscopic videos corresponding to the plurality of acquisition sub-modules; and the picture form includes at least one of the following: display position, display angle, and display size.
[0022] In a fourth aspect, an embodiment of the present application provides a video processing system, and the video processing system comprises an acquisition module and the video processing device of any one of the first aspect;
[0023] The acquisition module comprises a plurality of acquisition sub-modules, and different acquisition sub-modules are located at different positions of a scene in which the video processing device is located.
[0024] This application provides a video processing device, a terminal device, and a video processing method. The video processing device includes a processor. The processor can acquire images of the scene where the video processing device is located and generate a stereoscopic spatial model of the scene based on the images. It receives real-time images of the scene acquired by multiple acquisition sub-modules included in the acquisition module. Based on the real-time images of the scene acquired by each acquisition sub-module and the stereoscopic spatial model, it generates a single-point stereoscopic video corresponding to the acquisition sub-module. Based on the single-point stereoscopic videos corresponding to multiple acquisition sub-modules, it generates a target stereoscopic video set and sends the target stereoscopic video set to the terminal device. For the terminal device, when the user operates, it can adjust the image of the scene displayed on the terminal device based on the target stereoscopic video set and the user's operation. Since different acquisition sub-modules are located at different positions in the scene, the video of the current scene observed from different acquisition sub-modules will have certain differences. For the terminal device, when the user operates the terminal device, the terminal device can adjust the image of the scene displayed on the terminal device based on the received target stereoscopic video set and the user's operation, so that the displayed scene image presents a three-dimensional stereoscopic effect, allowing the user to feel as if they are in the current scene, enhancing the user's sense of immersion and effectively improving the user experience. Attached Figure Description
[0025] To more clearly illustrate the implementation methods in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0026] Figure 1 A schematic diagram illustrating an application scenario provided in an embodiment of this application;
[0027] Figure 2 This is a schematic diagram of the structure of a display device provided in an embodiment of this application;
[0028] Figure 3 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application;
[0029] Figure 4 A schematic diagram illustrating scene segmentation provided in an embodiment of this application;
[0030] Figure 5 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;
[0031] Figure 6 A flowchart illustrating a method for determining a display screen, as provided in an embodiment of this application;
[0032] Figure 7A schematic diagram of an architecture of a video processing system provided by an embodiment of the present application is shown in FIG. 1.
[0033] Figure 8 A flowchart of a video processing method provided by an embodiment of the present application is shown in FIG. 2. Figure 1
[0034] Figure 9 A flowchart of a video processing method provided by an embodiment of the present application is shown in FIG. 2. Figure 2
[0035] Figure 10 A schematic diagram of a structure of a video processing device provided by an embodiment of the present application is shown in FIG. 3. DETAILED DESCRIPTION
[0036] In order to make the objectives, implementations and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementations of the present application with reference to the accompanying drawings of the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application.
[0037] It should be noted that the brief descriptions of the terms in the present application are only for the convenience of understanding the following described embodiments, and are not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.
[0038] In addition, the terms "comprise" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device comprising a series of components does not have to be limited to the clearly listed components, but can include other components that are not clearly listed or inherent to these products or devices.
[0039] At present, in remote communication, the audio and video module usually captures a scene at a certain fixed position, at which time the user at the remote end sees the scene at the fixed position, which makes the user feel less immersive and leads to poor user experience.
[0040] Based on this, the present application provides a video processing device, a terminal device and a video processing method, by setting multiple acquisition sub-modules at different positions in a scene, the video processing device can generate a single-point stereoscopic video according to the real-time scene pictures captured by the multiple acquisition sub-modules, and realize integration of the single-point stereoscopic video into a target stereoscopic video set. When the user performs clicking or sliding operations on the terminal device, based on the target stereoscopic video set, the picture form of the scene displayed by the terminal device changes, and the presented scene picture can achieve a three-dimensional stereoscopic effect, so that the user feels the feeling of moving and walking in the scene or rotating the eyes, which increases the user's sense of on-site immersion and improves the user experience.
[0041] Figure 1 An application scenario provided for an embodiment of the present application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the collection module 101 and the video processing device 102 are located in a certain scene, such as a conference room, a classroom, etc., and further include a plurality of terminal devices 103a-103b, etc. that interact with the video processing device 102, collectively referred to as terminal devices 103. The collection module 101 includes a plurality of collection sub-modules located at different positions in the scene, collects real-time pictures and audio information of the current scene from different positions, and sends the collected real-time pictures and audio information of the scene to the video processing device 102 for processing. The video processing device 102 sends the processed information to each terminal device 103. Based on the operation of the user, the terminal device 103 determines the corresponding collection sub-module to display the real-time pictures and audio information collected by the collection sub-module, so as to change the picture form of the display picture and make the picture present a three-dimensional effect.
[0042] The video processing device 102 can be a display device or other processing device. The display device can have various implementations, such as a large screen, a monitor, an electronic bulletin board, an electronic table, etc. The number of terminal devices in communication with the video processing device can be one or more.
[0043] Figure 2 A structural diagram of a display device provided for an embodiment of the present application is shown in FIG. 2. Figure 2 As shown in FIG. 2, the display device 200 includes at least one of a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.
[0044] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, a RAM, a ROM, a first interface to an n-th interface for input / output.
[0045] The display 260 includes a display screen component for presenting a picture, a driving component for driving image display, a component for receiving an image signal output from the controller, and a component for displaying video content, image content, and a menu control interface, and a user control UI interface.
[0046] The display 260 can be a liquid crystal display, an OLED display, and a projection display, and can also be a projection device and a projection screen.
[0047] The communicator 220 is a component for communicating with external devices or servers according to various communication protocol types. For example, the communicator can include at least one of a Wifi module, a Bluetooth module, a wired Ethernet module, and other network communication protocol chips or near field communication protocol chips, and an infrared receiver. The display device 200 can establish transmission and reception of control signals and data signals with an external control device or a server through the communicator 220.
[0048] The user interface can be used to receive control signals of a control device such as an infrared remote controller.
[0049] The detector 230 is used to collect signals of an external environment or interaction with the outside. For example, the detector 230 includes a light receiver for collecting ambient light intensity, or an image collector such as a camera for collecting external environment scenes, user attributes, or user interaction gestures, or a sound collector such as a microphone for receiving external sounds.
[0050] The external device interface 240 can include, but is not limited to, any one or more of the following: a high-definition multimedia interface (HDMI), an analog or digital high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. It can also be a composite input / output interface formed by a plurality of the above interfaces.
[0051] The tuner demodulator 210 receives broadcast television signals through wired or wireless reception, and demodulates audio and video signals and EPG data signals from a plurality of wireless or wired broadcast television signals.
[0052] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different split devices, i.e., the tuner demodulator 210 can also be in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.
[0053] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored on the memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command for selecting a UI object displayed on the display 260, the controller 250 can perform an operation related to the object selected by the user command.
[0054] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), a RAM (random access memory), a ROM (read-only memory), a first interface to an n-th interface for input / output, a communication bus, and the like.
[0055] The user can input a user command through a graphic user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphic user interface (GUI). Alternatively, the user can input a user command by inputting a specific sound or gesture, and the user input interface receives the user input command by recognizing the sound or gesture through a sensor.
[0056] A "user interface" is a medium interface for interaction and information exchange between an application program or an operating system and a user, which realizes conversion between an internal form of information and a form acceptable by the user. A commonly used form of the user interface is a graphic user interface (GUI), which refers to a user interface related to computer operation displayed in a graphic manner. The user interface can be an icon, a window, a control, and the like interface elements displayed in a display screen of an electronic device, wherein the control can include an icon, a button, a menu, a tab, a text box, a dialog box, a status bar, a navigation bar, a widget, and the like visible interface elements.
[0057] The technical solutions of the present application will be described in detail below in combination with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in detail in some embodiments.
[0058] Figure 3 A structural schematic diagram of a video processing device according to an embodiment of the present application is shown in FIG. 3. As shown in FIG. 3, the video processing device 300 includes: Figure 3
[0059] The processor 301 is configured to acquire a picture of a scene in which the video processing device 300 is located, and generate a three-dimensional space model of the scene according to the picture.
[0060] The receiving module is configured to receive real-time pictures of the scene collected by a plurality of collection sub-modules included in the collection module, wherein different collection sub-modules are located at different positions in the scene.
[0061] generate a single-point stereoscopic video corresponding to the acquisition sub-module based on the real-time picture of the scene collected by the acquisition sub-module and the stereoscopic space model, wherein the single-point stereoscopic video is a video of the current scene observed from the position of the acquisition sub-module;
[0062] send the target stereoscopic video set to the terminal device, so that the terminal device adjusts the picture form of the scene displayed by the terminal device based on the target stereoscopic video set and the operation of the user when the user performs the operation;
[0063] wherein the target stereoscopic video set is obtained based on single-point stereoscopic videos corresponding to a plurality of acquisition sub-modules; and the picture form includes at least one of the following: display position, display angle, and display size.
[0064] In an implementation scenario, the video processing device 300 can include a camera configured to collect a picture of a scene in which the video processing device 300 is located and send the picture to the processor 301.
[0065] In another implementation scenario, there can be a separate camera that can communicate with the video processing device 300. The camera collects a picture of the scene and sends the picture to the video processing device 300.
[0066] wherein the scene is the surrounding environment of the location of the video processing device 300, such as a conference room, a classroom, etc., and the picture of the scene can be a plurality of images of the scene at different angles obtained by the camera shooting the scene in all directions, or a video containing all information of the scene.
[0067] The stereoscopic space model of the scene is a three-dimensional model corresponding to the scene, which can represent a plurality of characteristics of the scene, including but not limited to shape, size, volume, and spatial relationship between different objects in the scene.
[0068] In some embodiments, before receiving the real-time pictures of the scene collected by the plurality of acquisition sub-modules, the processor 301 is further configured to: segment the scene according to the stereoscopic space model, and determine the number of acquisition sub-modules and the placement positions of the plurality of acquisition sub-modules in the scene.
[0069] In an implementation scenario, the number and distribution of the acquisition sub-modules are related to the size and structure of the scene. For example, when the area of the scene is large, more acquisition sub-modules need to be set to obtain more real-time pictures of the scene collected from different positions of the scene.
[0070] In an implementation scenario, when the scene is segmented, the scene can be segmented by matrix points, i.e., the plurality of acquisition sub-modules are arranged in the form of a matrix. For details, refer to FIG. 1. Figure 4 FIG. 1 shows a schematic diagram of a scene segmented by matrix points.Figure 4 A schematic diagram of scene segmentation is provided for an embodiment of the present application, wherein 400 is used to represent a scene in which the video processing device 300 is located, and each matrix point in a11-a43 represents a collection submodule, Figure 4 Four rows and three columns, i.e., a total of 12 matrix points, are shown in the above table, i.e., 12 collection submodules are needed to be placed in the scene.
[0071] Figure 4 Only a case of scene segmentation and distribution of multiple collection submodules is exemplarily shown, and in addition to matrix point segmentation of the scene, the scene can also be segmented by other segmentation manners, which is not limited in the present application.
[0072] Since different collection submodules are located at different positions of the scene, the real-time pictures of the scene collected by different collection submodules also have certain differences.
[0073] In an implementation scenario, each collection submodule can include a camera, and the camera is used to take pictures of the scene in all directions to obtain real-time pictures of the scene.
[0074] In some embodiments, the real-time picture of the scene includes multiple images; and the processor 301 is specifically configured to:
[0075] project the images onto the stereoscopic space model to generate single-point stereoscopic images corresponding to the images;
[0076] synthesize the single-point stereoscopic images of the multiple images to obtain a single-point stereoscopic video corresponding to the collection submodule.
[0077] For any collection submodule, the real-time picture of the scene collected by the collection submodule includes multiple images, and each image can be projected onto the stereoscopic space model to obtain a single-point stereoscopic image corresponding to the image. The single-point stereoscopic image is an image of the current scene in three dimensions observed from the position of the collection submodule.
[0078] The target stereoscopic video set is obtained based on the single-point stereoscopic videos corresponding to the multiple collection submodules, and in an implementation scenario, the processor 301 is further configured to: crop, splice and synthesize the single-point stereoscopic videos corresponding to the multiple collection submodules to generate the target stereoscopic video set, and the target stereoscopic video set includes target stereoscopic videos corresponding to the multiple collection submodules.
[0079] For any acquisition sub-module, the corresponding target stereoscopic video is consistent with the perspective of the single-point stereoscopic video, and the single-point stereoscopic videos corresponding to different acquisition sub-modules are different, and thus the target stereoscopic videos corresponding to different acquisition sub-modules also have certain differences.
[0080] Since the target stereoscopic video fuses the single-point stereoscopic videos of other acquisition sub-modules, compared with the single-point stereoscopic video, the target stereoscopic video corresponding to the acquisition sub-module contains more information of the scene.
[0081] In another implementation scenario, the single-point stereoscopic videos corresponding to the multiple acquisition sub-modules can also be integrated to obtain a target stereoscopic video set. At this time, the target stereoscopic video corresponding to each acquisition sub-module is the same as the single-point stereoscopic video.
[0082] In some embodiments, the processor 301 is further configured to receive audio information collected by the multiple acquisition sub-modules, and send the audio information to the terminal device.
[0083] In an implementation scenario, each acquisition sub-module can further include a microphone that can collect audio information. Since different acquisition sub-modules are located at different positions, when a sound is emitted at a certain position, the audio information collected by different acquisition sub-modules will also have certain differences, such as the size and clarity of the sound.
[0084] After the audio information collected by the multiple acquisition sub-modules is sent to the terminal device, the terminal device can select the audio information collected by a certain acquisition sub-module to play according to requirements.
[0085] The embodiment of the present application provides a video processing device 300, which includes a processor 301. The processor 301 can obtain a picture of a scene where a video device is located, and generate a stereoscopic space model of the scene according to the picture. Real-time pictures of the scene collected by multiple acquisition sub-modules included in an acquisition module are received. For each acquisition sub-module, a single-point stereoscopic video corresponding to the acquisition sub-module is generated based on the real-time picture of the scene collected by the acquisition sub-module and the stereoscopic space model. A target stereoscopic video set obtained based on the single-point stereoscopic videos corresponding to the multiple acquisition sub-modules is sent to a terminal device. Since different acquisition sub-modules are located at different positions of the scene, the videos of the current scene observed from different acquisition sub-modules, i.e., the single-point stereoscopic videos, will have certain differences. When a user operates the terminal device, the picture form of the scene displayed by the terminal device is adjusted based on the target stereoscopic video set and the operation of the user, so that the displayed scene picture presents a stereoscopic effect, thereby improving the sense of on-site immersion and improving the user experience.
[0086] Figure 5 A structural schematic diagram of a terminal device provided by the embodiment of the present application is shown in FIG. 2.Figure 5 As shown, the terminal device comprises:
[0087] a display screen 501 for displaying a picture;
[0088] a processor 502 connected with the display screen 501, the processor 502 being configured to receive a target stereoscopic video set sent by a video processing device;
[0089] based on the target stereoscopic video set, responding to the operation of a user to adjust the picture form of a scene displayed by the display screen 501;
[0090] wherein the target stereoscopic video set is obtained based on single-point stereoscopic videos corresponding to a plurality of acquisition sub-modules in an acquisition module; the single-point stereoscopic video is a video of a current scene observed from the position of the acquisition sub-module; the scene is the surrounding environment of the video processing device; and the picture form comprises at least one of the following: display position, display angle, and display size.
[0091] The operation of the user includes but is not limited to clicking, sliding, etc. The terminal device of the present application can be a display device capable of displaying a picture. In an implementation scenario, the terminal device can be a display device with touch function, such as a mobile phone, a tablet computer, etc., at this time the user can click or slide the display screen 501 by fingers or a stylus, etc.
[0092] In another implementation scenario, the terminal device can also be a display device without touch function, such as a computer, etc., at this time the user can click or slide the display screen 501 by a mouse, etc.
[0093] In an implementation scenario, when the user clicks on the display screen 501 based on the demand, the click is on a certain position on the real-time picture of the current scene displayed, based on which the corresponding acquisition sub-module can be determined, and the real-time picture of the scene collected by the acquisition sub-module is displayed, realizing the conversion of the acquisition sub-module, so that the picture displayed by the display screen 501 is transformed and moved accordingly, the picture presents a three-dimensional stereoscopic effect, and the user has the experience of walking in the current scene.
[0094] In another implementation scenario, the terminal device can be combined with Figure 4 As shown, if the user slides along the X-axis, the picture can be rotated in the left-right direction, and if the user slides along the Y-axis, the picture can be rotated in the up-down direction. Further, if the user slides at a fixed position, the picture can be rotated, and the user can see the picture in multiple directions, effectively improving the on-site experience.
[0095] For example, still referring to Figure 4As shown, if the position clicked by the user on the display screen 501 is the position of the a12 matrix point, the a12 matrix point corresponding acquisition submodule can be converted to perform shooting and sound pickup, and the angle rotation can be realized by sliding at a12 to see the front, back, up, down, left and right pictures.
[0096] The terminal device provided by the embodiment of the present application comprises a display screen 501 and a processor 502, wherein the processor 502 can respond to the operation of the user based on the target stereoscopic video set sent by the video processing device, adjust the picture form of the scene displayed by the display screen 501, so that the displayed scene presents a three-dimensional stereoscopic effect, improves the on-site immersion of the user, and improves the user experience.
[0097] Figure 6 The method for determining the display picture provided by the embodiment of the present application is shown in a flowchart, and in some embodiments, the target stereoscopic video set comprises target stereoscopic videos corresponding to a plurality of acquisition submodules; when the processor 502 responds to the operation of the user based on the target stereoscopic video set to adjust the picture form of the scene displayed by the display screen 501, it is specifically used for executing the steps shown in the following. Figure 6
[0098] S601: based on the target stereoscopic video set, the position of the speaker and the face direction of the speaker are obtained, and according to the position of the speaker and the face direction of the speaker, a first target acquisition submodule is determined; or, based on the operation of the user, a first target acquisition submodule is determined.
[0099] S602: in the target stereoscopic video set, the target stereoscopic video corresponding to the first target acquisition submodule is obtained to display based on the target stereoscopic video, and the picture form of the scene displayed by the display screen is adjusted.
[0100] In an implementation scenario, taking a conference scene as an example, the face features and body features of the current participants can be obtained based on the target stereoscopic video set, and based on the face features and body features, the position of the current speaker and the face direction of the speaker are further determined among the participants.
[0101] The position of the speaker and the face direction thereof can also be determined by other ways, which are not limited by the present application.
[0102] For example, still referring to Figure 4 As shown, if the speaker is currently located at the a11 matrix point and faces the a21 matrix point, the optimal collection position can be determined as the a21 matrix point, and the collection submodule corresponding to the a21 matrix point is taken as the first target collection submodule. When the speaker walks to the a43 matrix point and faces the a32 matrix point, the optimal collection position is changed from the a21 matrix point to the a32 matrix point, and the collection submodule corresponding to the a32 matrix point is taken as the first target collection submodule.
[0103] In the above example, when the position and the facial orientation of the speaker change, the first target collection submodule also changes, and therefore the target stereoscopic video corresponding to the first target collection submodule also changes, so that the picture form of the scene displayed by the terminal device changes.
[0104] In the above example, the position adjacent to the current position of the speaker in the direction of the facial orientation of the speaker is taken as the optimal collection position, and the collection submodule corresponding to the optimal collection position is taken as the first target collection submodule. The first target collection submodule can also be set according to actual needs.
[0105] In another implementation scenario, the first target collection submodule can also be determined based on the operation of the user. For example, if the user needs to watch a part of the scene, the user can click the part of the scene displayed on the display screen 501 to determine the corresponding first target collection submodule, so as to change the picture form of the displayed scene and increase the flexibility of the display picture of the terminal device.
[0106] For example, still referring to Figure 4 As shown, if the speaker is located at the a41 matrix point, but wants to adjust the field of view to watch the content on the video processing device, the user can click the position of the video processing device in the scene displayed on the display screen 501 to switch to the real-time picture of the scene collected by the collection submodule corresponding to the a22 matrix point, that is, to display the target stereoscopic video corresponding to the collection submodule corresponding to the a22 matrix point.
[0107] In some embodiments, for the user of the terminal device, it is not only necessary to watch the picture displayed on the display screen 501, but also necessary to play the current audio information. The displayed target stereoscopic video and the played audio information can be provided by the same collection submodule, that is, the microphone and the camera are switched synchronously; or the displayed target stereoscopic video and the played audio information can be provided by different collection submodules, that is, the microphone and the camera are arranged respectively.
[0108] In an implementation scenario, the processor 502 is further configured to: acquire audio information collected by the first target collection submodule, and play the audio information collected by the first target collection submodule.
[0109] Or,
[0110] The second target acquisition sub-module closest to the position of the speaker is determined based on the position of the speaker, and audio information collected by the second target acquisition sub-module is acquired to play based on the audio information collected by the second target acquisition sub-module.
[0111] In an implementation scenario, the terminal device further includes a loudspeaker, and the audio information is played through the loudspeaker.
[0112] In an implementation scenario, if the audio information collected by the first target acquisition sub-module is played, the camera and the microphone are provided by the same acquisition sub-module at this time because the target stereoscopic video displayed is also provided by the first target acquisition sub-module. When the camera is switched, the microphone is also switched synchronously.
[0113] In another implementation scenario, to obtain clearer audio information, the second target acquisition sub-module closest to the position of the speaker is set to collect audio information, that is, the microphone follows the speaker, and when the position of the speaker changes, the second target acquisition sub-module also changes correspondingly. At this time, because the target stereoscopic video displayed is provided by the first target acquisition sub-module, the camera and the microphone are provided by different acquisition sub-modules.
[0114] For example, still referring to FIG. 4A, Figure 4 If the speaker is located at the position of the a41 matrix point, the acquisition sub-module at the position of the a41 matrix point can be set as the second acquisition sub-module, and the audio information collected by the second acquisition sub-module is played. The first acquisition sub-module can be determined based on the operation of the user or the current position and face orientation of the speaker.
[0115] For example, in a teaching scenario, a teacher walks in a classroom, and a blackboard is located at a fixed position. For a student using a terminal device to listen to a class, the teacher's voice needs to be heard clearly, and therefore the microphone needs to be fixed to follow the teacher, that is, the best second target acquisition sub-module is determined according to the position of the teacher, and the audio information collected by the second target acquisition sub-module is played. At the same time, if the content on the blackboard needs to be watched, the user can click the position of the blackboard in the picture displayed on the display screen, and the corresponding first target acquisition sub-module is determined based on the position, so that the target stereoscopic video corresponding to the first target acquisition sub-module displayed can display the content on the blackboard more clearly.
[0116] In summary, the first target acquisition submodule can be determined based on the position and face orientation of the speaker, or can be determined according to the operation of the user, so that the first target acquisition submodule corresponding to the target stereoscopic video is displayed, so that the picture form of the scene displayed by the display screen changes. For the audio information played, it can be provided by the first target acquisition submodule, or it can be provided by the second target acquisition submodule closest to the speaker, so as to obtain clearer audio information and increase flexibility.
[0117] Figure 7 An architecture schematic diagram of a video processing system provided by an embodiment of the present application is shown in FIG. 1. Figure 7 As shown in the figure, the video processing system includes an acquisition module 600 and the video processing device 300 described in the above embodiments.
[0118] The acquisition module 600 includes a plurality of acquisition submodules, and different acquisition submodules are located at different positions in the scene where the video processing device 300 is located.
[0119] In an implementation scenario, the acquisition module 600 can communicate with the video processing device 300 through the UVC protocol to realize transmission of audio and video data.
[0120] The video processing device 300 can run an Android system and / or a Windows system.
[0121] The specific processing process of the acquisition module 600 and the video processing device 300 can refer to the above embodiments, which will not be described in detail here.
[0122] Figure 8 A flowchart of a video processing method provided by an embodiment of the present application is shown in FIG. 2. Figure 1 The method can be executed by a video processing device or an acquisition module. As shown in the figure, the method includes the following steps: Figure 8
[0123] S801: Obtain a picture of a scene where the video processing device is located, and generate a stereoscopic space model of the scene according to the picture.
[0124] The picture of the scene can be an image including multiple angles of the scene, or a video containing all information of the scene, etc. The acquisition of the picture of the scene can be performed by a camera of the video processing device, or can be performed by an independent camera, which is not limited by the present application.
[0125] S802: Obtain real-time pictures of the scene collected by a plurality of acquisition submodules included in the acquisition module, and different acquisition submodules are located at different positions in the scene.
[0126] The scene real-time picture includes multiple frames of images, and the scene real-time pictures collected by different acquisition sub-modules are different due to different positions of the acquisition sub-modules.
[0127] S803: generating a single-point stereoscopic video corresponding to the acquisition sub-module based on the scene real-time picture collected by the acquisition sub-module and the three-dimensional space model.
[0128] The single-point stereoscopic video is a video of the current scene observed from the position of the acquisition sub-module.
[0129] In an implementation scenario, for any acquisition sub-module, each frame of image in the scene real-time picture collected by the acquisition sub-module can be projected to the three-dimensional space model to obtain a single-point stereoscopic image corresponding to the image, and the single-point stereoscopic videos corresponding to the multiple frames of images are integrated to obtain a single-point stereoscopic video corresponding to the acquisition sub-module.
[0130] The scene real-time pictures collected by different acquisition sub-modules are different, and thus the single-point stereoscopic videos corresponding to different acquisition sub-modules are also different.
[0131] S804: sending the target stereoscopic video set to a terminal device, so that the terminal device adjusts a picture form of the scene displayed by the terminal device based on the target stereoscopic video set and an operation of a user when the user performs the operation.
[0132] The target stereoscopic video set is obtained based on the single-point stereoscopic videos corresponding to the multiple acquisition sub-modules, and the picture form includes at least one of a display position, a display angle, and a display size.
[0133] The target stereoscopic video set includes target stereoscopic videos corresponding to the multiple acquisition sub-modules.
[0134] In an implementation scenario, for the terminal device, a corresponding acquisition sub-module can be determined based on an operation of a user, and a target stereoscopic video corresponding to the acquisition sub-module is obtained from the target stereoscopic video set for display. Due to the conversion of the acquisition sub-modules and the difference between the target stereoscopic videos corresponding to different acquisition sub-modules, the picture form of the scene displayed by the terminal device can be adjusted.
[0135] The embodiment of the present application provides a video processing method, a picture of a scene where a video processing device is located is acquired, and a three-dimensional space model of the scene is generated according to the picture. Real-time pictures of the scene collected by a plurality of collection sub-modules in a collection module are acquired, and for any collection sub-module, a single-point three-dimensional video corresponding to the collection sub-module is generated based on the real-time picture of the scene collected by the collection sub-module and the three-dimensional space model. A target three-dimensional video set obtained based on the single-point three-dimensional videos corresponding to the plurality of collection sub-modules is sent to a terminal device, and for the terminal device, when a user performs an operation, the picture form of the scene displayed by the terminal device is adjusted based on the target three-dimensional video set and the operation of the user, so that the displayed scene picture presents a three-dimensional effect, the user can have a feeling of being located in the current scene, the on-site identification of the user is improved, and the user experience is effectively improved.
[0136] Figure 9 A flowchart of a video processing method provided by the embodiment of the present application Figure 2 , as shown in the figure, the method can include: Figure 9
[0137] S901: The video processing device acquires a picture of a scene where the video processing device is located, and generates a three-dimensional space model of the scene according to the picture.
[0138] S902: The video processing device divides the scene according to the three-dimensional space model, determines the number of collection sub-modules included in a collection module and the placement positions of the plurality of collection sub-modules in the scene.
[0139] S903: Each collection sub-module collects a real-time picture of the scene, and sends the real-time picture of the scene to the video processing device.
[0140] S904: The video processing device generates a corresponding single-point three-dimensional video according to the real-time picture of the scene collected by each collection sub-module.
[0141] S905: The video processing device generates a target three-dimensional video set according to the single-point three-dimensional videos of the plurality of collection sub-modules, and sends the target three-dimensional video set to a terminal device.
[0142] S906: The terminal device determines whether it is a default setting at present. If yes, step S907 is performed; if no, step S908 is performed.
[0143] S907: The terminal device determines the position and face orientation of a speaker based on the target three-dimensional video set, to determine a first target collection sub-module, displays based on the target three-dimensional video of the first target collection sub-module, and plays based on audio information collected by the first target collection sub-module.
[0144] S908: The terminal device receives the user-selected pickup configuration and the user operation, determines a first target acquisition sub-module based on the user operation, and displays a target stereoscopic video based on the first target acquisition sub-module.
[0145] S909: It is determined whether the pickup and shooting are synchronously converted according to the pickup configuration. If yes, step S910 is performed; if no, step S911 is performed.
[0146] S910: The terminal device plays audio information acquired based on the first target acquisition sub-module.
[0147] S911: The terminal device determines a second target acquisition sub-module closest to the current speaker according to the position of the current speaker, and plays audio information acquired based on the second target acquisition sub-module.
[0148] The specific processing procedures of the acquisition module, the video processing device, and the terminal device can refer to the above embodiments, and will not be described in detail here.
[0149] Figure 10 A structural schematic diagram of a video processing apparatus provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the video processing apparatus 1000 includes an acquisition module 1001, a generation module 1002, and a sending module 1003. Figure 10
[0150] The acquisition module 1001 is configured to acquire a picture of a scene in which the video processing device is located.
[0151] The generation module 1002 is configured to generate a stereoscopic space model of the scene according to the picture.
[0152] The acquisition module 1001 is further configured to acquire real-time pictures of the scene acquired by a plurality of acquisition sub-modules included in an acquisition module; different acquisition sub-modules are located at different positions in the scene.
[0153] The generation module 1002 is further configured to generate a single-point stereoscopic video corresponding to each acquisition sub-module based on the real-time pictures of the scene acquired by the acquisition sub-modules and the stereoscopic space model; the single-point stereoscopic video is a video of the current scene observed from the position of the acquisition sub-module.
[0154] The sending module 1003 is configured to send a target stereoscopic video set to a terminal device, so that the terminal device adjusts a picture form of the scene displayed by the terminal device based on the target stereoscopic video set and a user operation when the user performs the operation.
[0155] The target stereoscopic video set is obtained based on the single-point stereoscopic videos corresponding to the plurality of acquisition sub-modules; the picture form includes at least one of the following: a display position, a display angle, and a display size.
[0156] The video processing apparatus provided in the embodiments of the present application can execute the video processing method in the method embodiments, and the implementation principles and technical effects are similar, which will not be repeated here. It should be noted that the above Figure 10 The division of each module shown is only a schematic, and the division of each module and the naming of each module by the present application are not limited.
[0157] The present application also provides a computer readable storage medium, the computer readable storage medium has computer execution instructions stored thereon, when the computer execution instructions are executed by a processor, the method described in any one of the above embodiments is implemented.
[0158] The computer readable storage medium can include a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes, and specifically, the computer readable storage medium has program instructions stored therein, and the program instructions are used for the method in the above embodiments.
[0159] The present application also provides a computer program product, including a computer program, which is executed by a processor to implement the method described in any one of the above embodiments.
[0160] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the above embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0161] In order to facilitate explanation, the above description has been made in combination with specific embodiments. However, the above exemplary discussion is not intended to exhaust or limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A video processing device, comprising: The video processing device comprises: a processor configured to acquire a picture of a scene in which the video processing device is located, and generate a three-dimensional space model of the scene according to the picture; receive real-time pictures of the scene collected by a plurality of collection sub-modules comprised in a collection module, wherein the different collection sub-modules are located at different positions in the scene; generate a single-point three-dimensional video corresponding to each of the collection sub-modules based on the real-time pictures of the scene collected by the collection sub-modules and the three-dimensional space model, wherein the single-point three-dimensional video is a video of the current scene observed from the position of the collection sub-module; send a target three-dimensional video set to a terminal device, so that the terminal device adjusts a picture form of the scene displayed by the terminal device based on the target three-dimensional video set and an operation of a user when the user performs the operation; wherein the target three-dimensional video set is obtained based on the single-point three-dimensional videos corresponding to the plurality of collection sub-modules, and the picture form comprises at least one of the following: a display position, a display angle, and a display size.
2. The video processing device of claim 1, wherein, The real-time pictures of the scene comprise a plurality of images, and the processor is specifically configured to: project the images onto the three-dimensional space model to generate single-point three-dimensional images corresponding to the images; and synthesize the single-point three-dimensional images of the plurality of images to obtain the single-point three-dimensional video corresponding to each of the collection sub-modules.
3. The video processing device of claim 1 or 2, wherein, The processor is further configured to: crop, splice, and synthesize the single-point three-dimensional videos corresponding to the plurality of collection sub-modules to generate the target three-dimensional video set, wherein the target three-dimensional video set comprises target three-dimensional videos corresponding to the plurality of collection sub-modules.
4. The video processing device of claim 1, wherein, The processor is further configured to: segment the scene according to the three-dimensional space model to determine the number of collection sub-modules and the placement positions of the plurality of collection sub-modules in the scene.
5. The video processing device of claim 1, wherein, The processor is further configured to: receive audio information collected by the plurality of collection sub-modules, and send the audio information to the terminal device.
6. A terminal device, characterized by comprising: The terminal device comprises: a display screen configured to display a picture; and a processor connected to the display screen, the processor being configured to receive a target three-dimensional video set sent by a video processing device, respond to an operation of a user based on the target three-dimensional video set to adjust a picture form of a scene displayed by the display screen, wherein the target three-dimensional video set is obtained based on single-point three-dimensional videos corresponding to a plurality of collection sub-modules in a collection module, the single-point three-dimensional video is a video of a current scene observed from the position of the collection sub-module, the scene is a surrounding environment in which the video processing device is located, and the picture form comprises at least one of the following: a display position, a display angle, and a display size.
7. The terminal device according to claim 6, characterized by The target three-dimensional video set comprises target three-dimensional videos corresponding to the plurality of collection sub-modules, and the processor is specifically configured to: The processor is further configured to: acquire audio information collected by the first target collection sub-module, and play the audio information based on the audio information collected by the first target collection sub-module; or, 8. The terminal device according to claim 7, characterized by acquire audio information collected by the second target collection sub-module closest to the position of the speaker based on the position of the speaker, and play the audio information based on the audio information collected by the second target collection sub-module. The method comprises: acquiring a picture of a scene in which the video processing device is located, and generating a three-dimensional space model of the scene according to the picture; acquiring real-time pictures of the scene collected by a plurality of collection sub-modules included in a collection module, different collection sub-modules being located at different positions in the scene; 9. A method for video processing, the method comprising: generating a single-point three-dimensional video corresponding to each collection sub-module based on the real-time pictures of the scene collected by the collection sub-modules and the three-dimensional space model, wherein the single-point three-dimensional video is a video of the current scene observed from the position of the collection sub-module; sending a target three-dimensional video set to a terminal device, so that the terminal device adjusts a picture form of the scene displayed by the terminal device based on the target three-dimensional video set and an operation of a user when the user performs the operation; wherein the target three-dimensional video set is obtained based on the single-point three-dimensional videos corresponding to the plurality of collection sub-modules, and the picture form comprises at least one of the following: a display position, a display angle, and a display size. The video processing system comprises a collection module and the video processing device according to any one of claims 1-5; wherein the collection module comprises a plurality of collection sub-modules, different collection sub-modules being located at different positions in a scene in which the video processing device is located. 10. A video processing system characterized by