Ultra-high-definition camera system with multi-format outputs
The ultra-high-definition camera system with multi-panel output uses an ultra-high-definition image sensor and target detection module to cut and output horizontal and vertical sub-screens, solving the problem that cameras cannot meet the requirements of both horizontal and vertical screens at the same time, thus improving the richness of video content and user experience.
Patent Information
- Application Number
- PCT/CN2024/123587
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-12
- Filing Date
- 2024-10-09
- Publication Date
- 2026-02-19
AI Technical Summary
Existing cameras cannot simultaneously meet the diverse needs of both landscape and portrait modes, resulting in limited richness and diversity of video content, and increased costs for additional equipment and manpower.
Design an ultra-high-definition camera system with multi-frame output, which adopts an ultra-high-definition image sensor, processor, memory and target detection module. It receives screen ratio instructions through user interface or network communication, cuts and outputs horizontal and vertical sub-screens, and supports multiple display ratios and custom settings.
It enables video output to adapt to diverse display needs, improves user experience, simplifies operation processes, and reduces equipment and labor costs.
Smart Images

Figure CN2024123587_19022026_PF_FP_ABST
Abstract
Description
A multi-aspect output ultra-high-definition camera system TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a multi-aspect output ultra-high-definition camera system. BACKGROUND
[0002] Under the background of the increasing popularity of current multimedia communication and live broadcast, the camera as the core equipment directly determines the diversity of video content and the viewing experience. However, most of the cameras on the market currently only support single horizontal screen or single vertical screen output, which cannot meet the diversified needs of users considering horizontal screen and vertical screen playback at the same time. This limitation is particularly prominent in multiple application scenarios, such as short video platforms (such as TikTok, Kuaishou), live streaming, online education, remote conference, etc. These scenarios often need to quickly switch between horizontal and vertical screen perspectives to optimize the viewing experience.
[0003] The existing solutions mainly have the following schemes, but all have disadvantages, specifically:
[0004] Second framing and cutting of horizontal screen to vertical screen: as shown in FIG. 1, this scheme takes a second framing in the video shot in horizontal screen to cut the vertical screen picture. However, due to the limitation of the framing range, the vertical screen picture has extremely limited space in the vertical direction, and can only capture close-up shots, making it difficult to show a wider scene or long shot, which severely limits the richness and diversity of video content.
[0005] Setting up a special vertical screen camera: in order to meet the needs of vertical screen playback, some users choose to use a special vertical screen camera or rotate a traditional camera by 90°, however, this method requires additional camera positions and camera personnel, increasing the cost of equipment and labor.
[0006] If a camera with an ultra-high-definition large-resolution sensor is designed, which can output multiple aspects simultaneously in horizontal and vertical screens, the above technical problems can be solved. SUMMARY
[0007] To solve the above technical problems, the present application provides a multi-aspect output ultra-high-definition camera system, which effectively solves the problems of current camera systems in adapting to multiple display ratios, and provides users with more rich and high-quality video content experience.
[0008] The multi-aspect output ultra-high-definition camera system provided by the present application comprises:
[0009] An ultra-high-definition image sensor for capturing image pictures;
[0010] A processor connected to the image sensor for processing the image pictures;
[0011] a memory, connected with the processor, for storing program instructions;
[0012] The program instructions, when executed, cause the processor to perform the following steps:
[0013] determining a target in the image frame and obtaining a position of the target;
[0014] receiving a frame output instruction by the processor, wherein the frame output instruction includes at least two output modes of landscape and portrait, and frame cutting position and size;
[0015] According to the at least two output modes and frame cutting position and size, at least two sub-frames including a landscape sub-frame and a portrait sub-frame are cut from the image frame captured by the image sensor according to the position of the target and output, wherein the at least two sub-frames include a landscape sub-frame and a portrait sub-frame to meet the display ratio requirements of different display devices;
[0016] Meanwhile, according to the target recognized by the target detection module, the position and size of the sub-frame are automatically moved and scaled in the image frame captured by the image sensor.
[0017] Preferably, the memory is further used for:
[0018] storing preset output mode templates, wherein the output mode templates include HD template, 4K template, 8K template, landscape template and portrait template, and a user can select an output mode template to generate the frame output instruction.
[0019] Preferably, the output mode templates further include templates for customizing frame width-height ratio and frame shape.
[0020] Preferably, the resolution of the image sensor is at least 4K, and the ratio of the image sensor includes 4:3, 16:9 and 1:1.
[0021] Preferably, the receiving mode of the processor receiving the frame output instruction includes one or a combination of both of user interface input and network communication receiving.
[0022] Preferably, the position of the target is obtained by using a preset target detection module automatically and / or based on a configured user interface manually.
[0023] Preferably, it further includes a network communication module for simultaneously transmitting the at least two sub-frames including a landscape sub-frame and a portrait sub-frame to corresponding display devices in real time to meet the display ratio of the corresponding display devices.
[0024] Preferably, the target detection module includes:
[0025] a target positioning unit configured to determine a target in a current image frame and obtain a position of the target;
[0026] a target tracking unit configured to track the target in successive image frames and update the position of the target.
[0027] Compared with the prior art, the super-high-definition camera system with multi-surface output provided by the application has the following beneficial effects:
[0028] The application realizes adaptive video output by setting the picture ratio and selecting the cutting position through user input or remote instruction.
[0029] Specifically, first, the picture ratio instruction and the picture cutting position are input through the user interface of the camera body or network communication; then, the processor decodes and analyzes the received picture ratio instruction in detail, and performs cutting processing according to the output mode and picture cutting position information in the picture ratio instruction to obtain sub-pictures including horizontal and vertical pictures; finally, the obtained sub-picture data is transmitted to the corresponding display device in real time, and necessary synchronization operation is performed to ensure that the picture can be accurately presented, thereby realizing adaptive video output, significantly improving user experience and being suitable for various display scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0030] FIG. 1 is a schematic diagram of the prior art in which a vertical picture is intercepted in a horizontal picture;
[0031] FIG. 2 is a flowchart of program instructions in the super-high-definition camera system with multi-surface output provided by the application;
[0032] FIG. 3 is a structural diagram of the super-high-definition camera system with multi-surface output provided by the application;
[0033] FIG. 4 is a schematic diagram of target movement in the second embodiment of the application;
[0034] FIG. 5 is a schematic diagram of the position of the target in the coordinate system in the second embodiment of the application. DETAILED DESCRIPTION
[0035] The application will be described in further detail below with reference to the drawings and embodiments. It is to be understood that the specific embodiments described herein are merely illustrative of the application and are not intended to limit the scope of the application. In addition, it should be noted that, for the sake of brevity, only structures related to the application are shown and described in the drawings. In addition, the embodiments in the application and the features in the embodiments can be combined with each other without conflict.
[0036] In addition, it should be noted that, for the sake of brevity, only structures related to the application are shown and described in the drawings. Before discussing the example embodiments in more detail, it should be mentioned that some of the example embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The processes can be terminated when their operations are completed, but can also have additional steps not included in the flowcharts. The processes can correspond to methods, functions, routines, subroutines, subprograms, etc. Embodiments
[0037] The application provides a multi-aspect output ultra-high-definition camera system. Referring to FIG. 3, the camera system 100 includes:
[0038] An ultra-high-definition image sensor 104 for capturing an image frame.
[0039] In the present application, the resolution of the image sensor 104 is at least 4K (3840x2160 pixels), and the image sensor 104 can flexibly adjust its effective pixel area to support different aspect ratio outputs, including but not limited to standard 4:3, 16:9 and user-defined aspect ratios such as 1:1, etc., so that the image sensor has the ability to cover all the required resolution pixels of the output mode templates.
[0040] A user interface 105 for receiving user input.
[0041] In the present application, the user interface 105 is a touch screen and supports multiple input methods such as mouse clicks, keyboard input and touch control.
[0042] A memory 102 connected to the processor 103 for storing program instructions.
[0043] In the present application, the memory not only stores the program instructions required for system operation, but also is responsible for storing a series of preset output mode templates. These output mode templates provide users with convenient choices, so that the camera can flexibly adapt to the needs of different scenes and display devices.
[0044] Specifically, a plurality of output mode templates are preset in the memory, including but not limited to standard HD (such as 720p, 1080p), 4K (3840x2160 pixels), 8K (7680x4320 pixels) resolution modes, and two basic picture orientation modes of landscape and portrait, which cover the mainstream display resolution and picture orientation requirements and provide a plug-and-play solution for users.
[0045] The user can intuitively browse and select the required output mode template through the user interface of the camera. When the user selects a certain output mode template, the system automatically generates the corresponding picture output instruction according to the resolution, frame ratio and other parameters defined in the output mode template, and sends it to the processor for execution. This process greatly simplifies the user's operation process and improves the usability of the system.
[0046] In addition, the output mode templates in the memory also support further custom settings. Users can adjust the width-height ratio and shape of the frame according to their own needs to create unique visual effects. These custom settings are also saved in the memory for subsequent selection and use by the user.
[0047] The network communication module 101 is configured to simultaneously transmit at least two sub-pictures including a landscape sub-picture and a portrait sub-picture to corresponding display devices in real time to meet the display ratio of the corresponding display devices.
[0048] In this embodiment, the network communication module 101 supports multiple protocols such as TCP / IP and HTTP to achieve seamless connection with remote devices. The network communication module can transmit image data including landscape sub-pictures and portrait sub-pictures to display devices in real time, and simultaneously receive and execute picture output instructions from the remote.
[0049] The processor 103 is connected with the image sensor 104 and is configured to process the image picture.
[0050] The camera system 100 further comprises a target detection module configured in the processor.
[0051] The target detection module is a key component of the camera system 100 responsible for identifying and tracking targets in the image picture. It can be divided into two main parts: a target positioning unit and a target tracking unit.
[0052] Specifically, the target positioning unit is configured to determine the target in the current image picture and obtain the position of the target.
[0053] The target positioning unit is based on deep learning technology, including but not limited to convolutional neural network (CNN), and the target positioning unit finds the targets of interest in each image frame and gives the accurate positions of the targets in the form of a bounding box.
[0054] The target tracking unit is used to track the targets in consecutive image frames and update the positions of the targets.
[0055] In the embodiment, the target tracking unit uses the target information identified in the current frame (i.e., the current image frame) to continuously track the targets in subsequent frames. The tracking algorithm can be, but is not limited to, Kalman filtering, particle filtering, optical flow method, etc.
[0056] The tracking algorithm keeps tracking the targets in consecutive image frames, so that the tracking of the targets can be completed even if the targets are occluded, deformed, or moved quickly during the movement.
[0057] Referring to FIG. 2, the program instructions, when executed, cause the processor 103 to perform the following steps:
[0058] S1: determining a target in the image frame and obtaining the position of the target.
[0059] S2: receiving a frame output instruction by the processor, wherein the frame output instruction includes at least two output modes of horizontal screen and vertical screen, and frame cutting position and size.
[0060] S3: according to the at least two output modes and the frame cutting position and size, cutting at least two sub-frames including a horizontal screen sub-frame and a vertical screen sub-frame according to the position of the target in the image frame captured by the image sensor and outputting, wherein the at least two sub-frames include a horizontal screen sub-frame and a vertical screen sub-frame to meet the display ratio requirements of different display devices.
[0061] More specifically, after the system is started, the processor 103 waits for the user to input a frame ratio instruction. The frame ratio instruction is generated by a user interface configured by the camera system 100 of the present application. The user can directly view the real-time image on the display screen of the camera body and select the required frame output mode and set the frame cutting position through the user interface. Specifically, the user interface is provided with clear menu options and a sliding bar / selector, allowing the user to easily switch the output mode and accurately adjust the cutting position. When the user completes the setting, the frame output instruction is sent to the processor for execution by clicking the confirmation button.
[0062] In addition to user interface input, the camera system also supports receiving picture output instructions through network communication. The network communication module 101 can transmit data with other network devices, including but not limited to smartphones, tablets, computers, or remote servers. Users can remotely access the camera system 100 and send picture output instructions through the supporting mobile phone APP, web-based control platform, or dedicated software client. After these instructions are transmitted to the camera system 100 through the network, they are received by the network communication module 101 and forwarded to the processor for processing.
[0063] In actual application, in order to provide more convenient and efficient operation experience, the camera system 100 also supports the combination of user interface input and network communication reception. For example, when the user is near the camera, he or she can directly operate through the user interface 105; when the user is at a remote location, he or she can send instructions through network communication. In addition, the user can also set the output mode and cutting position in advance through the user interface 105 and save them as preset configurations. When needed, the user can remotely call these preset configurations through network communication to achieve quick switching and output.
[0064] When the processor 103 receives a picture output instruction containing at least two output modes (landscape and portrait) and specific picture cutting positions, the processor 103 will first decode and analyze the picture output instruction in detail to ensure accurate extraction of the display requirements of landscape and portrait and how the picture should be accurately cut. Then, the processor will strictly verify these parameters to confirm their correctness and compatibility, avoiding problems in subsequent processing due to parameter errors.
[0065] Next, the processor 103 will process the real-time image captured by the image sensor 104 according to the analyzed output mode and picture cutting position information. Specifically, it will mark the areas that need to be cut on the image, which correspond to the display requirements of landscape and portrait, and perform optimization processing such as cropping, scaling, and color correction on these areas to ensure that the final generated sub-pictures meet the display ratio requirements and have excellent visual effects. After processing is completed, the processor will transmit these sub-picture data to the corresponding display device in real time through the network communication module or other output interfaces. During the entire transmission process, the processor will also perform necessary synchronization operations with the display device to ensure that the picture can be accurately displayed on the screen. At the same time, the processor will also pay attention to feedback information from the display device to adjust transmission parameters or processing strategies in a timely manner, ensuring the stable operation of the entire system.
[0066] In addition, the camera system 100 also includes an output interface supporting multiple output modes, such as IP, SDI, HDMI, etc., to adapt to different application scenarios, and in each sub-picture output, any one output mode is selected for transmission.
[0067] The working principle of the multi-aspect output ultra-high-definition camera system provided by the present application is as follows: after the system is started, the image sensor 104 captures a high-resolution image picture, at least reaching the 4K standard (3840x2160 pixels), the user selects the required output mode and picture cutting position through the touch screen user interface 105, and the processor 103 receives and processes these information. The user interface supports intuitive browsing and selection of preset output mode templates, which are stored in the memory 102, including common HD, 4K and 8K resolution modes, and basic direction modes of landscape and portrait screens. The user can also customize the aspect ratio of the output mode to meet special needs.
[0068] Once the user sets the output mode and cutting position, the processor 103 determines the target on the image picture according to the instructions and obtains the position of the target. Then, according to the requirements of the output mode and the cutting position, the processor cuts at least two sub-pictures (i.e. landscape and portrait sub-pictures) from the image picture, and these sub-pictures are processed as necessary, such as cropping, scaling and color correction, to ensure that they meet the display ratio requirements of the display device and have good visual effects.
[0069] The network communication module 101 is responsible for transmitting the processed sub-pictures to the corresponding display device in real time, ensuring seamless connection with the remote device. In addition, the system also supports receiving picture output instructions through the network, and the user can remotely send instructions through a mobile phone application, a web-based control platform or a dedicated software client.
[0070] During the entire process, the processor 103 is also responsible for coordinating the synchronization operation with the display device and adjusting the transmission parameters according to the feedback information to ensure stable and efficient image output. The output interface supports multiple formats, such as IP, SDI, HDMI, etc., to adapt to different application scenarios. Embodiment
[0071] In this embodiment, the picture output instruction only contains the requirements of landscape and portrait screens.
[0072] For example, when the resolution of the image sensor 104 is 8Kx8K, the maximum resolution of the horizontal cutting can be set to 7680x4320, and the minimum resolution can be set to 1920x1080; the maximum resolution of the vertical cutting can be set to 4320x7680, and the minimum resolution can be set to 1080x1920.
[0073] When the resolution of the image sensor 104 is 8Kx5K (8192x5556), the maximum resolution of the horizontal cut can be set to 8192x4608, and the minimum resolution can be set to 1920x1080; the maximum resolution of the vertical cut can be set to 3126x5556, and the minimum resolution can be set to 1080x1920.
[0074] Target detection and tracking: Referring to FIG. 4, the target detection module built-in the camera can automatically identify and track the target, determine the center position coordinates of the target (as shown in FIG. 5), and if the target moves, the cut sub-picture will move accordingly, always keeping the target in the cut sub-picture.
[0075] Output configuration: The user can select the output method (such as IP, SDI, HDMI, etc.) through the user interface 105, and the cut sub-picture corresponding to the horizontal screen and the cut sub-picture corresponding to the vertical screen can be output through any output method to adapt to different display device requirements. Specifically, the sub-picture corresponding to the horizontal screen can be used for television broadcast, and the sub-picture corresponding to the vertical screen can be used for live streaming on social media platforms.
[0076] In this way, the camera system 100 not only meets the traditional horizontal screen demand, but also flexibly responds to the emerging vertical screen application trend, thereby providing users with more diversified shooting and viewing experiences.
[0077] The present application is described with reference to flowcharts and / or block diagrams according to the method, device (system) and computer program product of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0078] Those skilled in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by instructing the relevant hardware by means of a program, and the program can be stored in a computer readable storage medium, including Read-Only Memory (ROM), Random Access Memory (RAM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), One-time Programmable Read-Only Memory (OTPROM), Electrically-Erasable Programmable Read-Only Memory (EEPROM), Compact Disc Read-Only Memory (CD-ROM) or other optical disk storage, magnetic disk storage, magnetic tape storage, or any other medium that can be used to carry or store data in a computer readable manner.
[0079] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.
Claims
1. A multi-aspect output ultra-high definition camera system, comprising: an ultra-high definition image sensor configured to capture an image frame; a processor connected to the image sensor and configured to process the image frame; a memory connected to the processor and configured to store program instructions; wherein the program instructions, when executed, cause the processor to perform the following steps: determining a target in the image frame and obtaining a position of the target; receiving a frame output instruction via the processor, wherein the frame output instruction comprises at least two output modes of a horizontal screen and a vertical screen, and a frame cutting position and size; cutting at least two sub-frames including a horizontal screen sub-frame and a vertical screen sub-frame according to the position of the target in the image frame captured by the image sensor and outputting the at least two sub-frames according to the at least two output modes and the frame cutting position and size, wherein the at least two sub-frames include the horizontal screen sub-frame and the vertical screen sub-frame to meet display ratio requirements of different display devices; simultaneously moving and scaling the position and size of the sub-frames in the image frame captured by the image sensor according to the target identified by the target detection module.
2. The multi-aspect output ultra-high definition camera system according to claim 1, wherein, The memory is further configured to: store preset output mode templates, wherein the output mode templates include an HD template, a 4K template, an 8K template, a horizontal screen template, and a vertical screen template, and a user can select an output mode template to generate the frame output instruction.
3. A multi-aspect output ultra-high definition camera system according to claim 2, wherein, The output mode templates further include templates for customizing an aspect ratio and an aspect shape.
4. The multi-aspect output ultra-high definition camera system according to claim 3, wherein, The image sensor has a resolution of at least 4K.
5. The multi-aspect output ultra-high definition camera system according to claim 4, wherein, The processor receives the frame output instruction in a manner including one or a combination of user interface input and network communication reception.
6. The multi-aspect output ultra-high definition camera system according to claim 5, wherein, The position of the target is obtained in a manner including automatic obtaining by using a preset target detection module and / or manual designation based on a configured user interface.
7. A multi-aspect output ultra-high definition camera system according to claim 6, wherein, The system further comprises a network communication module configured to simultaneously transmit the at least two sub-frames including the horizontal screen sub-frame and the vertical screen sub-frame to corresponding display devices in real time to meet display ratios of the corresponding display devices.
8. The multi-aspect output ultra-high definition camera system according to claim 7, wherein, The target detection module comprises: a target positioning unit configured to determine a target in a current image frame and obtain a position of the target; and a target tracking unit configured to track the target in consecutive image frames and update the position of the target.
Citation Information
Patent Citations
Video image segmentation method and device
CN108986117A
Method for simultaneously producing short videos with different breadth ratios
CN110418162A
Optimal view angle video processing system and method based on 8K video signal and AI technology
CN113055604A
Virtual multi-camera application method and system
CN116320214A
Image stitching method and apparatus for multi-camera device, storage medium, and terminal
WO2022111330A1