Image processing system and method for processing a video stream
By generating dual video streams through an image processing system, the problem of image distortion from wide-angle or fisheye cameras is solved, enabling flexible image processing and recognition detection, and making it suitable for a variety of video products.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GENESYS LOGIC INC
- Filing Date
- 2022-06-17
- Publication Date
- 2026-05-05
AI Technical Summary
Images captured by existing wide-angle or fisheye cameras are prone to distortion, making the content difficult to identify and causing discomfort to users. Furthermore, the back-end processing devices are highly complex, limiting the application scenarios.
The image processing system generates two video streams, one for detection and the other for display, including deformation correction, recognition detection, and parameter adjustment. These processes process the image data separately to meet different application requirements.
It achieves effective image deformation correction and recognition detection, improving the system's flexibility and applicability, and is suitable for a variety of video-related products.
Smart Images

Figure CN115564660B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an image processing technology, and more particularly, to an image processing system and a method for processing video streams. Background Technology
[0002] In conventional technology, while cameras equipped with wide-angle or fisheye lenses can capture images with a wider field of view (FoV), the edges of the images may be curved, resulting in an unnatural appearance. Distortion in wide-angle or fisheye images can make their content difficult to discern and may even cause eye discomfort to the user.
[0003] Conventional fisheye or wide-angle cameras transmit panoramic images to a powerful back-end system (such as a host computer) to process de-warping, detection, and the display of the dewarped image. However, the back-end devices are complex, and the application scenarios for these cameras are limited. Summary of the Invention
[0004] This invention relates to an image processing system and a video stream processing method that can generate dual video streams for detection and display applications, respectively.
[0005] According to embodiments of the present invention, a video stream processing method includes (but is not limited to) the following steps: obtaining a first image based on parameters; performing a deformation correction procedure on the first image and generating a second image; performing a recognition detection procedure on the second image and generating a detection result; generating control information based on the detection result; adjusting parameters based on the control information and generating a third image; and outputting the second and third images.
[0006] According to an embodiment of the present invention, the image processing system includes (but is not limited to) a controller. The controller is configured to acquire a first image based on parameters, perform a deformation correction procedure on the first image to generate a second image, adjust parameters according to control information to generate a third image, and output the second and third images. The control information is generated based on the detection results produced by a recognition and detection procedure on the second image.
[0007] Based on the above, according to the image processing system and video stream processing method of the present invention, the deformation correction procedure can generate two images respectively for two video streams to carry, and use them for different backend applications. This allows for effective division of labor and improves the flexibility of the architecture. Attached Figure Description
[0008] The accompanying drawings are included to further illustrate the invention, and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0009] Figure 1 This is a block diagram of components of an image processing system according to an embodiment of the present invention;
[0010] Figure 2 This is a block diagram of components of an image processing system according to another embodiment of the present invention;
[0011] Figure 3 This is a flowchart of a video stream processing method according to an embodiment of the present invention;
[0012] Figure 4 This is an example illustrating multi-window preview;
[0013] Figure 5 This is a flowchart of a dual video stream application according to an embodiment of the present invention.
[0014] Explanation of icon numbers
[0015] 1, 2: Image processing system;
[0016] 10: Image capture device;
[0017] 11: Lens;
[0018] 15: Image sensor;
[0019] 30, 30': Controller;
[0020] 31: Memory;
[0021] 39: Arithmetic unit;
[0022] 32, 33, 34, 35: Transmission interfaces;
[0023] 36, 37, 38: Endpoints;
[0024] 50, 70: Back-end devices;
[0025] 60: Display device;
[0026] 71: Processor;
[0027] 72: Image unit;
[0028] VS1: First video stream;
[0029] VS2: Second video stream;
[0030] CI: Control Information;
[0031] S310~S360, S510~S550: Steps;
[0032] TW1~TW3: Windows;
[0033] SIM: First Image;
[0034] TIM: Third Image. Detailed Implementation
[0035] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same component symbols are used in the drawings and description to denote the same or similar parts.
[0036] Figure 1 This is a block diagram of components of an image processing system 1 according to an embodiment of the present invention. Please refer to... Figure 1 The image processing system 1 includes (but is not limited to) an image capture device 10, a controller 30, a back-end device 50, and a display device 60.
[0037] Image capture device 10 may be a camera, video camera, monitor, or similar device. Image capture device 10 may include (but is not limited to) lens 11 and image sensor 15 (e.g., charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS)). In one embodiment, an image can be captured by lens 11 and image sensor 15. For example, light passes through lens 11 and is imaged onto image sensor 15.
[0038] In some embodiments, the specifications (e.g., imaging aperture, magnification, focal length, imaging viewing angle, size of image sensor 15, etc.) and the number of image capture devices 10 can be adjusted according to actual needs. For example, lens 11 is a fisheye or wide-angle lens, thereby generating fisheye or wide-angle images.
[0039] The controller 30 can be coupled to the image capture device 10 via a camera interface, I2C, and / or other transmission interfaces. The controller 30 includes (but is not limited to) a memory 31, transmission interfaces 32, 33, and 34, and a processing unit 39.
[0040] The memory 31 can be any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), or similar component. In one embodiment, the memory 31 is used to store program code, software modules, configuration settings, data, or files.
[0041] Transmission interfaces 32, 33, and 34 can be camera interfaces, I2C, or other transmission interfaces. For example, transmission interface 32 is an I2C interface, transmission interface 33 is a Mobile Industry Processor Interface (MIPI), and transmission interface 34 is a Universal Serial Bus (USB) interface. Alternatively, transmission interface 33 can be a Digital Parallel Bus (DPS) interface, a Low-Voltage Differential Signalling (LVDS) interface, a High-Speed Serial Pixel (HiSPi) interface, or a V-by-One interface. In one embodiment, transmission interfaces 32, 33, and 34 are used to connect to external devices. For example, transmission interfaces 32 and 33 are connected to backend device 50, and transmission interface 34 is connected to display device 60.
[0042] The processing unit 39 is coupled to the memory 31 and the transmission interfaces 32, 33, and 34. The processing unit 39 may be a graphics processing unit (GPU), or other programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or other similar components or combinations thereof. In one embodiment, the processing unit 39 is used to execute all or part of the operations of the controller 30, and can load and execute the program code, software modules, files, and data stored in the memory 31.
[0043] The backend device 50 and display device 60 may be a desktop computer, laptop computer, server, smartphone, or tablet computer. In one embodiment, the backend device 50 is used for image calculation and / or detection. In one embodiment, the display device 60 is used for displaying images. In some embodiments, the functionality of the backend device 50 and / or display device 60 may be implemented on a microcontroller, chip, SoC, or system.
[0044] Figure 2 This is a block diagram of components of an image processing system 2 according to another embodiment of the present invention. Please refer to... Figure 2 In this embodiment, the controller 30' and Figure 1 The controller 30 differs from the controller 30 in that it includes a transmission interface 35, and the processing unit 39 is coupled to the transmission interface 35. The transmission interface 35 may be a USB interface and provides endpoints 36, 37, and 38. It is worth mentioning that in this embodiment, endpoints 36, 37, and 38 are used to connect to the back-end device 70, wherein endpoints 36 and 37 are used to connect to the processor 71 of the back-end device 70, and endpoint 38 is used to connect to the image unit 72 of the back-end device 70.
[0045] The back-end device 70 can be a desktop computer, laptop computer, server, smartphone, or tablet computer. The processor 71 can be a graphics processor or GPU, or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), programmable controller, field-programmable gate array, application-specific integrated circuit, or other similar components or combinations thereof. The image unit 72 can be processing or control circuitry for image output / display / playback.
[0046] The methods described in the embodiments of the present invention will be described below in conjunction with the various devices, components, and modules in image processing system 1 or image processing system 2. The various processes of this method may be adjusted according to the implementation situation, and are not limited thereto.
[0047] Figure 3 This is a flowchart of a video stream processing method according to an embodiment of the present invention. Please refer to... Figure 3The controllers 30 and 30' acquire the image from the first image according to the parameters (step S310). The parameters include setting parameters for de-wrapping, panning, tilting, zooming, rotation, shifting, and / or field of view (FoV) adjustment. In one embodiment, the parameters are setting parameters used by the image capture device 10 or other external image capture device to perform image capturing operations. For example, the controllers 30 and 30' may send instructions or messages to the image capture device 10 or other external image capture device, and these instructions or messages are related to the setting parameters. For example, changing the image capturing range. The image capture device 10 or other external image capture device can perform image capturing operations according to the setting parameters to acquire an image.
[0048] In one embodiment, controllers 30 and 30' acquire a first image from image capture device 10. Specifically, the first image is an image captured by image capture device 10 or other external image capture device of one or more targets. In one embodiment, the target is a human body. In some embodiments, the first image is of the upper body of a human (e.g., waist, shoulders, or above the chest). In other embodiments, the target may be various types of living or non-living objects. Controllers 30 and 30' can acquire the first image captured by image capture device 10 via a camera interface and / or I2C.
[0049] Controllers 30 and 30' perform a deformation correction procedure on the first image and generate a second image (step S330). The deformation correction procedure is used to adjust the deformation in the first image. In one embodiment, the deformation correction procedure may be anti-twist, left and right rotation, tilt, scaling, rotation, translation, and / or viewpoint adjustment.
[0050] In one embodiment, the image sensor 15 may output raw data of a first image (e.g., the sensing intensity of multiple primary colors) to the controllers 30, 30'. The controllers 30, 30' may process the raw data according to an image signal processing (ISP) program and output a visible full-color image to a distortion correction program.
[0051] In another embodiment, image sensor 15 may output lightness-chroma-saturation (YUV) or other color-coded data of the first image to controllers 30, 30'. Controllers 30, 30' may ignore or disable the ISP program for processing this color-coded data (i.e., skip the ISP program for the color-coded data) and directly perform a deformation correction procedure on the color-coded data.
[0052] In one embodiment, the first image is a wide-angle or fisheye image, and the controllers 30, 30' can generate a dewarped panoramic image (i.e., the second image) through a deformation correction procedure. This panoramic image can be used for subsequent detection of one or more targets. For example, the panoramic image can be used to detect faces, gestures, or body parts. In another embodiment, the controllers 30, 30' can rotate the viewpoint of the first image along any axis, zoom in or out over all or part of the area, translate in or out over all or part of the area, and / or adjust the viewpoint size to generate the second image.
[0053] In one embodiment, controller 30 may transmit the second image to back-end device 50 via transmission interface 33, while controller 30' may transmit the second image to processor 71 of back-end device 70 via endpoint 37. For example, controllers 30 and 30' may output the second image via a first video stream VS1.
[0054] The backend device 50 or processor 71 performs a recognition and detection procedure (or object detection procedure) on the second image and generates a detection result (step S330). The recognition and detection procedure, for example, determines one or more regions of interest (RoI) (or bounding boxes, or bounding rectangles) or pinots (which may be located on the outline, center, or any position on the target object) in the second image corresponding to the target object (e.g., a person, animal, inanimate object, or part of it), and then identifies the type of the target object (e.g., human, male or female, dog or cat, table or chair, etc.).
[0055] In one embodiment, the detection result is related to one or more regions of interest (RoIs) (or bounding boxes, or bounding rectangles) in the second image. For example, the detection program can identify regions of interest in the second image that correspond to a target object, and these regions of interest can encompass all or part of the target object.
[0056] The backend device 50 or processor 71 generates control information CI based on the detection results (step S340). The control information is related to the detection results of the second image. The detection results are, for example, identifying one or more faces, gestures, body parts, or other objects in the second image (i.e., object detection), or determining the position / motion of one or more objects (i.e., object tracking). In one embodiment, the control information is based on the detection results and adjusts parameters for a deformation correction procedure. For example, the position of a face in the image is used to set the viewpoint turning angle in the deformation correction procedure. As another example, the size of a face in the image is used to set the region magnification in the deformation correction procedure.
[0057] In one embodiment, the back-end device 50 may transmit control information CI to the controller 30 via the transmission interface 32, while the processor 71 may transmit control information CI to the controller 30' via the endpoint 36.
[0058] Controllers 30 and 30' adjust parameters according to control information and generate a third image (step S350). For example, controllers 30 and 30' generate instructions or messages related to the setting parameters of the image capture device 10 according to control information, and accordingly capture the first image again through the image capture device 10. In one embodiment, controllers 30 and 30' obtain the recaptured first image, adjust the first image according to control information CI through a deformation correction procedure, and generate a third image.
[0059] In one embodiment, the third image comprises one or more windows. That is, the third image is divided into one or more windows. Controllers 30, 30' can assign one or more regions of interest to these windows. For example, controllers 30, 30' can scale the regions of interest in the panoramic image to fit the size of a specified window and shift the scaled regions of interest (by means of left and right rotation, tilt, rotation, translation, and / or viewpoint adjustment) to this window.
[0060] For example, Figure 4 This is an example illustrating multi-window preview. Please refer to it. Figure 4 The controllers 30 and 30' divide the third image TIM (as shown below) into three windows TW1, TW2, and TW3. Window TW3 is a capture of four figures from the first image SIM (as shown above), while windows TW1 and TW2 are each captured from only one figure from the first image SIM.
[0061] In one embodiment, controllers 30 and 30' can simultaneously or time-divisionally output the second and third images through a deformation correction procedure. For example, controllers 30 and 30' can generate a first third image based on the detection result of the first second image. Then, controllers 30 and 30' can output the second and third images (step S360). For example, controllers 30 and 30' can generate the second and third images sequentially through time-division multitasking. Furthermore, after generating the first second image, controllers 30 and 30' then generate the third image based on the detection result of this second image. In addition, the controllers 30 and 30' of the present invention can simultaneously or time-divisionally output the second and third images, and no restrictions are placed on the timing of the output herein.
[0062] In one embodiment, controllers 30 and 30' can output a second image via a first video stream VS1 and a third image via a second video stream VS2. Specifically, this second video stream is different from the first video stream. The first video stream VS1 carries the second image, and the second video stream VS2 carries the third image. That is, controllers 30 and 30' can provide dual video stream output. Furthermore, depending on the instruction cycle, these two video streams can be output to an external device simultaneously or at different times.
[0063] In one embodiment, with Figure 1 Taking the architecture as an example, the processing unit 39 outputs a first video stream VS1 through the transmission interface 33, outputs a second video stream VS2 through the transmission interface 34, and inputs control information CI through the transmission interface 32. For example, MIPI outputs the first video stream VS1, one endpoint of the USB interface outputs the second video stream VS2, and the I2C interface inputs control information CI.
[0064] In another embodiment, with Figure 2 Taking the architecture as an example, the processing unit 39 outputs a first video stream VS1 through endpoint 37 of the transmission interface 35, outputs a second video stream VS2 through endpoint 38, and inputs control information CI through endpoint 36. For example, the two endpoints of the USB interface output the first video stream VS1 and the second video stream VS2, and the other endpoint inputs control information CI.
[0065] The applications of the first video stream VS1 and the second video stream VS2 may differ. Figure 5 This is a flowchart illustrating the application of dual video streams according to an embodiment of the present invention. Please refer to it as well. Figure 1 and Figure 5In one embodiment, a first video stream VS1 is output to a backend device 50, and the backend device 50 performs an identification and detection procedure on the first video stream VS1 (i.e., calculation and detection of the second image) (step S510) to generate a detection result. This identification and detection procedure can be the aforementioned object detection and / or tracking. Furthermore, the backend device 50 can generate control information CI based on the detection result. It should be noted that the detection result and control information CI can be referred to in the descriptions of steps S330 and S340, and will not be repeated here. The backend device 50 can further feed back the control information CI to the controllers 30 and 30' (step S530), enabling the controllers 30 and 30' to generate a third image based on the control information CI. Similarly, in Figure 2 In the implementation architecture, the processor 71 of the back-end device 70 further feeds back control information CI to the controller 30' (step S530), so that the controller 30' can generate a third image according to the control information CI.
[0066] In one embodiment, the second video stream VS2 is output to the display device 60, and the display device 60 displays the third image (i.e., image preview) carried by the second video stream VS2 through a display program (step S550). Figure 4 For example, the third image (TIM) is used for multi-window preview and display. It should be noted that the isochronous and bulk modes provided by USB's UVC (USB video device class or UVC) allow the display device 60 to display the third image in real time. However, other real-time video output protocols can still be used in embodiments of the present invention.
[0067] In summary, the image processing system and video stream processing method of this invention provide dual video streaming. One video stream can be used for the detection of faces, gestures, or body parts, and the other video stream can be used to control the display of dewarped multiple regions of interest images.
[0068] This invention provides a flexible architecture with wide applicability (e.g., it can be applied to various types of video-related products). The system of this invention allows for efficient division of labor. For example, the backend or computer system can be responsible for detecting faces, gestures, or body parts, while the display device or system can be responsible for displaying images of multiple regions of interest without distortion. In this way, manufacturers of conventional cameras can easily upgrade their products to fisheye or wide-angle cameras, providing multi-target detection and tracking capabilities such as faces, gestures, or body parts.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing a video stream, characterized in that, include: Obtain the first image based on the parameters; The first image is subjected to a deformation correction procedure, and a second image is generated; The second image is subjected to a recognition and detection procedure, and the detection results are generated; Control information is generated based on the detection results; Adjust the parameters according to the control information; Based on the adjusted parameters, another first image is captured; According to the control information, the deformation correction procedure is used to adjust the other first image and generate a third image; and Output the second image and the third image.
2. The video stream processing method according to claim 1, characterized in that, include: The first image is acquired from the image capture device; The first image is subjected to the deformation correction procedure according to the image signal processing procedure to generate the second image, wherein the deformation correction procedure is used to adjust the deformation in the first image; The third image is generated according to the control information through the deformation correction procedure; as well as Simultaneously outputting the second image and the third image, wherein the second image is output through a first video stream and the third image is output through a second video stream, wherein the second video stream is different from the first video stream.
3. The video stream processing method according to claim 2, characterized in that, Also includes: The deformation correction procedure simultaneously generates the second image and the third image.
4. The video stream processing method according to claim 2, characterized in that, Also includes: The first video stream is subjected to the recognition and detection procedure by a backend device to generate the detection result, wherein the detection result is related to at least one region of interest in the second image.
5. The video stream processing method according to claim 4, characterized in that, Also includes: Arrange the at least one region of interest into at least one window, wherein the third image includes the at least one window; as well as The third image is displayed via a display device.
6. The video stream processing method according to claim 2, characterized in that, Outputting the second image through the first video stream and outputting the third image through the second video stream includes: The first video stream is output through a mobile industry processor interface, and the second video stream is output through a universal serial bus.
7. The video stream processing method according to claim 2, characterized in that, Outputting the second image through the first video stream and outputting the third image through the second video stream includes: The first video stream and the second video stream are output from the two endpoints of the Universal Serial Bus interface, respectively.
8. The video stream processing method according to claim 7, characterized in that, Also includes: The control information is input through the other end of the universal serial bus interface.
9. The video stream processing method according to claim 2, characterized in that, The deformation correction procedure includes at least one of anti-torsion, left and right rotation, tilt, scaling, rotation, translation and view adjustment, and the parameters include setting parameters for at least one of anti-torsion, left and right rotation, tilt, scaling, rotation, translation and view adjustment.
10. The video stream processing method according to claim 1, characterized in that, The first image is a wide-angle image or a fisheye image, and the second image is an anti-distortion panoramic image.
11. An image processing system, characterized in that, include: The controller is configured to: Obtain the first image based on the parameters; The first image is subjected to a deformation correction procedure, and a second image is generated; Adjust the parameters according to the control information; Based on the adjusted parameters, another first image is captured; According to the control information, the deformation correction procedure adjusts the other first image and generates a third image, wherein the control information is generated based on the detection result of the recognition and detection procedure applied to the second image; and Output the second image and the third image.
12. The image processing system according to claim 11, characterized in that, Also includes: An image capture device, coupled to the controller, includes a lens and an image sensor, and is configured to capture a first image through the lens and the image sensor, wherein the controller is further configured to: The first image is acquired from the image capture device; The first image is subjected to a deformation correction procedure according to the image signal processing procedure to generate the second image, wherein the deformation correction procedure is used to adjust the deformation in the first image; The third image is generated according to the control information through the deformation correction procedure; as well as Simultaneously outputting the second image and the third image, wherein the second image is output through a first video stream and the third image is output through a second video stream, wherein the second video stream is different from the first video stream.
13. The image processing system according to claim 12, characterized in that, The controller is also used for: The deformation correction procedure simultaneously generates the second image and the third image.
14. The image processing system according to claim 12, characterized in that, Also includes: A back-end device, coupled to the controller, is used to perform the recognition and detection procedure on the first video stream to generate the detection result, wherein the detection result is related to at least one region of interest in the second image.
15. The image processing system according to claim 14, characterized in that, The detection result is related to at least one region of interest in the second image, and the image processing system further includes: a display device coupled to the controller, wherein the controller is further configured to: Arrange the at least one region of interest into at least one window, wherein the third image includes the at least one window; and The third image is displayed through the display device.
16. The image processing system according to claim 12, characterized in that, The controller includes: A mobile industry processor interface for outputting the first video stream; and A universal serial bus is used to output the second video stream.
17. The image processing system according to claim 12, characterized in that, The controller includes: A universal serial bus interface includes two endpoints for outputting the first video stream and the second video stream, respectively.
18. The image processing system according to claim 17, characterized in that, The universal serial bus interface also includes: The other endpoint is used to input the control information.
19. The image processing system according to claim 12, characterized in that, The deformation correction procedure includes at least one of anti-torsion, left and right rotation, tilt, scaling, rotation, translation and view adjustment, and the parameters include setting parameters for at least one of anti-torsion, left and right rotation, tilt, scaling, rotation, translation and view adjustment.
20. The image processing system according to claim 11, characterized in that, The first image is a wide-angle image or a fisheye image, and the second image is an anti-distortion panoramic image.
Citation Information
Patent Citations
Method, apparatus and system for image processing
US20160112652A1
Method and apparatus for processing wide angle image
US20180089795A1