Region of interest based video encoding method, apparatus, device and medium

By configuring detection sensors and lens imaging principles to determine the location of the region of interest, and employing differential coding measures, the problem of blurry face recordings in video door lock devices was solved, achieving low-power video transmission.

CN118075474BActive Publication Date: 2026-01-02ZHEJIANG UNIVIEW TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211485872.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2026-01-02
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

Existing video door lock devices suffer from blurry face recordings due to reduced video frame encoding quality during short-range detection, and AI algorithms for analyzing face positions increase power consumption, making them unsuitable for battery-powered devices.

Method used

By configuring detection sensors to acquire distance and angle information of the target object, and combining the lens imaging principle and camera attribute information, the position of the region of interest in the image is determined, and different encoding methods are used to encode the region of interest and non-regions of interest.

Benefits of technology

While ensuring clarity in the region of interest, it reduces video size and transmission power consumption, making it suitable for low-power video devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118075474B_ABST
    Figure CN118075474B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a video coding method and device based on a region of interest, equipment and a medium. The method comprises: obtaining distance information and angle information of a target object relative to a detection sensor based on the detection sensor; determining imaging position information of the target object on an imaging sensor of a camera based on a lens imaging principle, the distance information, the angle information and camera attribute information; determining final position information of the region of interest on an image based on a region of interest feature, the distance information, the angle information, the imaging position information, imaging size information of the imaging sensor and image resolution information; and performing ROI coding on image frames in a to-be-coded video according to the final position information. The present scheme can assist a low-power video device in identifying a region of interest through a detection sensor, and can take different coding measures on the region of interest and a non-region of interest, thereby ensuring the definition of the region of interest while reducing the video size and video transmission power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, in particular to a video coding method based on a region of interest, an apparatus, a device and a medium. BACKGROUND

[0002] With the rapid development of Internet of Things technology, video door locks, video cat eyes and other devices are becoming more and more popular. Currently, video door locks usually use lithium batteries or dry batteries for power supply, and the power consumption of the video module is relatively large, so the video module is usually in a dormant state. An external detection sensor is used to detect whether there is a person in front of the door lock, and if there is a person, the video module is awakened to record the video. At the same time, since the video door lock usually uses wifi or operator network for video data transmission, it is required that the video captured by the video module is as small as possible. At this time, in order to reduce the video size, it is often necessary to reduce the frame rate and reduce the encoding quality of the video frame, resulting in blurred and unclear face recording.

[0003] In related technologies, a region of interest (ROI) is circled in a video frame by manually setting a static position. However, this solution is only suitable for long-distance monitoring with a large field of view, and for short-distance detection such as door locks, the face position cannot be accurately determined by manually setting a static position, so the problem of blurred and unclear face recording caused by reducing the encoding quality of the video frame in short-distance detection cannot be solved.

[0004] In related technologies, a video AI algorithm is used to analyze the position of a face in a video. In this solution, the AI algorithm analysis requires hardware with computing power, and at the same time, it greatly increases the power consumption, which is not suitable for door lock devices powered by batteries. SUMMARY

[0005] The present application provides a video coding method based on a region of interest, an apparatus, a device and a medium, which can assist low-power video devices in video region of interest positioning through a detection sensor, and take different coding measures for the region of interest and the non-region of interest, so as to reduce the video size and video transmission power consumption while ensuring the clarity of the region of interest.

[0006] According to an aspect of the present application, a video coding method based on a region of interest is provided, a detection sensor is configured for a camera for collecting a video to be coded, the detection sensor has the ability to collect distance and angle information of a target object in the field of view of the camera, and the method comprises:

[0007] obtaining distance information and angle information of the target object relative to the detection sensor based on the detection sensor;

[0008] determine, based on the lens imaging principle, the imaging position information of the target object on the imaging sensor of the camera according to the distance information, the angle information and camera attribute information;

[0009] determine, based on the region of interest feature, the final position information of the region of interest on the image according to the distance information, the angle information, the imaging position information, imaging size information of the imaging sensor and image resolution information;

[0010] perform ROI encoding on the image frames in the video to be encoded according to the final position information.

[0011] According to another aspect of the present application, a video encoding device based on a region of interest is provided, which is configured with a detection sensor for a camera collecting a video to be encoded, the detection sensor having the ability to collect distance and angle information of a target object in the field of view of the camera, and the device comprises:

[0012] a detection sensor information determination module configured to acquire distance information and angle information of the target object relative to the detection sensor based on the detection sensor;

[0013] an imaging position information determination module configured to determine, based on the lens imaging principle, the imaging position information of the target object on the imaging sensor of the camera according to the distance information, the angle information and camera attribute information;

[0014] a final position information determination module configured to determine, based on the region of interest feature, the final position information of the region of interest on the image according to the distance information, the angle information, the imaging position information, imaging size information of the imaging sensor and image resolution information;

[0015] a ROI encoding module configured to perform ROI encoding on the image frames in the video to be encoded according to the final position information.

[0016] According to another aspect of the present application, a video encoding electronic device based on a region of interest is provided, which comprises:

[0017] at least one processor; and

[0018] a memory in communication connection with the at least one processor; wherein,

[0019] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the video encoding method based on a region of interest according to any one of the embodiments of the present application.

[0020] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for causing a processor to implement the region of interest based video encoding method according to any of the embodiments of the present application when executed.

[0021] The technical solution of the embodiment of the present application is based on a detection sensor to obtain distance information and angle information of a target object relative to the detection sensor; based on a lens imaging principle, to determine imaging position information of the target object on an imaging sensor of a camera according to the distance information, the angle information and camera attribute information; based on a region of interest feature, to determine final position information of the region of interest on an image according to the distance information, the angle information, the imaging position information, imaging size information of the imaging sensor and image resolution information; and to perform ROI encoding on image frames in a to-be-encoded video according to the final position information. The technical solution can assist a low-power video device to perform video region of interest positioning through a detection sensor, and can take different encoding measures on the region of interest and a non-region of interest, so as to ensure the definition of the region of interest while reducing the video size and video transmission power consumption.

[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0024] Figure 1 is a flow chart of a region of interest based video encoding method according to the first embodiment of the present application;

[0025] Figure 2 is a schematic diagram of the working principle of a detection sensor according to the first embodiment of the present application;

[0026] Figure 3 is a schematic diagram of a detection space coordinate system of a video door lock according to the first embodiment of the present application;

[0027] Figure 4A is a horizontal cross-sectional schematic diagram of a detection space coordinate system according to the first embodiment of the present application;

[0028] Figure 4BIt is a vertical profile diagram of detecting space coordinate system according to the embodiment one of the present application;

[0029] Figure 5A It is a horizontal profile diagram of lens imaging according to the embodiment one of the present application;

[0030] Figure 5B It is a vertical profile diagram of lens imaging according to the embodiment one of the present application;

[0031] Figure 6A It is an imaging diagram of the region of interest on the imaging sensor according to the embodiment one of the present application;

[0032] Figure 6B It is an imaging diagram of the region of interest on the image picture according to the embodiment one of the present application;

[0033] Figure 7 It is a structure diagram of the video coding device based on the region of interest according to the embodiment two of the present application;

[0034] Figure 8 It is a structure diagram of the electronic device of the video coding method based on the region of interest. DETAILED DESCRIPTION

[0035] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by the person of ordinary skill in the art without making creative labor should belong to the protection scope of the present application.

[0036] It should be noted that the terms "first", "second", "target" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0037] Embodiment one

[0038] Figure 1 A flowchart of a video coding method based on a region of interest is provided for the first embodiment of the present application. The first embodiment can be applied to the case of video region of interest positioning assisted by a detection sensor for a low-power video device, such as a low-power video door lock or a video peephole including a camera. The method can be performed by a video coding device based on a region of interest. The video coding device based on a region of interest can be implemented in the form of hardware and / or software and can be configured in an electronic device with data processing capability. As shown in Figure 1 The method comprises the following steps.

[0039] In S110, distance information and angle information of a target object relative to a detection sensor are obtained based on the detection sensor.

[0040] The detection sensor can be a sensor capable of detecting an object. For example, the detection sensor can include a laser sensor, a radar sensor, an infrared sensor, an ultrasonic sensor, and the like. Specifically, the detection sensor has the capability of collecting distance information and angle information of a target object in the field of view of the camera. The configuration position of the detection sensor is not limited. The detection sensor can be integrated in the camera, or the detection sensor can be installed in a certain range around the camera, i.e., the detection sensor and the camera are two independent devices. Figure 2 A working principle diagram of a detection sensor is provided for the first embodiment of the present application. As shown in Figure 2 The detection sensor on the door lock can emit a detection signal, such as a laser signal, an ultrasonic signal, a radar signal, an infrared signal, and the like, at low power. The detection signal can detect the distance information and angle information of the target object in the field of view of the camera.

[0041] In this embodiment, the distance information and angle information of the target object relative to the detection sensor are first obtained by the detection sensor. Optionally, the distance information includes a horizontal distance of the target object from the detection sensor, a first height distance of a top end of the target object from the detection sensor, and a second height distance of a bottom end of the target object from the detection sensor. The angle information includes a horizontal angle between the target object and a first reference line, a first vertical angle between the top end of the target object and the first reference line, and a second vertical angle between the bottom end of the target object and the first reference line. The first reference line has the detection sensor as the origin and is perpendicular to the plane on which the detection sensor is located.

[0042] In this embodiment, a detection space coordinate system is first established with the detection sensor as the origin. In this system, the positive X-axis is perpendicular to the plane containing the detection sensor and points towards the center of the camera's field of view; the Y-axis lies on a plane perpendicular to the plane containing the detection sensor and is perpendicular to the X-axis; and the Z-axis is perpendicular to both the X-axis and Y-axis. Figure 3 This is a schematic diagram of the detection spatial coordinate system of a video door lock according to Embodiment 1 of the present invention. Figure 3 As shown, the video door lock includes a detection sensor and a camera, with the dotted line indicating the direction of the center of the camera's field of view. Figure 3 For example, the plane where the video door lock is located is the same plane where the detection sensor is located, and this plane is perpendicular to the ground. Taking the location of the detection sensor as the origin O, and a direction perpendicular to the plane where the video door lock is located (the plane where the detection sensor is located) and pointing towards the center of the field of view (… Figure 3 The direction indicated by the dashed line is the positive X-axis. The Y-axis is the direction perpendicular to the X-axis on a plane that is perpendicular to the plane where the detection sensor is located (parallel to the ground). The Z-axis is the direction perpendicular to both the X-axis and the Y-axis (perpendicular to the ground and upwards).

[0043] Figure 4A and Figure 4B These are, respectively, a horizontal cross-sectional view and a vertical cross-sectional view of a detection spatial coordinate system provided in Embodiment 1 of the present invention. Figure 4A As shown, the positive X-axis direction is the direction of the first reference line, L1 represents the horizontal distance between the target and the detection sensor, and θ1 represents the horizontal angle between the target and the first reference line. Figure 4B As shown, L2 represents the first height distance between the top of the target and the detection sensor, L3 represents the second height distance between the bottom of the target and the detection sensor, θ2 represents the first vertical angle between the top of the target and the first reference line, and θ3 represents the second vertical angle between the bottom of the target and the first reference line.

[0044] S120, based on the lens imaging principle, determines the imaging position information of the target object on the imaging sensor in the camera according to distance information, angle information and camera attribute information.

[0045] The lens imaging principle can refer to the principle where the target object is inverted and imaged on the imaging sensor in the camera, and the imaging sensor converts the target object's light signal into an electrical signal. Camera attribute information can be used to characterize the camera's properties. Optionally, the camera attribute information includes the equivalent focal length of the lens assembly and the imaging sensor. Imaging position information can be used to characterize the imaging position of the target object on the imaging sensor in the camera.

[0046] In this embodiment, optionally, based on the lens imaging principle, the imaging position information of the target object on the imaging sensor of the camera is determined according to the distance information, the angle information and the camera attribute information, including: based on the lens imaging principle, the imaging horizontal distance, the first imaging height and the second imaging height of the target object in the first imaging on the imaging sensor are determined according to the equivalent focal length and the horizontal angle, the first vertical angle and the second vertical angle respectively; wherein the imaging horizontal distance is the horizontal distance of the target object from the first imaging center; and the imaging position information is determined according to the imaging horizontal distance, the first imaging height and the second imaging height.

[0047] Wherein, the first imaging can refer to the imaging of the target object on the imaging sensor. The imaging horizontal distance can refer to the horizontal distance of the target object from the first imaging center. The first imaging center can refer to the imaging center of the imaging sensor. The first imaging height can refer to the length of the imaging sensor in the vertical direction from the vertical center downward. The second imaging height can refer to the length of the imaging sensor in the vertical direction from the vertical center upward.

[0048] Figure 5A And Figure 5B are respectively a horizontal cross-sectional view and a vertical cross-sectional view of lens imaging provided by the first embodiment of the present application. As shown in Figure 5A , f represents the equivalent focal length, w1 represents the imaging horizontal distance, and w and h respectively represent the imaging width and the imaging height of the imaging sensor. The imaging horizontal distance can be calculated by the formula w1=fxtan(θ1). As shown in Figure 5B , h12 and h11 respectively represent the first imaging height and the second imaging height. The first imaging height and the second imaging height can be calculated by the formulas h12=fxtan(θ2) and h11=fxtan(θ3) respectively. After the imaging horizontal distance, the first imaging height and the second imaging height are determined, the imaging position information can be determined.

[0049] S130, based on the feature of the region of interest, the final position information of the region of interest on the image is determined according to the distance information, the angle information, the imaging position information, the imaging size information of the imaging sensor and the image resolution information.

[0050] The region of interest can refer to an imaging region containing part or all of the target object. The region of interest feature can refer to a feature of part or all of the target object contained in the region of interest. For example, assuming that the target object is a person, the region of interest can be an imaging region containing a face or a human body; accordingly, the region of interest feature can be a face region feature or a human body region feature. Taking the face region feature as an example, the face region feature can include size information of the head, such as the maximum width and the maximum height of the head. According to the summary statistics of the face in GB 2428-1998 standard (the maximum head of a woman is smaller than that of a man, and here it is assumed that the head of a man is estimated), the maximum head height of a man is 241 mm, and the maximum head width is 164 mm, so a rectangular region containing the head of the target object with a height greater than 241 mm and a width greater than 164 mm can be set as the region of interest. For example, a rectangular region of interest with a height of 300 mm and a width of 200 mm can be set. It should be noted that the shape and size of the region of interest in the embodiment are not limited, and can be set according to actual needs. For example, the region of interest can be a rectangle or a square.

[0051] The imaging size information can be used to represent the imaging size of the imaging sensor. For example, the imaging size information can include the imaging width (such as w in Figure 5A ) and the imaging height (such as h in Figure 5A ). The image resolution information can be used to represent the real resolution of the image picture. For example, the image resolution information can include the horizontal resolution and the vertical resolution of the image. For example, assuming that the image resolution information is 1080P (1920x1080), the horizontal resolution of the image is 1920, and the vertical resolution of the image is 1080. The final position information can be used to represent the final position of the region of interest on the image picture. It should be noted that the imaging result of the target object on the imaging sensor is inversely related to the imaging result of the target object on the image picture.

[0052] Figure 6A An imaging schematic diagram of a region of interest on an imaging sensor is provided for the first embodiment of the present application. As shown in Figure 6A , the center position of the imaging sensor is used as the coordinate origin to establish an imaging coordinate system, and the X1 axis and the Y1 axis represent the horizontal axis and the vertical axis of the imaging coordinate system, respectively. Wherein a and b represent the imaging height and the imaging width of the region of interest on the imaging sensor, respectively. Assuming that the region of interest includes a face, the coordinates (x, y) represent the position of the head in the imaging coordinate system. Figure 6B An imaging schematic diagram of a region of interest on an image picture is provided for the first embodiment of the present application. As shown in Figure 6BAs shown, an image coordinate system is established with the image center position as the coordinate origin, and the X2 axis and the Y2 axis represent the horizontal axis and the vertical axis of the image coordinate system, respectively. Wherein, W and H are the horizontal resolution and the vertical resolution of the image, respectively, W1 represents the horizontal distance of the target object to the coordinate origin of the image coordinate system (the horizontal distance of the final image), H1 represents the distance from the bottom end of the target object to the X2 axis (the second height of the final image), H2 represents the distance from the top end of the target object to the X2 axis (the first height of the final image), and A1 and B1 represent the height and the width of the region of interest on the image, respectively. Assuming that the region of interest includes a face, the coordinates (X, Y) represent the position of the head in the image coordinate system.

[0053] In the embodiment, optionally, the final position information of the region of interest on the image is determined according to the distance information, the angle information, the imaging position information, the imaging size information of the imaging sensor and the image resolution information, including: determining the target object reference size information based on the region of interest feature; determining the imaging region size information of the region of interest on the first image according to the target object reference size information, the distance information, the angle information and the imaging position information; and determining the final position information according to the imaging region size information, the imaging position information, the imaging size information of the imaging sensor and the image resolution information.

[0054] Wherein, the target object reference size information can be used to represent the real size of the target object. For example, the target object reference size information can include head reference width information and head reference height information. Wherein, the head reference width information and the head reference height information can be used to represent the real width and height of the head, respectively. The imaging region size information can be used to represent the size of the region of interest on the imaging sensor. For example, the imaging region size information can include the imaging height (such as a) and the imaging width (such as b) of the region of interest on the imaging sensor. Figure 6A Figure 6A

[0055] In the embodiment, optionally, the region of interest feature is a face region feature; and the target object reference size information is determined based on the region of interest feature, including: determining the target object reference size information based on the head size statistical information in the face region feature. Wherein, the head size statistical information can include the maximum width and the maximum height of the head, which can be specifically referred to the summary statistics of the face in the GB2428-1998 standard.

[0056] ​​In the embodiment, the imaging area size information of the region of interest on the first image is determined according to the target reference size information, the distance information, the angle information and the imaging position information, which includes: determining the imaging area width information according to the head reference width information, the imaging horizontal distance, the horizontal distance between the target and the sensor and the horizontal angle; determining the imaging area height information according to the head reference height information, the first imaging height, the second imaging height, the first height distance, the second height distance, the first vertical angle and the second vertical angle; and determining the imaging area size information according to the imaging area width information and the imaging area height information.

[0057] In the embodiment, the imaging area height information and the imaging area width information can be determined according to the following proportional relationship respectively:

[0058] a / A=(h11+h12) / (L2×sin(θ2)+L3×sin(θ3)), b / B=w1 / (L1×sin(θ1));

[0059] wherein A and B represent the head reference height information and the head reference width information respectively.

[0060] Therefore, a=A×(h11+h12) / (L2×sin(θ2)+L3×sin(θ3)), and b=B×w1 / (L1×sin(θ1)).

[0061] In the embodiment, the final position information is determined according to the imaging area size information, the imaging position information, the imaging size information of the imaging sensor and the image resolution information, which includes: determining the proportional information according to the imaging size information of the imaging sensor and the image resolution information; determining the final image horizontal distance and the final image first height according to the imaging horizontal distance and the first imaging height and the proportional information respectively; determining the final image area size information according to the imaging area size information and the proportional information; determining the reference position information of the top end of the target according to the final image horizontal distance and the final image first height; and determining the final position information of the region of interest on the image according to the reference position information and the final image area size information.

[0062] The proportion information can include width proportion information and height proportion information. The width proportion information can be determined according to the imaging width information of the imaging sensor and the horizontal resolution of the image, and can be represented as w / W. The height proportion information can be determined according to the imaging height information of the imaging sensor and the vertical resolution of the image, and can be represented as h / H. The final image region size information can be used to represent the size of the region of interest on the image. For example, the final image region size information can include final image region height information and final image region width information. The final image region height information and the final image region width information can be used to represent the height (such as A1 in Figure 6B ) and the width (such as B1 in Figure 6B ) of the region of interest on the image, respectively. The reference position information can refer to the position information of the top end of the target object, such as the position information of the coordinates (X, Y) in Figure 6B

[0063] In this embodiment, first, the width proportion information w / W and the height proportion information h / H are determined according to the imaging size information of the imaging sensor and the image resolution information. Then, the final image horizontal distance W1 can be determined as W1=(w1 / w)×W according to the proportion relationship W1 / W=w1 / w, and the final image first height H2 can be determined as H2=(h12 / h)×H according to the proportion relationship H2 / H=h12 / h. Further, the final image region height information A1 can be determined as A1=(a / h)×H according to the imaging region height information and the height proportion information, and the final image region width information B1 can be determined as B1=(b / w)×W according to the imaging region width information and the width proportion information. The reference position information of the top end of the target object can be determined according to the final image horizontal distance and the final image first height. For example, the reference position information (X, Y) of the top end of the target object is (W1, H2). Figure 6B Finally, the final position information of the region of interest on the image can be determined according to the reference position information and the final image region size information. For example, the final position information of the region of interest on the image is a rectangular region position formed by taking the coordinates (W1, H2) as the origin, moving to the left and right sides by B1×1 / 2 pixel points, and moving downward by A1 pixel points. The rectangular region is the ROI encoding region. Figure 6B

[0064] S140, according to the final position information, the image frames in the video to be encoded are ROI encoded.

[0065] ​​The to-be-encoded video can be a video to be encoded and can be captured by a camera configured with a probe sensor. The ROI encoding can be a process of losslessly or lossily encoding the region of interest and highly compressively encoding other regions, and can be used to retain key information (i.e., information of the region of interest) in the image frame while reducing the size of the video and the transmission bandwidth.

[0066] In this embodiment, after the final position information of the region of interest on the image is determined, the image frame in the to-be-encoded video can be encoded according to the final position information. Specifically, the ROI encoding technology in the prior art can be used to losslessly or lossily encode the region of interest in the image frame in the to-be-encoded video according to the final position information, and to highly compressively and highly lossily encode the non-region of interest (i.e., other regions outside the region of interest) in the image frame in the to-be-encoded video. In this way, the features of the region of interest in the image frame can be ensured to be clear and lossless, and the useless information picture can be compressed to the greatest extent, thereby effectively reducing the size of the to-be-encoded video and reducing the transmission bandwidth and storage pressure.

[0067] The technical scheme of the embodiment of the application comprises the following steps: obtaining distance information and angle information of a target object relative to a probe sensor based on the probe sensor; determining imaging position information of the target object on an imaging sensor in a camera based on a lens imaging principle and according to the distance information, the angle information and camera attribute information; determining final position information of a region of interest on an image based on a region of interest feature and according to the distance information, the angle information, the imaging position information, imaging size information of the imaging sensor and image resolution information; and performing ROI encoding on an image frame in a to-be-encoded video according to the final position information. The technical scheme can assist a low-power video device in positioning a video region of interest through a probe sensor, and can take different encoding measures on the region of interest and the non-region of interest, thereby ensuring the definition of the region of interest while reducing the size of the video and the video transmission power consumption.

[0068] Embodiment Two

[0069] Figure 7 A structure schematic diagram of a video encoding device based on a region of interest provided in the second embodiment of the application is shown in FIG. 2. The device can execute the video encoding method based on a region of interest provided in any embodiment of the application, and has the corresponding function modules and beneficial effects of the execution method. As shown in FIG. 2, the device comprises: Figure 7

[0070] The probe sensor information determination module 210 is configured to obtain distance information and angle information of a target object relative to a probe sensor based on the probe sensor.

[0071] ​The imaging position information determination module 220 is configured to determine imaging position information of the target object on an imaging sensor of the camera based on a lens imaging principle according to the distance information, the angle information and camera attribute information.

[0072] The final position information determination module 230 is configured to determine final position information of the region of interest on an image based on a region of interest feature according to the distance information, the angle information, the imaging position information, imaging size information of the imaging sensor and image resolution information.

[0073] The ROI encoding module 240 is configured to perform ROI encoding on an image frame in the video to be encoded according to the final position information.

[0074] The camera for collecting the video to be encoded is provided with a detection sensor, and the detection sensor has the capability of collecting distance and angle information of the target object in a field of view of the camera.

[0075] Optionally, the distance information includes a horizontal distance of the target object from the detection sensor, a first height distance of a top end of the target object from the detection sensor and a second height distance of a bottom end of the target object from the detection sensor.

[0076] The angle information includes a horizontal angle between the target object and a first reference line, a first vertical angle between a top end of the target object and the first reference line and a second vertical angle between a bottom end of the target object and the first reference line.

[0077] The first reference line takes the detection sensor as an origin and is perpendicular to a plane where the detection sensor is located.

[0078] Optionally, the camera attribute information includes an equivalent focal length of a lens group and an imaging sensor.

[0079] The imaging position information determination module 220 is configured to:

[0080] determine, based on a lens imaging principle, an imaging horizontal distance, a first imaging height and a second imaging height in first imaging of the target object on the imaging sensor according to the equivalent focal length and the horizontal angle, the first vertical angle and the second vertical angle respectively, wherein the imaging horizontal distance is a horizontal distance of the target object from a first imaging center.

[0081] determine the imaging position information according to the imaging horizontal distance, the first imaging height and the second imaging height.

[0082] Optionally, the final position information determination module 230 includes:

[0083] The target object reference size information determination unit is configured to determine target object reference size information based on the region of interest feature;

[0084] The imaging region size information determination unit is configured to determine imaging region size information of the region of interest on the first image according to the target object reference size information, the distance information, the angle information and the imaging position information;

[0085] The final position information determination unit is configured to determine final position information according to the imaging region size information, the imaging position information, imaging size information of the imaging sensor and image resolution information.

[0086] Optionally, the region of interest feature is a face region feature.

[0087] The target object reference size information determination unit is configured to:

[0088] Determine the target object reference size information based on head size statistical information in the face region feature.

[0089] Optionally, the target object reference size information includes head reference width information and head reference height information.

[0090] The imaging region size information determination unit is configured to:

[0091] Determine imaging region width information according to the head reference width information, the imaging horizontal distance, the horizontal distance between the target object and the detection sensor and the horizontal angle.

[0092] Determine imaging region height information according to the head reference height information, the first imaging height, the second imaging height, the first height distance, the second height distance, the first vertical angle and the second vertical angle.

[0093] Determine the imaging region size information according to the imaging region width information and the imaging region height information.

[0094] Optionally, the final position information determination unit is configured to:

[0095] Determine scale information according to the imaging size information of the imaging sensor and the image resolution information.

[0096] Determine a final image horizontal distance and a final image first height according to the imaging horizontal distance and the first imaging height and the scale information respectively.

[0097] Determine final image region size information according to the imaging region size information and the scale information.

[0098] The reference position information of the top end of the target object is determined according to the final image horizontal distance and the final image first height;

[0099] The final position information of the region of interest on the image is determined according to the reference position information and the final image region size information.

[0100] The video encoding device based on the region of interest provided by the embodiment of the present application can execute the video encoding method based on the region of interest provided by any embodiment of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0101] Embodiment three

[0102] Figure 8 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0103] As shown in Figure 8 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0104] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0105] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The processor 11 performs various methods and processes described above, such as the region-of-interest based video encoding method.

[0106] In some embodiments, the region-of-interest based video encoding method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the region-of-interest based video encoding method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the region-of-interest based video encoding method by any other appropriate means, such as by means of firmware.

[0107] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0108] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, enables the functions / acts specified in the flowcharts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package and partially on a remote machine or entirely on a remote machine or server.

[0109] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0111] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0112] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0113] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in different orders, as long as the desired results of the technical solutions of the present disclosure can be achieved, and the present disclosure is not limited herein.

[0114] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method of video coding based on a region of interest, characterized in that, A detection sensor is configured for a camera for collecting a video to be encoded, the detection sensor having the capability of collecting distance and angle information of a target object in a field of view of the camera, the method comprising: obtaining distance information and angle information of the target object relative to the detection sensor based on the detection sensor; determining imaging position information of the target object on an imaging sensor of the camera based on the distance information, the angle information and camera attribute information according to a lens imaging principle; determining final position information of a region of interest on an image based on the distance information, the angle information, the imaging position information, imaging size information of the imaging sensor and image resolution information according to a region of interest feature; performing ROI encoding on image frames in the video to be encoded according to the final position information; wherein the ROI encoding refers to a process of lossless or low-loss encoding for the region of interest and high compression encoding for other regions; wherein the distance information comprises a horizontal distance of the target object from the detection sensor, a first height distance of a top end of the target object from the detection sensor and a second height distance of a bottom end of the target object from the detection sensor; the angle information comprises a horizontal angle between the target object and a first reference line, a first vertical angle between the top end of the target object and the first reference line and a second vertical angle between the bottom end of the target object and the first reference line; wherein the first reference line is perpendicular to a plane where the detection sensor is located and takes the detection sensor as an origin; a detection space coordinate system is established taking the detection sensor as an origin; wherein the camera attribute information comprises an equivalent focal length of a lens group and an imaging sensor; determining imaging position information of the target object on the imaging sensor of the camera based on the distance information, the angle information and camera attribute information according to a lens imaging principle, comprising: determining an imaging horizontal distance, a first imaging height and a second imaging height in a first imaging of the target object on the imaging sensor based on the equivalent focal length and the horizontal angle, the first vertical angle and the second vertical angle respectively according to a lens imaging principle; wherein the imaging horizontal distance is a horizontal distance of the target object from a first imaging center; determining imaging position information according to the imaging horizontal distance, the first imaging height and the second imaging height; wherein the first imaging height refers to a length of the imaging sensor in a vertical direction from a vertical center downward, and the second imaging height refers to a length of the imaging sensor in the vertical direction from the vertical center upward.

2. The method of claim 1, wherein, determining final position information of a region of interest on an image based on the distance information, the angle information, the imaging position information, imaging size information of the imaging sensor and image resolution information according to a region of interest feature, comprising: determining target object reference size information based on a region of interest feature; determining imaging area size information of the region of interest on the first imaging according to the target object reference size information, the distance information, the angle information and the imaging position information; The final position information is determined according to the imaging area size information, the imaging position information, imaging size information of an imaging sensor and image resolution information.

3. The method of claim 2, wherein, The region-of-interest feature is a human face region feature; The target object reference size information is determined based on the region-of-interest feature, including: The target object reference size information is determined based on head size statistical information in the human face region feature.

4. The method of claim 3, wherein, The target object reference size information includes head reference width information and head reference height information; The imaging area size information of the region-of-interest on the first image is determined according to the target object reference size information, the distance information, the angle information and the imaging position information, including: The imaging area width information is determined according to the head reference width information, the imaging horizontal distance, the horizontal distance of the target object from the detection sensor and the horizontal angle; The imaging area height information is determined according to the head reference height information, the first imaging height, the second imaging height, the first height distance, the second height distance, the first vertical angle and the second vertical angle; The imaging area size information is determined according to the imaging area width information and the imaging area height information.

5. The method of claim 4, wherein, The final position information is determined according to the imaging area size information, the imaging position information, imaging size information of an imaging sensor and image resolution information, including: The scale information is determined according to the imaging size information of the imaging sensor and the image resolution information; The final image horizontal distance and the final image first height are respectively determined according to the imaging horizontal distance and the first imaging height and the scale information; The final image area size information is determined according to the imaging area size information and the scale information; The reference position information of the top end of the target object is determined according to the final image horizontal distance and the final image first height; The final position information of the region-of-interest on the image is determined according to the reference position information and the final image area size information.

6. A video encoding apparatus based on a region of interest, characterized in that, A detection sensor is configured for a camera for collecting a video to be encoded, the detection sensor has the capability of collecting distance and angle information of a target object in a field of view of the camera, and the device includes: A detection sensor information determination module is configured to acquire distance information and angle information of the target object relative to the detection sensor based on the detection sensor; An imaging position information determination module is configured to determine imaging position information of the target object on an imaging sensor of the camera based on lens imaging principles and according to the distance information, the angle information and camera attribute information; A final position information determination module is configured to determine final position information of the region-of-interest on an image based on a region-of-interest feature and according to the distance information, the angle information, the imaging position information, imaging size information of an imaging sensor and image resolution information; An ROI encoding module is configured to perform ROI encoding on image frames in the video to be encoded according to the final position information, wherein the ROI encoding refers to a process of performing lossless or low-loss encoding on a region-of-interest and high-compression encoding on other regions. The distance information includes a horizontal distance of the target object from the detection sensor, a first height distance of a top end of the target object from the detection sensor, and a second height distance of a bottom end of the target object from the detection sensor. The angle information includes a horizontal angle between the target object and a first reference line, a first vertical angle between a top end of the target object and the first reference line, and a second vertical angle between a bottom end of the target object and the first reference line. The first reference line is perpendicular to a plane on which the detection sensor is located. The camera attribute information includes an equivalent focal length of a lens group and an imaging sensor. The imaging position information determination module is configured to: based on a lens imaging principle, determine, according to the equivalent focal length and the horizontal angle, the first vertical angle, and the second vertical angle, an imaging horizontal distance, a first imaging height, and a second imaging height of the target object in a first image on the imaging sensor, wherein the imaging horizontal distance is a horizontal distance of the target object from a first image center. determine imaging position information according to the imaging horizontal distance, the first imaging height, and the second imaging height. The first imaging height refers to a length of the imaging sensor in a vertical direction from a vertical center downward, and the second imaging height refers to a length of the imaging sensor in the vertical direction from the vertical center upward.

7. A video coding electronic device based on a region of interest, characterized by, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the video coding method based on the region of interest according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to execute the video coding method based on the region of interest according to any one of claims 1-5 when executed.

Citation Information

Patent Citations

  • Coal mine underground miner face area detecting method

    CN104657719A

  • Position information determination method and device and storage medium

    CN112955711A