Virtual double-camera technology of intelligent lock cat eye

By using wide-angle lens and virtual dual camera technology in the smart lock cat eye, the dual-screen correction and fusion processing of one machine is solved, and the existing smart lock cat eye device requires two cameras is reduced, cost and power consumption is simplified, and data analysis and mutual passage process is simplified.

CN119996834APending Publication Date: 2025-05-13HANGZHOU XINQIAO INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510131057.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing smart lock cat eye device requires two cameras, which leads to higher cost and power consumption, and makes algorithm data analysis and interoperability more difficult.

Method used

It adopts wide-angle lens and virtual dual camera technology, and through one-machine dual-screen correction and fusion processing, dynamic dual-screen display is achieved, covering all important areas in front of the door.

Benefits of technology

It effectively reduces the number and power consumption of cameras, simplifies algorithm data analysis and mutual passage, and solves the problems of fixed cameras being compatible with children's height and incomplete express delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996834A_ABST
    Figure CN119996834A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual double-shooting technology of an intelligent lock cat eye. The virtual double-shooting technology comprises the following steps: S1, selecting a wide-angle lens: selecting the wide-angle lens; s2, correcting a designated area right in front of the current picture; correcting a middle picture; s3, performing human shape detection and parcel detection: performing human shape detection and parcel detection on the corrected picture to obtain a position block diagram of a child, a position block diagram of an adult and a position diagram of an express parcel; s4, video judgment: judging the number of children and adults in the picture, and when only children exist in the picture and the door is opened by the children, reporting an alarm that the children go out independently; s5, dynamic snapshot and recording: after detecting that the child opens the door, recording the whole process of going out by the child; s6, self-adaptively generating dynamic double pictures: self-adaptively generating the dynamic double pictures according to the activity area of the person and the placement area of the express; and S7, fusing the pictures: combining the two pictures into one on the camera side to form a fused picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart door locks, and in particular to a virtual dual-camera technology for a smart lock cat's eye. Background Art

[0002] Smart lock cat's eye is a smart door lock that integrates the cat's eye function (i.e. door mirror function). It not only has the security protection function of traditional door locks, but also incorporates modern technological elements to achieve multiple functions such as remote monitoring and intelligent alarm. With the advancement of technology and people's increasing attention to family security, smart lock cat's eye, as an important part of smart home, has an increasingly strong market demand. Especially in first-tier cities and high-end residential areas, smart lock cat's eye has become a standard feature.

[0003] After searching, the Chinese patent number CN219262029U discloses a multi-angle cat-eye smart door lock, which belongs to the field of smart door lock technology. It is composed of an outer lock frame and an inner lock frame, and the outer lock frame is deployed on the outer side of the home door panel, and the inner lock frame is deployed on the inner side of the home door panel: the outer lock frame is composed of an outer lock body and a collection beam; the collection beam has a first collection part and a second collection part, the first collection part is located in the upper half of the collection beam facing the outside of the door, and the second collection part is located directly below the first collection part; the first collection part is provided with a first electronic cat's eye, and the second collection part is provided with a second electronic cat's eye; the first electronic cat's eye and the collection beam face the outside of the door and keep parallel photography; the second electronic cat's eye and the collection beam face the outside of the door and keep an angle of 25° to 50° to achieve the supervision effect of express delivery or take-out on the premise of distinguishing users and visitors.

[0004] However, the above device requires two cameras during use, which is more costly and consumes more power. In addition, algorithm data analysis and intercommunication of two independent cameras are more difficult. Therefore, a virtual dual-camera technology for a smart lock cat's eye is proposed. Summary of the invention

[0005] The purpose of the present invention is to reduce the number of cameras and power consumption of the front panel of a smart door lock to achieve a dual-screen camera, and a virtual dual-camera technology of a smart lock cat's eye is proposed.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A virtual dual-camera technology for a smart lock peephole includes the following steps:

[0008] S1: Select a wide-angle lens: Select a wide-angle lens, which captures all important areas in front of the door, including the ground in front of the door and the area where the express delivery is placed;

[0009] S2: Correct the designated area in front of the current image: Correct the middle image to obtain a normal distorted image; Human detection and package detection:

[0010] S3: Human shape detection and package detection: By performing human shape detection and package detection on the corrected image, we can obtain the position frame diagram of children, the position frame diagram of adults, and the position map of express packages. We can track the detected targets and obtain their motion information such as position, speed, and acceleration. Based on the motion information of the targets, we can determine whether they have entered or left the specified area. We can adjust the human figure area so that the human face is in the upper middle area of ​​the image.

[0011] S4: Video judgment: The door unlocking message is used as the cutoff point for video judgment. When the door is unlocked, judgment is performed 2 to 5 seconds before the door is unlocked to determine the number of children and adults in the picture. If there are only children in the picture and the door is opened by a child, an alarm of a child going out alone is reported;

[0012] S5: Dynamic capture and recording: After detecting that a child opens the door, the entire process is captured as the person moves, generating a dynamic image to record the entire process of the child going out, ensuring that there is no blind spot in monitoring;

[0013] S5: Adaptively generate dynamic dual screens: Adaptively generate dynamic dual screens according to the activity area of ​​people and the area where the express delivery is placed. The dynamic dual screens respectively display the people in front of the door and the express delivery under the door. The display range of the express delivery under the door is relatively dynamic. According to the distribution of the express delivery, the largest area is selected to cover the express delivery area.

[0014] S6: Picture fusion: The two pictures are combined into one on the camera side to form a third picture, which covers all important areas in front of the door and displays the information of the person and the courier at the same time.

[0015] The above further includes:

[0016] Furthermore, in S1, when selecting a wide-angle lens, the relationship between the focal length and the viewing angle of the lens is considered. The relationship between the viewing angle (θ), the focal length (f), and the sensor size (D) is expressed by the following formula:

[0017] θ=2*arctan(D / (2*f))

[0018] Where D is the diagonal length of the sensor (unit: mm), and f is the focal length of the lens (unit: mm). In order to obtain a wider viewing angle, a lens with a shorter focal length needs to be selected.

[0019] The wide-angle lens is installed above or on the side of the smart door lock to capture all important areas in front of the door. The wide-angle lens is tilted downward to capture details on the ground. If there is a spacious open space in front of the door, the lens can be tilted slightly to both sides to expand the monitoring range. At the same time, the vertical angle of the lens needs to be adjusted to ensure that the ground and steps in front of the door can be captured.

[0020] Further, in S1, the wide-angle lens is installed above or on the side of the smart door lock to capture all important areas in front of the door. At the same time, the angle of the wide-angle lens is adjusted according to the actual situation in front of the door to obtain the best monitoring effect. According to the actual situation in front of the door, the horizontal angle of the lens is adjusted to ensure that all important areas in front of the door can be captured. For example, if there is a spacious open space in front of the door, the lens can be tilted slightly to both sides to expand the monitoring range, and the vertical angle of the lens can be adjusted to ensure that the ground and steps in front of the door can be captured. Usually, the lens should be tilted slightly downward to capture the details of the ground. At the same time, fine-tuning is required according to the actual situation in front of the door to obtain the best monitoring effect.

[0021] Furthermore, in S2, the specific steps of the correction process are:

[0022] Image acquisition and preprocessing: obtain the original image data of the wide-angle lens, and preprocess the original image, such as denoising and grayscale, to improve the accuracy of subsequent processing;

[0023] Distortion model establishment: The wide-angle lens distortion includes barrel distortion and pincushion distortion. The barrel distortion is manifested as the center of the picture being enlarged compared to the edge, while the pincushion distortion is the opposite. According to the distortion type, a mathematical model is selected to describe it;

[0024] Feature point extraction and matching: Select a set of feature points in the original image. The feature points should have obvious geometric features, such as corners and edges. Use SIFT to extract the positions and descriptors of the feature points. Find the feature points corresponding to the feature points in the original image in an ideal distortion-free image or an image with known distortion parameters.

[0025] Distortion parameter solution: According to the matching results of the feature points, the least square method is used to solve the distortion parameters, which include radial distortion coefficients (k1, k2, k3, etc.) and tangential distortion coefficients (p1, p2, etc.);

[0026] Image correction: According to the obtained distortion parameters, the original image is corrected. The correction process involves coordinate transformation and pixel resampling. The coordinate transformation formula is:

[0027]

[0028]

[0029] Among them, x original ,y original is the coordinate in the original image, x corrected ,y corrected are the corrected coordinates, r is the distance from the point to the image center, k1, k2, ... are the radial distortion coefficients, and p1, p2 are the tangential distortion coefficients.

[0030] Post-processing: interpolate the corrected image to obtain smooth pixel values, and perform operations such as cropping and scaling on the corrected image as needed.

[0031] In the distortion model establishment, the barrel distortion is manifested as the center of the picture being enlarged compared to the edge, that is, the lines in the center area of ​​the image are bent outward, making the image barrel-shaped. The barrel distortion mathematical model is described using the radial distortion model, and the mathematical formula is:

[0032] x distorted =x(1+k1r 2 +k2r 4 +k3r 6 )

[0033] y distorted =y(1+k1r 2 +k2r 4 +k3r 6 )

[0034] Where (x, y) is the pixel coordinate in the original undistorted image coordinate system, (x distorted ,y distorted ) is the pixel coordinate in the distorted image coordinate system, r 2 =x 2 +y 2 It represents the square of the distance from the pixel to the center of the image. k1, k2, and k3 are radial distortion coefficients that represent the degree of distortion. The radial distortion coefficient is a negative value. Especially when k1 is dominant, it determines the degree and type of distortion. The coefficient is obtained through actual measurement and calibration.

[0035] The pincushion distortion is opposite to the barrel distortion, and is manifested as the edge of the picture being enlarged compared to the center, that is, the lines in the edge area of ​​the image are bent inward, making the image appear pincushion-shaped. The mathematical model of the pincushion distortion is also described using the radial distortion model. The sign of the distortion coefficient of the mathematical model of the pincushion distortion is opposite to that of the barrel distortion, and the mathematical formula is the same as that of the barrel distortion, but the value and sign of the coefficient are different, and the radial distortion coefficient is positive, especially when k1 is dominant.

[0036] In feature point extraction and matching, the specific steps of using SIFT to extract the position and descriptor of feature points are as follows:

[0037] Construct scale space: Construct the scale space of the image and detect feature points that are stable at different scales. The scale space is obtained by convolving the original image with Gaussian kernels of different scales. The calculation formula is: L(x,y,σ)=G(x,y,σ)*I(x,y)

[0038] Where L(x,y,σ) is the scale space image, G(x,y,σ) is the Gaussian function, I(x,y) is the original image, * represents the convolution operation, σ is the standard deviation of the Gaussian kernel, and represents the scale;

[0039] Detect extreme points: In scale space, extreme points are determined by comparing each pixel with 26 points in its neighborhood (8 neighboring points of the current scale and 9×2 points at corresponding positions of the upper and lower scales). If a pixel has a grayscale value greater or smaller than that of all its neighboring points, the point is considered an extreme point.

[0040] Accurate key point positioning: Use Taylor expansion to fit the scale space function and determine the exact position and scale of the key points. At the same time, the contrast of the fitted function value is calculated to eliminate the key points with low contrast, and the edge response is calculated to eliminate the key points with strong edge effects.

[0041] Calculate the direction of key points: Calculate one or more directions for each key point, maintain rotation invariance, and achieve this by calculating the gradient direction and amplitude of the pixels in the neighborhood of the key point. The calculation formula is:

[0042] Gradient Amplitude:

[0043] Where m(x, y) represents the gradient amplitude of the pixel at position (x, y) in the image. The gradient amplitude reflects the intensity of the change of the image at that point, that is, how fast the brightness of the image changes at that point. L(x+1, y)-L(xl, y) represents the difference in grayscale values ​​between the pixel at position (x, y) and its left and right adjacent pixels in the x direction. This difference reflects the rate of change of the image at that point along the x direction. L(x, y+1)-L(x, y-1) represents the difference in grayscale values ​​between the pixel at position (x, y) and its upper and lower adjacent pixels in the y direction. This difference reflects the rate of change of the image at that point along the y direction.

[0044] Gradient direction:

[0045] Wherein, θ(x, y) represents the gradient direction of the pixel at position (x, y) in the image. The gradient direction is the direction in which the image changes fastest at that point, that is, the brightness gradient direction of the image at that point. Arctan represents the inverse tangent function, which is used to calculate the angle of the gradient direction.

[0046] Then, the histogram is used to count the gradient direction and amplitude of the pixels in the neighborhood of the key point. The direction corresponding to the peak of the histogram is the main direction of the key point.

[0047] Generate feature descriptors: After determining the position, scale, and orientation of the key points, generate feature descriptors by selecting a square area (e.g., 16x16 pixels) around the key points and dividing it into smaller sub-areas (e.g., 4x4 pixels). For each sub-area, calculate the histogram of its gradient direction and magnitude, and combine these histograms into a vector as the feature descriptor of the key points.

[0048] Furthermore, in S3, the specific steps of using the YOLO algorithm for human detection and package detection are as follows:

[0049] The YOLO algorithm divides the input image into multiple grid cells, and each grid predicts multiple bounding boxes and category probabilities;

[0050] Collect a dataset of images or video data containing human figures and packages to train the detection model so that the detection model can recognize human figures and packages. During the training process, the YOLO algorithm learns the features in the image and establishes a mapping relationship between the features and the target. At the same time, a loss function is defined to measure the difference between the detection results output by the model and the true label. The loss function usually includes the positioning error of the bounding box and the classification error of the category probability.

[0051] In the inference phase, the input image is fed into the trained detection model to obtain the detection results of the human figure and the package, that is, the location information of the human figure and the package. These results include the coordinates, size, and category probability of the bounding box.

[0052] The detection results of the human figure and the package are jointly post-processed. For example, the detection frame can be adjusted and optimized according to the relative position and size relationship of the human figure and the package to obtain a more accurate detection result and obtain the final detection result.

[0053] Furthermore, in S4, the specific steps of video judgment are:

[0054] Feature extraction and classification: Extract feature information of humanoid targets tracked in S3, such as height, body proportions, facial features, etc., and use support vector machine to classify humanoid targets as adults or children based on the extracted feature information;

[0055] Number counting and behavior analysis: Count the number of adults and children in the video frame, analyze the door opening behavior, and determine whether it is performed by a child by detecting the movement of the door handle and the opening angle of the door;

[0056] Alarm triggering: If there is only a child in the picture and the door is opened by the child, an alarm will be triggered that the child is going out alone, and the alarm information will be sent to the preset parent's mobile phone or smart home control center.

[0057] Furthermore, in S5, the specific steps of dynamic capture and recording are:

[0058] Trigger condition: When the smart lock peephole system recognizes that a child opens the door alone through human detection technology, the process is triggered;

[0059] Detection and tracking: When a child opens the door, the dynamic capture function is immediately activated, and the tracking algorithm is used to track the child in real time to ensure that the target is not lost;

[0060] Capture and record: As the child moves, continuous capture begins, and each frame records the child's movements. These captured images are combined into a dynamic image in real time to show the entire continuous process of the child going out;

[0061] Generate animated image: These continuous captured images are processed to generate a smooth animated image, which records the entire process from the child opening the door to leaving the monitoring area, ensuring that there is no blind spot in monitoring;

[0062] Storage and viewing: The generated animated images are stored in the built-in storage of the smart lock or in the cloud server. Users can view these animated images through the mobile phone APP or other authorized devices to understand the specific situation of children going out.

[0063] Alarms and notifications: When a child goes out alone, the system sends an alarm notification to a preset mobile phone number or APP. The notification content includes the time and place of the child going out, as well as a link or preview of an animated picture, so that users can understand the situation in time and take appropriate measures.

[0064] Furthermore, in S6, the specific steps of adaptively generating the dynamic dual screen are:

[0065] Based on the results of target detection and tracking, two images are dynamically generated;

[0066] The first screen shows the area where people are active. When someone enters the area, the screen automatically switches to the image of that person.

[0067] The second screen shows the delivery area. When a delivery is placed or taken away, the screen will dynamically adjust as the number of deliveries increases or decreases. When a new delivery is placed in the monitoring area, the change is automatically detected and the display range is expanded accordingly to include the new delivery. When the delivery is taken away, the display range will automatically shrink.

[0068] The present invention has the following beneficial effects:

[0069] 1. In the present invention, a wide-angle lens is used to collect and correct the image of the area facing the screen and the area under the door, and the fusion processing is performed to produce the effect of one machine with two images, which effectively reduces the number of cameras and power consumption. At the same time, through the fusion processing, the two images are combined into a unique third image, which simplifies the algorithm data analysis and intercommunication process.

[0070] 2. In the present invention, a wide-angle lens is used to superimpose dual-image correction on the device side to achieve the effect of dynamic image adjustment, thereby solving the problem that fixed cameras are not compatible with children's height and that express delivery cannot be fully seen. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 This is a step diagram of a virtual dual-camera technology for a smart lock peephole proposed by the present invention;

[0072] Figure 2 It is a schematic diagram of the character area mentioned in the present invention. DETAILED DESCRIPTION

[0073] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0074] See also Figure 1 As shown, the present invention is a virtual dual-camera technology for a smart lock peephole, comprising the following steps:

[0075] S1: Select a wide-angle lens: Select a wide-angle lens, which captures all important areas in front of the door, including the ground in front of the door and the area where the express delivery is placed;

[0076] S2: Correct the designated area in front of the current image: Correct the middle image to obtain a normal distorted image; Human detection and package detection:

[0077] S3: Human shape detection and package detection: By performing human shape detection and package detection on the corrected image, we can obtain the position frame diagram of children, the position frame diagram of adults, and the position map of express packages. We can track the detected targets and obtain their motion information such as position, speed, and acceleration. Based on the motion information of the targets, we can determine whether they have entered or left the specified area. We can adjust the human figure area so that the human face is in the upper middle area of ​​the image.

[0078] S4: Video judgment: The door unlocking message is used as the cutoff point for video judgment. When the door is unlocked, judgment is performed 2 to 5 seconds before the door is unlocked to determine the number of children and adults in the picture. If there are only children in the picture and the door is opened by a child, an alarm of a child going out alone is reported;

[0079] S5: Dynamic capture and recording: After detecting that a child opens the door, the entire process is captured as the person moves, generating a dynamic image to record the entire process of the child going out, ensuring that there is no blind spot in monitoring;

[0080] S5: Adaptively generate dynamic dual screens: Adaptively generate dynamic dual screens according to the activity area of ​​people and the area where the express delivery is placed. The dynamic dual screens respectively display the people in front of the door and the express delivery under the door. The display range of the express delivery under the door is relatively dynamic. According to the distribution of the express delivery, the largest area is selected to cover the express delivery area.

[0081] S6: Picture fusion: The two pictures are combined into one on the camera side to form a third picture, which covers all important areas in front of the door and displays the information of the person and the courier at the same time.

[0082] In one embodiment, for the above S1, in S1, when a wide-angle lens is selected, the relationship between the focal length and the viewing angle of the lens is considered, and the relationship between the viewing angle (θ), the focal length (f), and the sensor size (D) is expressed by the following formula:

[0083] θ=2*arctan(D / (2*f))

[0084] Where D is the diagonal length of the sensor (unit: mm), and f is the focal length of the lens (unit: mm). In order to obtain a wider viewing angle, a lens with a shorter focal length needs to be selected.

[0085] Assuming the diagonal length of the sensor is 22 mm and you want to get a 180-degree viewing angle, the focal length f can be calculated to be approximately 11 mm according to the formula;

[0086] The wide-angle lens is installed above or on the side of the smart door lock to capture all important areas in front of the door. The wide-angle lens is tilted downward to capture details on the ground. If there is a spacious open space in front of the door, the lens can be tilted slightly to both sides to expand the monitoring range. At the same time, the vertical angle of the lens needs to be adjusted to ensure that the ground and steps in front of the door can be captured.

[0087] In one embodiment, for the above S1, in S1, the wide-angle lens is installed above or on the side of the smart door lock, so as to capture all important areas in front of the door. At the same time, the angle of the wide-angle lens is adjusted according to the actual situation in front of the door to obtain the best monitoring effect. According to the actual situation in front of the door, the horizontal angle of the lens is adjusted to ensure that all important areas in front of the door can be captured. For example, if there is a spacious open space in front of the door, the lens can be tilted slightly to both sides to expand the monitoring range, and the vertical angle of the lens can be adjusted to ensure that the ground and steps in front of the door can be captured. Usually, the lens should be tilted slightly downward to capture the details of the ground. At the same time, it is also necessary to make fine adjustments according to the actual situation in front of the door to obtain the best monitoring effect.

[0088] In one embodiment, for the above S2, in S2, the specific steps of the correction process are:

[0089] Image acquisition and preprocessing: obtain the original image data of the wide-angle lens, and preprocess the original image, such as denoising and grayscale, to improve the accuracy of subsequent processing;

[0090] Distortion model establishment: The wide-angle lens distortion includes barrel distortion and pincushion distortion. The barrel distortion is manifested as the center of the picture being enlarged compared to the edge, while the pincushion distortion is the opposite. According to the distortion type, a mathematical model is selected to describe it;

[0091] Feature point extraction and matching: Select a set of feature points in the original image. The feature points should have obvious geometric features, such as corners and edges. Use SIFT to extract the positions and descriptors of the feature points. Find the feature points corresponding to the feature points in the original image in an ideal distortion-free image or an image with known distortion parameters.

[0092] Distortion parameter solution: According to the matching results of the feature points, the least square method is used to solve the distortion parameters, which include radial distortion coefficients (k1, k2, k3, etc.) and tangential distortion coefficients (p1, p2, etc.);

[0093] Image correction: According to the obtained distortion parameters, the original image is corrected. The correction process involves coordinate transformation and pixel resampling. The coordinate transformation formula is:

[0094]

[0095]

[0096] Among them, x original ,y original is the coordinate in the original image, x corrected ,y corrected are the corrected coordinates, r is the distance from the point to the image center, k1, k2, ... are the radial distortion coefficients, and p1, p2 are the tangential distortion coefficients.

[0097] Post-processing: interpolate the corrected image to obtain smooth pixel values, and perform operations such as cropping and scaling on the corrected image as needed.

[0098] Suppose we have an image taken with a wide-angle lens, and there is obvious barrel distortion in the image. Follow the above steps to correct it.

[0099] Image acquisition and preprocessing: read the original image and perform denoising and grayscale processing.

[0100] Distortion model establishment: The radial distortion model is selected for description, and it is assumed that the distortion is mainly determined by the two radial distortion coefficients k1 and k2.

[0101] Feature point extraction and matching: Select a set of corner points in the original image as feature points, and use the SIFT algorithm to extract the location and descriptor of the feature points. Then, find the feature points corresponding to the feature points in the original image in an image with known distortion parameters.

[0102] Distortion parameter solution: Use the least squares method to solve the values ​​of k1 and k2.

[0103] Image correction: According to the obtained distortion parameters, the original image is corrected using the above coordinate transformation formula.

[0104] Post-processing: Perform bilinear interpolation on the rectified image to obtain smooth pixel values. Then, perform cropping and scaling operations on the rectified image to obtain the final correction result.

[0105] For the above distortion model establishment, in the distortion model establishment, the barrel distortion is manifested as the center of the picture is enlarged than the edge, that is, the lines in the center area of ​​the image are bent outward, making the image barrel-shaped. The barrel distortion mathematical model is described using the radial distortion model, and the mathematical formula is:

[0106] x distorted =x(1+k1r 2 +k2r 4 +k3r 6 )

[0107] y distorted =y(1+k1r 2 +k2r 4 +k3r 6 )

[0108] Where (x, y) is the pixel coordinate in the original undistorted image coordinate system, (x distorted ,y distorted ) is the pixel coordinate in the distorted image coordinate system, r 2 =x 2 +y 2 It represents the square of the distance from the pixel to the center of the image. k1, k2, and k3 are radial distortion coefficients that represent the degree of distortion. The radial distortion coefficient is a negative value. Especially when k1 is dominant, it determines the degree and type of distortion. The coefficient is obtained through actual measurement and calibration.

[0109] Suppose there is a point (x, y) = (100, 100), which is r = 141.42 from the center of the image. If k1 = -0.001, k2 = 0.00001, k3 = 0, the distorted coordinates (x_distorted, y_distorted) will be close to (100, 100) but slightly offset, because the negative value of k1 causes the coordinates to be slightly reduced, while the positive value of k2 makes a slight adjustment;

[0110] The pincushion distortion is opposite to the barrel distortion, and is manifested as the edge of the picture being enlarged compared to the center, that is, the lines in the edge area of ​​the image are bent inward, making the image appear pincushion-shaped. The mathematical model of the pincushion distortion is also described using the radial distortion model. The sign of the distortion coefficient of the mathematical model of the pincushion distortion is opposite to that of the barrel distortion, and the mathematical formula is the same as that of the barrel distortion, but the value and sign of the coefficient are different, and the radial distortion coefficient is positive, especially when k1 is dominant.

[0111] For the above feature point extraction and matching, in the feature point extraction and matching, the specific steps of using SIFT to extract the position and descriptor of the feature point are:

[0112] Construct scale space: Construct the scale space of the image and detect feature points that are stable at different scales. The scale space is obtained by convolving the original image with Gaussian kernels of different scales. The calculation formula is: L(x,y,σ)=G(x,y,σ)*I(x,y)

[0113] Where L(x,y,σ) is the scale space image, G(x,y,σ) is the Gaussian function, I(x,y) is the original image, * represents the convolution operation, σ is the standard deviation of the Gaussian kernel, and represents the scale;

[0114] Detect extreme points: In scale space, extreme points are determined by comparing each pixel with 26 points in its neighborhood (8 neighboring points of the current scale and 9×2 points at corresponding positions of the upper and lower scales). If a pixel has a grayscale value greater or smaller than that of all its neighboring points, the point is considered an extreme point.

[0115] Accurate key point positioning: Use Taylor expansion to fit the scale space function and determine the exact position and scale of the key points. At the same time, the contrast of the fitted function value is calculated to eliminate the key points with low contrast, and the edge response is calculated to eliminate the key points with strong edge effects.

[0116] Calculate the direction of key points: Calculate one or more directions for each key point, maintain rotation invariance, and achieve this by calculating the gradient direction and amplitude of the pixels in the neighborhood of the key point. The calculation formula is:

[0117] Gradient Amplitude:

[0118] Where m(x, y) represents the gradient amplitude of the pixel at position (x, y) in the image. The gradient amplitude reflects the intensity of the change of the image at that point, that is, how fast the brightness of the image changes at that point. L(x+1, y)-L(x-1, y) represents the difference in grayscale values ​​between the pixel at position (x, y) and its left and right adjacent pixels in the x direction. This difference reflects the rate of change of the image at that point along the x direction. L(x, y+1)-L(x, y-1) represents the difference in grayscale values ​​between the pixel at position (x, y) and its upper and lower adjacent pixels in the y direction. This difference reflects the rate of change of the image at that point along the y direction.

[0119] Gradient direction:

[0120] Wherein, θ(x, y) represents the gradient direction of the pixel at position (x, y) in the image. The gradient direction is the direction in which the image changes fastest at that point, that is, the brightness gradient direction of the image at that point. Arctan represents the inverse tangent function, which is used to calculate the angle of the gradient direction.

[0121] Then, the histogram is used to count the gradient direction and amplitude of the pixels in the neighborhood of the key point. The direction corresponding to the peak of the histogram is the main direction of the key point.

[0122] Generate feature descriptors: After determining the position, scale, and orientation of the key points, generate feature descriptors by selecting a square area (e.g., 16x16 pixels) around the key points and dividing it into smaller sub-areas (e.g., 4x4 pixels). For each sub-area, calculate the histogram of its gradient direction and magnitude, and combine these histograms into a vector as the feature descriptor of the key points.

[0123] Suppose we have a keypoint with position (x, y), scale σ, and direction θ. We select a 16x16 pixel square region as the neighborhood of the keypoint and divide it into 4x4=16 subregions. For each subregion, we calculate the histogram of its gradient direction and magnitude, and the histogram has 8 direction buckets (for example, one bucket every 45 degrees). Therefore, each subregion has an 8-dimensional vector representation. Combining these vectors, we get a 128-dimensional feature descriptor (16 subregions x 8-dimensional vector).

[0124] This feature descriptor is robust to changes in image scale, rotation, and illumination, and can be used for subsequent image matching and recognition tasks.

[0125] In one embodiment, for the above S3, in S3, the specific steps of using the YOLO algorithm to perform human detection and package detection are:

[0126] The YOLO algorithm divides the input image into multiple grid cells, and each grid predicts multiple bounding boxes and category probabilities;

[0127] Collect a dataset of images or video data containing human figures and packages to train the detection model so that the detection model can recognize human figures and packages. During the training process, the YOLO algorithm learns the features in the image and establishes a mapping relationship between the features and the target. At the same time, a loss function is defined to measure the difference between the detection results output by the model and the true label. The loss function usually includes the positioning error of the bounding box and the classification error of the category probability.

[0128] In the inference phase, the input image is fed into the trained detection model to obtain the detection results of the human figure and the package, that is, the location information of the human figure and the package. These results include the coordinates, size, and category probability of the bounding box.

[0129] The detection results of the human figure and the package are jointly post-processed. For example, the detection frame can be adjusted and optimized according to the relative position and size relationship of the human figure and the package to obtain a more accurate detection result and obtain the final detection result.

[0130] In one embodiment, for the above S4, in S4, the specific steps of video determination are:

[0131] Feature extraction and classification: Extract feature information of humanoid targets tracked in S3, such as height, body proportions, facial features, etc., and use support vector machine to classify humanoid targets as adults or children based on the extracted feature information;

[0132] Number counting and behavior analysis: Count the number of adults and children in the video frame, analyze the door opening behavior, and determine whether it is performed by a child by detecting the movement of the door handle and the opening angle of the door;

[0133] Alarm triggering: If there are only children in the picture and the door is opened by the child, an alarm will be triggered that the child is going out alone, and the alarm information will be sent to the preset parent's mobile phone or smart home control center.

[0134] In one embodiment, for the above S5, in S5, the specific steps of dynamic capture and recording are:

[0135] Trigger condition: When the smart lock peephole system recognizes that a child opens the door alone through human detection technology, the process is triggered;

[0136] Detection and tracking: When a child opens the door, the dynamic capture function is immediately activated, and the tracking algorithm is used to track the child in real time to ensure that the target is not lost;

[0137] Capture and record: As the child moves, continuous capture begins, and each frame records the child's movements. These captured images are combined into a dynamic image in real time to show the entire continuous process of the child going out;

[0138] Generate animated image: These continuous captured images are processed to generate a smooth animated image, which records the entire process from the child opening the door to leaving the monitoring area, ensuring that there is no blind spot in the monitoring;

[0139] Storage and viewing: The generated animated images are stored in the built-in storage of the smart lock or in the cloud server. Users can view these animated images through the mobile phone APP or other authorized devices to understand the specific situation of children going out.

[0140] Alarms and notifications: When a child goes out alone, the system sends an alarm notification to a preset mobile phone number or APP. The notification content includes the time and place of the child going out, as well as a link or preview of an animated picture, so that users can understand the situation in time and take appropriate measures.

[0141] In one embodiment, for the above S6, in S6, the specific steps of adaptively generating a dynamic dual screen are:

[0142] Based on the results of target detection and tracking, two images are dynamically generated;

[0143] The first screen shows the area where people move. When someone enters the area, the screen automatically switches to the image of that person.

[0144] The second screen shows the delivery area. When a delivery is placed or taken away, the screen will dynamically adjust as the number of deliveries increases or decreases. When a new delivery is placed in the monitoring area, the change is automatically detected and the display range is expanded accordingly to include the new delivery. When the delivery is taken away, the display range will automatically shrink.

[0145] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A virtual dual-camera technology for a smart lock peephole, characterized in that: The following steps are involved: S1: Select a wide-angle lens: Select a wide-angle lens, which captures all important areas in front of the door, including the ground in front of the door and the area where the express delivery is placed; S2: Correct the designated area in front of the current image: Correct the middle image to obtain a normal distorted image; S3: Human figure detection and package detection: By performing human figure detection and package detection on the corrected image, we can obtain the position frame map of children, the position frame map of adults, and the position map of express packages, track the detected targets, obtain their motion information, and adjust the human figure area so that the human face is in the upper middle area of ​​the image; S4: Video judgment: The door unlocking message is used as the cutoff point for video judgment. When the door is unlocked, judgment is performed 2 to 5 seconds before the door is unlocked to determine the number of children and adults in the picture. If there are only children in the picture and the door is opened by a child, an alarm of a child going out alone is reported; S5: Dynamic capture and recording: After detecting that a child opens the door, the entire process is captured as the person moves, and a dynamic image is generated to record the entire process of the child going out; S6: Adaptively generate dynamic dual screens: Adaptively generate dynamic dual screens according to the activity area of ​​people and the area where the express delivery is placed. The dynamic dual screens respectively display the people in front of the door and the express delivery under the door. The display range of the express delivery under the door is relatively dynamic. According to the distribution of the express delivery, the largest area is selected to cover the express delivery area. S7: Image fusion: The two images are combined into one on the camera side to form a fused image, which covers all important areas in front of the door and displays the information of the person and the courier at the same time.

2. According to claim 1, the virtual dual-camera technology of the smart lock cat eye is characterized in that: In S1, the wide-angle lens is installed above or on the side of the smart door lock to capture all important areas in front of the door, and the wide-angle lens is tilted downward to capture details on the ground.

3. According to claim 1, the virtual dual-camera technology of the smart lock cat eye is characterized in that: In S1, the wide-angle lens is installed above or on the side of the smart door lock to capture all important areas in front of the door. At the same time, the angle of the wide-angle lens is adjusted according to the actual situation in front of the door.

4. The virtual dual-camera technology of the smart lock cat eye according to claim 1 is characterized in that: In S2, the specific steps of the correction process are: Image acquisition and preprocessing: obtain the original image data of the wide-angle lens and preprocess the original image; Distortion model establishment: The wide-angle lens distortion includes barrel distortion and pincushion distortion. The barrel distortion is manifested as the center of the picture being enlarged compared to the edge, while the pincushion distortion is the opposite. According to the distortion type, a mathematical model is selected to describe it; Feature point extraction and matching: Select a set of feature points in the original image. The feature points should have geometric features. Use SIFT to extract the positions and descriptors of the feature points. Find the feature points corresponding to the feature points in the original image in an ideal distortion-free image or an image with known distortion parameters. Distortion parameter solution: Based on the matching results of feature points, the least squares method is used to solve the distortion parameters; Image correction: According to the obtained distortion parameters, the original image is corrected. The correction process involves coordinate transformation and pixel resampling. The coordinate transformation formula is: Among them, x original ,y original is the coordinate in the original image, x corrected ,y corrected are the corrected coordinates, r is the distance from the point to the image center, k1, k2, ... are the radial distortion coefficients, and p1, p2 are the tangential distortion coefficients; Post-processing: Interpolate the rectified image to obtain smooth pixel values.

5. The virtual dual-camera technology of the smart lock cat eye according to claim 1 is characterized in that: In S3, the specific steps of using the YOLO algorithm for human detection and package detection are: The YOLO algorithm divides the input image into multiple grids, each of which predicts multiple bounding boxes and class probabilities; Collect a dataset of images or videos containing human figures and packages to train the detection model so that the detection model can recognize human figures and packages. During the training process, the YOLO algorithm learns the features in the image and establishes a mapping relationship between the features and the target. At the same time, a loss function is defined to measure the difference between the detection results output by the model and the true label. In the inference phase, the input image is fed into the trained detection model to obtain the detection results of the human figure and the package, that is, the location information of the human figure and the package; The detection results of the human figure and the package are jointly post-processed to obtain the final detection result.

6. The virtual dual-camera technology of the smart lock cat eye according to claim 1, characterized in that: In S4, the specific steps of video judgment are: Feature extraction and classification: Extract feature information of the humanoid target tracked in S3. Using support vector machine to classify humanoid targets into adults or children based on the extracted feature information; Number counting and behavior analysis: Count the number of adults and children in the video frame, analyze the door opening behavior, and determine whether it is performed by a child by detecting the movement of the door handle and the opening angle of the door; Alarm triggering: If there are only children in the picture and the door is opened by the child, an alarm will be triggered that the child is going out alone, and the alarm information will be sent to the preset parent's mobile phone or smart home control center.

7. The virtual dual-camera technology of the smart lock cat eye according to claim 1, characterized in that: In S5, the specific steps of dynamic capture and recording are: Trigger condition: When the smart lock peephole system recognizes that a child opens the door alone through human detection technology, the process is triggered; Detection and tracking: When a child opens the door, the dynamic capture function is immediately activated to track the child in real time; Capture and record: As the child moves, continuous capture begins, and each frame records the child's movements; Generate a dynamic image: These continuous captured images are processed to generate a smooth dynamic image, which records the entire process from the child opening the door to leaving the monitoring area; Storage and viewing: The generated animated images are stored in the built-in storage of the smart lock or in the cloud server. Users can view these animated images through the mobile phone APP or other authorized devices to understand the specific situation of children going out. Alarm and notification: When a child goes out alone, the system sends an alarm notification to a preset mobile phone number or APP. The notification content includes the time and place of the child going out, as well as a link or preview of an animated picture.

8. The virtual dual-camera technology of the smart lock cat eye according to claim 1, characterized in that: In S6, the specific steps of adaptively generating dynamic dual screens are: Based on the results of target detection and tracking, two images are dynamically generated; The first screen shows the area where people move. When someone enters the area, the screen automatically switches to the image of that person. The second screen shows the delivery area. When a delivery is placed or taken away, the screen will dynamically adjust as the number of deliveries increases or decreases. When a new delivery is placed in the monitoring area, the change is automatically detected and the display range is expanded accordingly to include the new delivery. When the delivery is taken away, the display range will automatically shrink.

Citation Information

Patent Citations

  • Multi-angle cat eye intelligent door lock

    CN219262029U