Focusing method, focusing device and storage medium
By obtaining the camera preview image in real time, using motion detection and human body semantic segmentation algorithms to extract the target object pixels, determine the target image area, and set the focus mode according to the positional relationship, the problem that the camera cannot focus on the object to be displayed is solved, and automatic switching and clear imaging of the focus mode are realized.
Patent Information
- Application Number
- CN202410064789.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art cannot effectively realize the camera focusing on the object to be displayed instead of the whole human body, and the method based on infrared ranging is not very accurate, so the function of focusing on the object to be displayed in live broadcast cannot be realized.
By acquiring the camera preview image in real time, using motion detection and human body semantic segmentation algorithms to extract the target object pixels, determine the target image area, and set the focus mode according to the positional relationship between the target image area and the preset image area, so as to realize automatic switching of focusing on the target object.
In live broadcast, smooth switching between focus products and default focus methods is achieved, improving the recording experience, ensuring that the objects to be displayed are clearly imaged, and the background is blurred, meeting the needs of live broadcast goods and display.
Smart Images

Figure CN120343396A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of camera, and particularly to a focusing method, a focusing device, and a storage medium. Background Art
[0002] With the development of industries such as live product recommendation, some users have higher requirements for the camera function of the terminal. When it is necessary to show some products to the users, it is required to make the foreground object form a clear image, and it is required that the terminal camera focuses on the object to be shown after the object to be shown enters the display area, and the terminal camera no longer focuses on the object to be shown after the object to be shown enters the display area.
[0003] In related technologies, there is a technical solution that based on image semantic segmentation technology and target tracking technology to track the human body to determine whether it exceeds the area, and then guide the camera to rotate. This solution can only make the camera focus on the whole person. In related technologies, there is also a technology that based on infrared ranging technology to detect the movement of people and perform focusing. This solution has low accuracy and cannot focus on the object to be shown. Summary of the Invention
[0004] To overcome the problems existing in the related technologies, the present disclosure provides a focusing method, a focusing device, and a storage medium.
[0005] According to the first aspect of the embodiments of the present disclosure, a focusing method is provided, including: obtaining a preview image currently captured by a camera in real time, and determining a target image area in the preview image, where the target image area is an image area corresponding to a target object in the preview image; setting a current focusing mode according to a positional relationship between the target image area and a preset image area, where the preset image area is located in a central area of the preview image.
[0006] In an implementation, the preview image includes a human image, a target object image, and a background image; the determining the target image area in the preview image includes: obtaining a set of moving pixels in the target image area according to a motion detection algorithm, where the set of moving pixels corresponds to the human image and the target object image; obtaining a set of target object pixels in the set of moving pixels according to a human body semantic segmentation algorithm, where the set of target object pixels corresponds to the target object image; determining the target image area according to pixel coordinates corresponding to the set of target object pixels in the preview image.
[0007] In one implementation, determining the target image area according to the pixel coordinates corresponding to the target object pixel set in the preview image includes: determining a first coordinate, a second coordinate, a third coordinate, and a fourth coordinate in the target object pixel set, where the first coordinate is the pixel coordinate closest to the upper boundary of the preview image corresponding to the target object pixel set, the second coordinate is the pixel coordinate closest to the lower boundary of the preview image corresponding to the target object pixel set, the third coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the target object pixel set, and the fourth coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the target object pixel set; determining a first line parallel to the upper and lower boundaries of the preview image and passing through the first coordinate in the preview image, determining a second line parallel to the upper and lower boundaries of the preview image and passing through the second coordinate, determining a third line parallel to the left and right boundaries of the preview image and passing through the third coordinate, and determining a fourth line parallel to the upper and lower boundaries of the preview image and passing through the fourth coordinate; and determining the rectangular area corresponding to the first line, the second line, the third line, and the fourth line in the preview image as the target image area.
[0008] In one implementation, setting the current focusing mode according to the positional relationship between the target image area and the preset image area includes: determining the area of the intersection region between the target image area and the preset image area, and determining the area of the union region between the target image area and the preset image area; determining the current ratio between the area of the intersection region and the area of the union region, and setting the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio, where the preset ratio is a critical ratio for focusing mode adjustment.
[0009] In one implementation, setting the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio includes: in response to the current ratio being greater than or equal to the preset ratio, setting the current focusing mode to the first mode, in which the focusing area of the camera corresponds to the target image area; in response to the current ratio being less than the preset ratio, setting the current focusing mode to the second mode, in which the focusing area of the camera corresponds to the face image area in the preview image or the central area of the preview image.
[0010] In one implementation, in the first mode, the following method is used to make the focusing area of the camera correspond to the target image area: obtaining the coordinate set corresponding to the boundary of the target image area, and controlling the camera lens module according to the coordinate set so that the camera focus falls on the target object corresponding to the target image area.
[0011] In one implementation, the method further includes: in response to the current ratio being greater than or equal to a preset ratio, setting label information for the object image enclosed by the target image region, where the label information is used to identify the correspondence between the object image and the object.
[0012] According to a second aspect of the embodiments of the present disclosure, a focusing device is provided, including: a determining unit, configured to obtain in real time a preview image currently captured by a camera and determine a target image region in the preview image, where the target image region is an image region corresponding to an object in the preview image; a processing unit, configured to set a current focusing mode according to a positional relationship between the target image region and a preset image region, where the preset image region is located in a central region of the preview image.
[0013] In one implementation, the preview image includes a person image, an object image, and a background image; the determining unit determines the target image region in the preview image in the following manner: obtaining a set of moving pixels in the target image region according to a motion detection algorithm, where the set of moving pixels corresponds to the person image and the object image; obtaining a set of object pixels in the set of moving pixels according to a human semantic segmentation algorithm, where the set of object pixels corresponds to the object image; and determining the target image region according to pixel coordinates corresponding to the set of object pixels in the preview image.
[0014] In one implementation, the determining unit determines the target image region according to pixel coordinates corresponding to the set of object pixels in the preview image in the following manner: determining a first coordinate, a second coordinate, a third coordinate, and a fourth coordinate in the set of object pixels, where the first coordinate is the pixel coordinate closest to the upper boundary of the preview image corresponding to the set of object pixels, the second coordinate is the pixel coordinate closest to the lower boundary of the preview image corresponding to the set of object pixels, the third coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the set of object pixels, and the fourth coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the set of object pixels; determining a first straight line parallel to the upper and lower boundaries of the preview image and passing through the first coordinate in the preview image, determining a second straight line parallel to the upper and lower boundaries of the preview image and passing through the second coordinate, determining a third straight line parallel to the left and right boundaries of the preview image and passing through the third coordinate, and determining a fourth straight line parallel to the upper and lower boundaries of the preview image and passing through the fourth coordinate; and determining the rectangular region corresponding to the first straight line, the second straight line, the third straight line, and the fourth straight line in the preview image as the target image region.
[0015] In one implementation, the processing unit sets the current focusing mode according to the positional relationship between the target image area and the preset image area in the following manner, including: determining the area of the intersection area between the target image area and the preset image area, and determining the area of the union area between the target image area and the preset image area;
[0016] determining the current ratio between the area of the intersection area and the area of the union area, and setting the current focusing mode according to the magnitude relationship between the current ratio and a preset ratio, where the preset ratio is a critical ratio for focusing mode adjustment.
[0017] In one implementation, the processing unit sets the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio in the following manner: in response to the current ratio being greater than or equal to the preset ratio, setting the current focusing mode to the first mode, in which the focusing area of the camera corresponds to the target image area; in response to the current ratio being less than the preset ratio, setting the current focusing mode to the second mode, in which the focusing area of the camera corresponds to the face image area in the preview image or the central area of the preview image.
[0018] In one implementation, in the first mode, the processing unit makes the focusing area of the camera correspond to the target image area in the following manner: obtaining the coordinate set corresponding to the boundary of the target image area, and controlling the camera lens module according to the coordinate set so that the camera focus falls on the target corresponding to the target image area.
[0019] In one implementation, the processing unit is further configured to: in response to the current ratio being greater than or equal to the preset ratio, set label information for the target object image enclosed by the target image area, where the label information is used to identify the corresponding relationship between the target object image and the target object.
[0020] According to a third aspect of the embodiments of the present disclosure, a focusing device is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to: execute the focusing method described in the first aspect or any one of the implementations of the first aspect.
[0021] According to a fourth aspect of the embodiments of the present disclosure, a storage medium is provided, in which instructions are stored, and when the instructions in the storage medium are executed by a processor, the processor is enabled to execute the focusing method described in the first aspect or any one of the implementations of the first aspect.
[0022] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: obtaining in real time the target image area corresponding to the target object in the current preview image of the camera, and setting the current focusing mode according to the positional relationship between the target image area and the preset image area, where the preset image area is located in the central area of the preview image. Through the present disclosure, automatic and smooth switching between focusing on the product and the default focusing method can be achieved in the item display scenario.
[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0025] Figure 1A and Figure 1B are schematic diagrams of a shooting scenario shown according to an exemplary embodiment of the present disclosure.
[0026] Figure 2 Schematic diagram of the structure of a focusing lens shown according to an exemplary embodiment
[0027] Figure 3 is a flowchart of a focusing method shown according to an exemplary embodiment.
[0028] Figure 4 is a schematic diagram of a preset image area shown according to an exemplary embodiment.
[0029] Figure 5 is a flowchart of a method for determining the target image area in the preview image shown according to an exemplary embodiment.
[0030] Figure 6 is a schematic diagram of obtaining a set of motion pixels based on a motion detection algorithm shown according to an exemplary embodiment of the present disclosure.
[0031] Figure 7 is a schematic diagram of obtaining a set of pixels corresponding to the human body based on a human semantic segmentation algorithm shown according to an exemplary embodiment of the present disclosure.
[0032] Figure 8 is a flowchart of a method for determining the target image area shown according to an exemplary embodiment.
[0033] Figure 9 is a flowchart of a method for setting the current focusing mode shown according to an exemplary embodiment.
[0034] Figure 10It is a flowchart of a method for setting a current focusing mode shown according to an exemplary embodiment.
[0035] Figure 11 It is a flowchart of a method for making a focusing area of a camera correspond to a target image area shown according to an exemplary embodiment.
[0036] Figure 12 It is a schematic structural diagram of a camera lens module shown according to an exemplary embodiment of the present disclosure.
[0037] Figures 13A to 13D It is a schematic diagram of a camera imaging optical principle shown according to an exemplary embodiment of the present disclosure.
[0038] Figure 14 It is a schematic diagram of the relationship between camera imaging and depth of field shown according to an exemplary embodiment of the present disclosure.
[0039] Figure 15 It is a schematic diagram of a camera imaging principle shown according to an exemplary embodiment of the present disclosure.
[0040] Figure 16A and Figure 16B A schematic diagram of a camera imaging principle shown according to an exemplary embodiment of the present disclosure.
[0041] Figure 17 It is a schematic diagram of a camera imaging principle shown according to an exemplary embodiment of the present disclosure.
[0042] Figure 18 It is a schematic diagram of a phase parameter calculation method shown according to an exemplary embodiment of the present disclosure.
[0043] Figure 19 It is a flowchart of a focusing method shown according to an exemplary embodiment.
[0044] Figure 20 It is a flowchart of a focusing method shown according to an exemplary embodiment of the present disclosure.
[0045] Figure 21 It is a block diagram of a focusing device shown according to an exemplary embodiment.
[0046] Figure 22 It is a block diagram of a device for a focusing device shown according to an exemplary embodiment. Detailed implementation manners
[0047] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure.
[0048] The focusing method provided in the embodiments of the present disclosure is applied to the scenario of live recommending items, and controls the camera to switch between two modes: focusing on the item and not focusing on the item.
[0049] The focusing method provided in the embodiments of the present disclosure is mainly applied to the field of autofocus for terminal camera preview during shooting and video recording, and can be extended to the field of camera imaging technology.
[0050] With the development of live broadcast industries that require item display, such as live e-commerce, live unboxing, and live item recommendation, higher requirements are put forward for the terminal video recording function. Especially when the live broadcaster needs to show some products to users, it is necessary to make the displayed products clearly imaged when they are in the display area, and the camera no longer focuses on the products after they leave the display area. For example Figure 1A and Figure 1B As shown in the schematic diagram of the shooting scene in, the black rectangular frame in the figure is the display area. After the product enters the rectangular frame, the camera needs to focus on the product to clearly display the product; after the product leaves the rectangular frame, the camera no longer focuses on the product and adopts the default focusing mode (focusing on the center of the preview image or focusing on the face). In summary, when the above-mentioned live broadcast industries that require item display are recommending product displays, the focusing mode of the camera is switched between the default focusing mode and the mode of focusing on the product. That is, it is required that when the object to be displayed (product) enters the display area, the focusing area is on the item, and when the object to be displayed (product) leaves the display area, the default mode is adopted.
[0051] In the related art, there is a camera control method based on image semantic segmentation technology and target tracking technology, that is, using image semantic segmentation technology and target tracking technology to perform human body tracking, judging whether the human body exceeds a pre-set area, and guiding the rotation of the camera based on the relationship between the area where the human body is located and the pre-set area to achieve tracking of the user. However, using semantic segmentation and human body tracking technology can only ensure real-time acquisition of the human body position and tracking of the human body. The main purpose of the camera control method in the above-mentioned related art is to move the camera so that the person being photographed is still in the state of being photographed during the moving state. The above camera control method cannot achieve autofocus, and it is even more impossible to separately obtain the area of the product to be displayed and focus on the product to be displayed. It can only make the camera focus on the whole person.
[0052] In the related art, there is also a focusing method based on infrared ranging. An infrared rangefinder is used to detect the movement of the human body. After detecting the movement of the human body, the corresponding information is transmitted to a micro programmable controller. The programmable controller controls the movement of the mobile device based on the corresponding information. After the mobile device moves, it drives the transmission part to move, and then drives the focusing gear to rotate. After the focusing gear rotates, it focuses on the imaging picture inside the lens. However, the accuracy of using infrared ranging is not as good as that of images, and the function of product display focusing cannot be performed either.
[0053] In the traditional focusing lens scheme in the related art, the entire lens is pushed by an external motor to achieve focusing. As Figure 2 shown in the structural schematic diagram of the focusing lens, the overall structure of the traditional focusing lens includes: a lens, a voice coil motor (VCM), a base bracket, a photosensitive chip (Sensor), a driver chip (Driver IC), and an output interface. The entire lens of the traditional focusing lens is connected to the voice coil motor, and the voice coil motor with the lens is installed on the magnetic module bracket. When focusing, the voice coil is energized, and the energized coil moves under the action of the magnetic field, and then drives the lens to move on the module bracket to achieve focusing. The general characteristics of the above lens are large length, the field of view (FOV) changes during the focusing process, and the closest focusing distance is difficult to reach the centimeter level and cannot focus on close objects.
[0054] In view of this, the present disclosure proposes a focusing method. When a user carries a product for display, moving pixels including human pixels and product pixels in the preview image are extracted. The human pixels included in the moving pixels are extracted and the human pixels are excluded to obtain product pixels. A tracking frame surrounding the product pixels is obtained, the tracking frame surrounding the product pixels is continuously tracked, and it is determined whether an object enters the display area. When the tracking frame enters the display area, focus on the object, and when it leaves the display area, restore the default focusing distance. Through the present disclosure, smooth switching between focusing on the product and the default focusing method can be achieved in the item display scenario.
[0055] Figure 3 is a flowchart of a focusing method shown according to an exemplary embodiment. As Figure 3 shown, the method includes steps S101 to step S102.
[0056] In step S101, the preview image currently captured by the camera is obtained in real time, and the target image area in the preview image is determined.
[0057] Among them, the target image area is the image area corresponding to the target object in the preview image.
[0058] In step S102, the current focusing mode is set according to the positional relationship between the target image area and the preset image area.
[0059] Among them, the preset image area is located in the central area of the preview image.
[0060] In the embodiment of the present disclosure, in the camera shooting scenario where the user carries the product to be displayed (the target object) for display, the preview image collected by the camera is continuously obtained. It can be understood that the preview image mainly includes a human image (an image of a part of the human body), a product image, and a background image. After the present disclosure obtains the target image area corresponding to the product image from the preview image, the current focusing mode is set according to the position of the product image in the entire preview image. It can be understood that when the user needs to display the product, the product to be displayed will be located in the central area of the preview image. Therefore, the present disclosure sets a preset image area at the central position of the preview image, and then sets the current focusing mode according to the positional relationship between the preset image area and the above-mentioned target object area.
[0061] In an exemplary embodiment of the present disclosure, as Figure 4 shown in the schematic diagram of the preset image area in, the present disclosure sets the preset image area (the display area of the product) in the preview image in the following manner, making the aspect ratio of the preset image area consistent with the aspect ratio of the preview image, and making the preset image area located in the central area of the preview image, and making the preset image area proportionally reduced compared to the preview image border, such as making the length ratio of the preset image area to the preview image border be 5.5:10 or 3.3:10. The user can select a preset image area corresponding to the product size from a variety of preset image area setting strategies based on the product to be displayed before the product display. It can be understood that the user can customize more types of preview image area configurations in terms of size, shape, etc. based on personal needs.
[0062] In the embodiment of the present disclosure, when the user carries the product (the target object) for display, the target image area corresponding to the product is obtained. According to the positional relationship between the target image area and the preset image area in the preview image, the current focusing mode is set. Through the present disclosure, setting the focusing mode based on the position area of the target object in the preview image in the article display scenario can achieve automatic and smooth switching between focusing on the product and the default focusing method.
[0063] In the embodiment of the present disclosure, when the user carries the product (the target object) for display, the preview image includes a background image, a human body image (human image), and a product image. The present disclosure can separate the product image from the preview image and determine the target image area corresponding to the product image based on the differences in the pixels corresponding to the above three images. The following embodiments of the present disclosure illustrate the method for determining the target image area in the preview image.
[0064] Figure 5 is a flowchart of a method for determining the target image area in the preview image shown according to an exemplary embodiment. As Figure 5As shown, the method includes steps S201 to S203.
[0065] In step S201, according to the motion detection algorithm, a set of motion pixels in the target image region is obtained, and the set of motion pixels corresponds to the human image and the target object image.
[0066] In step S202, according to the human semantic segmentation algorithm, a set of target object pixels in the set of motion pixels is obtained, and the set of target object pixels corresponds to the target object image.
[0067] In step S203, according to the pixel coordinates corresponding to the set of target object pixels in the preview image, the target image region is determined.
[0068] In the embodiment of the present disclosure, the preview image includes a human image, a target object image, and a background image. The pixels corresponding to the background image basically do not change much and can be regarded as non-moving pixels. Compared with the pixels corresponding to the background object, the pixels corresponding to the human image and the target object image are continuously in a changing state and can be regarded as moving pixels. Based on the difference in the changing state between the above-mentioned moving pixels and non-moving pixels, the present disclosure obtains the moving pixels in the preview image through a motion detection algorithm (such as background modeling method, frame difference method, optical flow method, etc.), that is, the pixels corresponding to the human image and the target object image.
[0069] In an exemplary embodiment of the present disclosure, the background modeling method (Background subtraction, BGS, that is, background difference method) is used to obtain the moving pixels in the preview image. As Figure 6 shown in the schematic diagram of obtaining the set of motion pixels based on the motion detection algorithm, the principle of obtaining the motion pixels using the background modeling method is: a background model is established in advance, that is, a background model that does not include humans and moving target objects. In practical applications, the difference is taken between the currently acquired current frame of the image (that is, the current preview image) and the pre-established background model. When the difference is greater than the threshold (THRESHOLD) (in view of the slow change of the image background, a threshold is set when taking the difference), the pixels with a difference greater than the threshold are determined as motion pixels, so as to obtain the pixels corresponding to the motion image, and the foreground mask corresponding to the motion pixels is obtained synchronously.
[0070] In the embodiment of the present disclosure, the obtained motion pixels include the pixels corresponding to the human image and the pixels corresponding to the target object image. The present disclosure further processes the motion pixels using algorithms such as human semantic segmentation, such as Figure 7As shown in the schematic diagram of obtaining the human body corresponding pixel set based on the human body semantic segmentation algorithm, the pixels corresponding to the person in the image can be obtained through the human body semantic segmentation algorithm. In the present disclosure, the human body pixels in the preview image are obtained through the human body semantic segmentation algorithm, the human body pixels in the motion pixels are deleted, and the pixels corresponding to the target object included in the motion pixels are obtained. Furthermore, the noise in the pixels corresponding to the target object is removed through image morphology processing algorithms such as erosion and dilation to obtain the target object pixels. It can be understood that if the above human body segmentation process needs to use machine learning and deep learning related technologies, then the requirement is that the object pixels are not classified into the human body category when segmenting.
[0071] In the embodiments of the present disclosure, the target image area is delimited by a rectangular frame surrounding the target object pixels. The rectangular frame is the smallest bounding box surrounding the target object pixels, and it can be regarded as a tracking frame for the target object. The present disclosure determines the above rectangular frame according to the spatial coordinates of the target object pixels in the preview image and determines the target object area. The following embodiments of the present disclosure illustrate the method for determining the target image area.
[0072] Figure 8 is a flowchart of a method for determining a target image area shown according to an exemplary embodiment. As Figure 8 shown, the method includes steps S301 to S303.
[0073] In step S301, the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate in the target object pixel set are determined.
[0074] Among them, the first coordinate is the pixel coordinate of the target object pixel set closest to the upper boundary of the preview image, the second coordinate is the pixel coordinate of the target object pixel set closest to the lower boundary of the preview image, the third coordinate is the pixel coordinate of the target object pixel set closest to the left boundary of the preview image, and the fourth coordinate is the pixel coordinate of the target object pixel set closest to the left boundary of the preview image.
[0075] In step S302, in the preview image, a first straight line parallel to the upper and lower boundaries of the preview image and passing through the first coordinate is determined, a second straight line parallel to the upper and lower boundaries of the preview image and passing through the second coordinate is determined, a third straight line parallel to the left and right boundaries of the preview image and passing through the third coordinate is determined, and a fourth straight line parallel to the upper and lower boundaries of the preview image and passing through the fourth coordinate is determined.
[0076] In step S303, the rectangular area in the preview image corresponding to the first straight line, the second straight line, the third straight line, and the fourth straight line is determined as the target image area.
[0077] In the embodiments of the present disclosure, after determining the set of target object pixels corresponding to the target object image in the preview image, the tracking frame surrounding the target object pixels and the target object area corresponding to the target object are determined according to the spatial coordinates of the pixels in the pixel set: determine the pixel coordinates corresponding to the uppermost pixel (the pixel coordinate closest to the upper boundary of the preview image), the lowermost pixel (the pixel coordinate closest to the lower boundary of the preview image), the leftmost pixel (the pixel coordinate closest to the left boundary of the preview image), and the rightmost pixel (the pixel coordinate closest to the right boundary of the preview image) in the pixel set, that is, the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate. Further, determine the straight lines passing through the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate and parallel to the boundaries of the preview image, that is, the first straight line parallel to the upper and lower boundaries of the preview image and passing through the first coordinate, the second straight line parallel to the upper and lower boundaries of the preview image and passing through the second coordinate, the third straight line parallel to the left and right boundaries of the preview image and passing through the third coordinate, and the fourth straight line parallel to the upper and lower boundaries of the preview image and passing through the fourth coordinate. The area enclosed by the above four straight lines is determined as the target object area, and the border of the target object area can be regarded as the tracking frame for the target object.
[0078] In the embodiments of the present disclosure, the tracking frame surrounding the target object image is determined based on the spatial coordinates of the pixels in the set of target object pixels corresponding to the target object, and then the target image area corresponding to the target object is obtained, which is convenient for automatically switching the focus mode according to the spatial position of the target image area in the preview image in the subsequent process.
[0079] In the embodiments of the present disclosure, when the user needs to display the product, the product to be displayed is placed in the central area of the preview image. Therefore, the present disclosure sets a preset image area at the center position of the preview image, and then sets the current focus mode according to the positional relationship between the preset image area and the above target object area. The following embodiments of the present disclosure illustrate the method for setting the current focus mode.
[0080] Figure 9 is a flowchart of a method for setting the current focus mode shown according to an exemplary embodiment. As Figure 9 shown, the method includes steps S401 to S402.
[0081] In step S401, determine the area of the intersection region between the target image area and the preset image area, and determine the area of the union region between the target image area and the preset image area.
[0082] In step S402, determine the current ratio between the area of the intersection region and the area of the union region, and set the current focus mode according to the magnitude relationship between the current ratio and the preset ratio.
[0083] Wherein, the preset ratio is a critical ratio for focus mode adjustment.
[0084] In the embodiments of the present disclosure, the preset image area is located in the central area of the preview image. When the user displays the target object, the target object image will move closer to the central area of the preview image. In this case, the preset image area and the target image area corresponding to the target object will overlap, and the higher the degree of overlap, the closer the target object image is to the center of the preview image. In this case, the intersection-over-union ratio of the target image area and the preset image area (i.e., the ratio of the area of the intersection region of the two to the area of the union region) can reflect the degree of coincidence between the preset image area and the target image area.
[0085] It can be understood that when the user displays the target object, the image corresponding to the target object often does not completely lie in the center of the preview image, but only makes the image corresponding to the target object close to the center of the preview image. Based on this, the present disclosure pre-sets a critical ratio for focus mode adjustment, that is, the preset ratio. The current focus mode of the camera is set based on the magnitude relationship between the current ratio and the preset ratio.
[0086] The following embodiments of the present disclosure further illustrate the method for setting the current focus mode.
[0087] Figure 10 is a flowchart of a method for setting the current focus mode shown according to an exemplary embodiment. As Figure 9 shown, the method includes step S501, step S502A, and step S502B.
[0088] In step S501, determine the current ratio between the area of the intersection region and the area of the union region.
[0089] In step S502A, in response to the current ratio being greater than or equal to the preset ratio, set the current focus mode to the first mode. In the first mode, the focus area of the camera corresponds to the target image area.
[0090] In step S502B, in response to the current ratio being less than the preset ratio, set the current focus mode to the second mode. In the second mode, the focus area of the camera corresponds to the face image area in the preview image or the central area of the preview image.
[0091] In an embodiment of the present disclosure, after determining the current ratio between the area of the intersection region and the area of the union region, if the current ratio is greater than or equal to a preset ratio, it indicates that the coincidence degree between the preset image region and the target image region reaches the coincidence degree threshold, indicating that the target object image is close to the central region of the preview image, indicating that the user needs to display the target object. At this time, the current focus mode of the camera is set to the focus mode (the first mode) for focusing on the target object to ensure that the target object image in the preview image is clear. In the above first mode, the target tracking algorithm is used to perform real-time tracking on the detection frame of the displayed item to ensure that the camera focuses on the target object and the target object image is clearly displayed when the spatial position of the target object changes and the position of the target object image in the preview image changes.
[0092] In an embodiment of the present disclosure, after determining the current ratio between the area of the intersection region and the area of the union region, if the current ratio is greater than or equal to a preset ratio, it indicates that the coincidence degree between the preset image region and the target image region does not meet the coincidence degree threshold, indicating that the target object image is not in the central region of the preview image, indicating that the user no longer needs to display the target object. At this time, the default focus mode is adopted to make the camera focus on the region corresponding to the center of the preview image or to make the camera focus on the user's face.
[0093] In an embodiment of the present disclosure, after determining the current ratio between the area of the intersection region and the area of the union region, the focus mode is switched based on the size relationship between the current ratio and the preset ratio, realizing the automatic smooth switching between the camera focusing on the product and the default focus mode.
[0094] In an embodiment of the present disclosure, the camera focus area is adjusted based on phase detection autofocus (PDAF). The focusing principle of the phase focusing method is mainly to reserve some masked pixel points on the photosensitive element as an autofocus sensor to obtain phase data, and then through a distance mapping model, convert the phase data into corresponding focusing parameters such as depth of field (defocus), and control the lens module to the optimal focusing position based on the obtained focusing parameters to achieve autofocus. The following embodiments of the present disclosure illustrate the method for making the focus area of the camera correspond to the target image area.
[0095] Figure 11 is a flowchart of a method for making the focus area of the camera correspond to the target image area shown according to an exemplary embodiment. As Figure 11 shown, the method includes step S601 and step S602.
[0096] In step S601, in response to the current ratio being greater than or equal to the preset ratio, a coordinate set corresponding to the boundary of the target image area is obtained.
[0097] In step S602, the camera lens module is controlled according to the coordinate set so that the camera focus falls on the target object corresponding to the target image area.
[0098] In the embodiments of the present disclosure, as Figure 12 is a schematic structural diagram of the camera lens module in, the focusing lens solution adopted by the present disclosure is different from the traditional focusing lens solution. The camera lens module of the camera in the present disclosure includes structures such as a voice coil motor VCM, a fixed group lens set G1, and an autofocus moving group lens set G2.
[0099] In an exemplary embodiment of the present disclosure, as Figures 13A to 13D shown in the schematic diagram of the camera imaging optical principle. Figure 13B Characterizes that the camera is in a focused state for the photographed object, that is, the light rays entering from above and below coincide at a point, Figure 13A the distance between the lens and the sensor is greater than the focusing distance, and the light rays entering from above and below do not coincide at a point. Figure 13C and Figure 13D the distance between the lens and the sensor is less than the focusing distance, and the light rays entering from above and below coincide at a point. Figure 13C and Figure 13D comparison, Figure 13D in which the distance between the lens and the sensor is smaller, and the spacing of the light ray landing points entering from above and below is larger. The spacing of the light ray landing points on the sensor entering from above and below can not only reflect the magnitude of defocus, but also the direction of defocus. The present disclosure controls the autofocus moving group in the camera lens module so that the distance between the lens and the sensor conforms to the focusing distance to ensure clear imaging.
[0100] In an exemplary embodiment of the present disclosure, as Figure 14 is a schematic diagram of the relationship between camera imaging and depth of field. The range of the front and rear distances of the photographed object measured by the front edge of the camera lens or other imager to obtain a clear image (the depth of the object space that can form a clear image on the plane). There is a range from a certain place in front of the focal plane to a certain place behind it, within which the scenes can all form clear images. This range is called the depth of field, such as a narrow depth of field or a large depth of field. Combining Figure 15 the schematic diagram of the camera imaging principle in, and Figure 16A and Figure 16BAs shown in the schematic diagram of the camera imaging principle. When the camera is taking pictures, if the distance between the lens and the sensor is constant, there is actually only one plane in the imaging system that is truly in focus, and this plane is the focal plane. For a certain point on an object in the focal plane, the light rays emitted at different angles pass through the lens and converge at a single point on the image plane (film or sensor plane), resulting in the clearest image. For a point on a plane that is not in focus, the light rays emitted at different angles strike different points on the image plane, forming a circle of confusion, and the closer the lens is to the focal plane, the smaller the circle of confusion. When the circle of confusion is smaller than a certain size (the allowable circle of confusion), the human eye cannot distinguish it. Therefore, there is a certain distance in front of and behind the focal plane where the image on the image plane is relatively clear and within the range of the allowable circle of confusion, forming the depth of field. For two adjacent object points in the focal plane, their images each converge at a single point and can be clearly distinguished. However, for two adjacent object points on a plane that is not in focus, after each forming a circle of confusion and then overlapping, they are not easily distinguishable. The farther the plane that is not in focus is from the focal plane, the larger the circle of confusion formed, and the more difficult it is to distinguish the overlapping images of different image points. When it exceeds the human eye's allowable resolution (beyond the range defined by the allowable circle of confusion), it is considered indistinguishable. In the depth direction in front of and behind the object being photographed (the focus point or the focal plane), there is a certain distance that is within the range defined by the allowable circle of confusion, and the distance between them is called the depth of field.
[0101] In an exemplary embodiment of the present disclosure, as Figure 17 shown in the schematic diagram of the camera imaging principle, the tolerable spot size (diameter of the circle of confusion) on the image plane is c, and for a plane at a distance D from the lens F and a plane at a distance D from the lens N where the spot size formed on the image plane is exactly c, then D F -D N is the depth of field (DOF) of the camera. In the figure, s is the focused object distance, V is the corresponding focused distance (the distance between the lens and the image plane) for the focused object distance s, V F is the focused distance corresponding to the focused object distance D F , and V N is the focused distance corresponding to the focused object distance D N . According to the Gaussian formula, the imaging formula for an object in front of the lens can be obtained:
[0102]
[0103]
[0104] Among them, c is the allowable circle of confusion diameter, F is the f-number, f is the lens focal length, and s is the focused object distance. Based on the above formula, the following rules can be obtained: the larger the aperture, the smaller the depth of field; the smaller the aperture, the larger the depth of field; the longer the lens focal length, the smaller the depth of field; the shorter the focal length, the larger the depth of field; the farther the distance, the larger the depth of field; the closer the distance, the smaller the depth of field. In the embodiments of the present disclosure, in the scenario where the user displays the target object, the following requirements exist for the camera used for real-time shooting: Since the distance between the target object and the lens is relatively close when displaying the target object, the minimum focusing distance of the lens needs to be small enough, reaching the centimeter level. During the stage when the user introduces the target object and the stage when the user displays the target object, the focusing area will switch between the background and the foreground object (the target object), and the lens will have a large focusing movement. The present disclosure requires that there is no defocusing during the focusing process, and the preview field of view angle of the real-time preview image remains stable. When the focusing area is an item, to highlight the displayed product, it is required to have the effect that the foreground object (the target object) is clear and the background is blurred, so as to produce an impact visually. It is required that the depth of field of the lens is small. According to the principle of depth of field, when the aperture is fixed and the object distance is small, the larger the focal length, the smaller the depth of field. Therefore, the present disclosure requires a larger lens focal length. To meet the above requirements, the in-camera focusing in the present disclosure is designed with a two-group lens assembly, which can achieve precise focusing by moving the group lenses inside the lens structure. The characteristics are that the focusing range is large, reaching the centimeter level at least (such as 10 cm), the focusing speed is faster, the preview field of view angle does not change during the focusing process, and the requirement of a small depth of field can be achieved when the object distance is small.
[0105] In the embodiments of the present disclosure, when the focusing mode is to focus on the target object, the coordinate set corresponding to the boundary of the target image area (the target object tracking frame) is obtained in real time, and the phase data for phase focusing is obtained based on the coordinate set corresponding to the boundary of the target image area, and the movement of the autofocus moving group lens assembly in the above camera lens module is controlled, so that the camera focus falls on the target object corresponding to the target image area, making the target object image in the preview image clear and the person image and the background image blurred, thereby highlighting the target object image.
[0106] In the embodiments of the present disclosure, when performing phase focusing of the camera, the phase (Phase Detection, PD) parameter for zooming needs to be obtained. In one example, as Figure 18 As shown in the schematic diagram of the phase parameter calculation method, the present disclosure uses the sum of difference (SAD) algorithm to obtain the phase parameter. The grid image corresponding to the preview image is obtained, and M*N phase results P1 are calculated on the obtained grid image. The calculation formula is: Wherein, d is the parallax value, D is the depth value, z is the focusing distance, α is a constant coefficient, L represents the lens aperture size, and f represents the lens focal length. When performing phase calculation, a window of size n*n is selected to determine the maximum parallax dmax; windows are constructed centered at (x, y) on the left window image and the right window image. On the left window image and the right window image, windows are respectively constructed at (x, y - d) ((x, y + d)), where d = 0 - Dmax and (x, y + d), where d = 0 - Dmax, and the pixel brightness difference within the left and right image windows within the range of parallax 0 - Dmax is calculated; the brightness differences within the window are summed, and the d that minimizes the sum of the absolute brightness differences is selected as the parallax value at (x, y). When the optical module is determined, the constant coefficient α, the lens aperture size L, and the lens focal length f are fixed values, and when the camera motor is not rotating, the focusing distance z is also fixed. Therefore, the parallax d and the depth value D have an inverse proportional relationship. Thus, the above phase calculation result can be considered as the parallax, which also represents the depth.
[0107] It can be understood that after obtaining the target image region, the present disclosure tracks the target image region through a target tracking algorithm. To ensure that the target tracking algorithm can identify the target object corresponding to the target image region, label information for identifying the target object needs to be set for the target object. The following embodiments of the present disclosure further illustrate the focusing method in the present disclosure.
[0108] Figure 19 is a flowchart of a focusing method shown according to an exemplary embodiment. As Figure 19 shown, the method includes steps S701 to step S702.
[0109] In step S701, the current ratio between the area of the intersection region and the area of the union region is determined.
[0110] In step S702, in response to the current ratio being greater than or equal to a preset ratio, label information is set for the target object image surrounded by the target image region, and the label information is used to identify the corresponding relationship between the target object image and the target object.
[0111] In the embodiments of the present disclosure, during the entire process of displaying the target object, label information (target ID) is set for the target object, and the label information is used to identify the corresponding relationship between the target object image and the target object. The target tracking algorithm is enabled to identify the corresponding relationship between the target object and the target image region, thereby tracking the target image region, so that the camera continuously focuses on the target object during the entire target object display process, and a clear target object image is continuously displayed in the preview image.
[0112] In an exemplary embodiment of the present disclosure, as Figure 20Flowchart of the middle focusing method, which realizes the automatic focusing of the camera in the following way, so that the focusing mode of the camera automatically switches between the focusing mode for the product to be displayed and the default focusing mode: Obtain the preview image captured by the camera in real time. Based on the running pixel detection technology and human body segmentation technology, obtain the pixels corresponding to the product, that is, use a motion detection algorithm (such as the background modeling method) to obtain all the pixels of the moving objects in the scene (including human body pixels and product pixels, and the product is carried and displayed by a person), and use algorithms such as human body semantic segmentation to separate the human body pixels and non-human body pixels (that is, the pixels corresponding to the product) from the above-obtained motion pixels. Then delete the human body pixels separated from the above motion pixels to obtain the pixels corresponding to the product. Through morphological processing methods (such as image morphological processing algorithms such as erosion and dilation), remove the noise of the pixels corresponding to the product to obtain the product pixels for further processing. Obtain the tracking frame surrounding the product pixels, use the target tracking algorithm to continuously track the tracking frame surrounding the product pixels (the detection frame of the displayed item), and assign a target ID to the product pixels. Then, based on the size relationship between the intersection over union (the ratio of the area of the intersection region to the area of the union region) of the tracking frame and the display area in the preview image and a preset ratio, determine whether the tracking frame enters the display area. When the intersection over union is less than the preset ratio, it is considered that the tracking frame has not entered the display area, and the default focusing mode (i.e., face focusing or center focusing) is adopted and the process ends. When the intersection over union is greater than or equal to the preset ratio, it is considered that the tracking frame has entered the display area, input the tracking frame coordinates into the automatic focusing system, control the camera lens to continuously focus on the product, and end the process.
[0113] In the embodiments of the present disclosure, a moving object detection scheme is applied to extract the motion pixels of the entire scene, and a human body segmentation technology is used to extract the human body pixels of the scene. The human body pixels are excluded from the motion pixels, and only the product pixels are retained. Obtain the tracking frame surrounding the product pixels, continuously track the tracking frame surrounding the product pixels through the target tracking algorithm, and determine whether the object enters the display area. When the tracking frame enters the display area, focus on the object, and restore the default focusing distance when leaving the display area. Through the present disclosure, smooth switching between focusing on the product and the default focusing method can be achieved in the item display scene. The video recording experience of specific users can be improved.
[0114] Based on the same concept, the embodiments of the present disclosure also provide a focusing device 100.
[0115] It can be understood that in order to implement the above functions, the focusing device 100 provided in the embodiments of the present disclosure includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of the various examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described function, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.
[0116] Figure 21 is a block diagram of a focusing device 100 shown according to an exemplary embodiment. Referring to Figure 21 , the device includes a determination unit 101 and a processing unit 102.
[0117] The determination unit 101 is configured to obtain in real time a preview image currently captured by the camera and determine a target image area in the preview image.
[0118] Wherein, the target image area is an image area corresponding to the target object in the preview image.
[0119] The processing unit 102 is configured to set the current focusing mode according to the positional relationship between the target image area and a preset image area.
[0120] Wherein, the preset image area is located in the central area of the preview image.
[0121] In one implementation, the preview image includes a human image, a target object image, and a background image. The determination unit 101 determines the target image area in the preview image in the following manner: obtaining a set of moving pixels in the target image area according to a motion detection algorithm, and the set of moving pixels corresponds to the human image and the target object image. Obtaining a set of target object pixels in the set of moving pixels according to a human semantic segmentation algorithm, and the set of target object pixels corresponds to the target object image. Determining the target image area according to the pixel coordinates corresponding to the set of target object pixels in the preview image.
[0122] In one implementation, the determining unit 101 determines the target image area according to the pixel coordinates corresponding to the target object pixel set in the preview image in the following manner: Determine the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate in the target object pixel set. The first coordinate is the pixel coordinate closest to the upper boundary of the preview image corresponding to the target object pixel set. The second coordinate is the pixel coordinate closest to the lower boundary of the preview image corresponding to the target object pixel set. The third coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the target object pixel set. The fourth coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the target object pixel set. Determine a first straight line parallel to the upper and lower boundaries of the preview image and passing through the first coordinate in the preview image, determine a second straight line parallel to the upper and lower boundaries of the preview image and passing through the second coordinate, determine a third straight line parallel to the left and right boundaries of the preview image and passing through the third coordinate, and determine a fourth straight line parallel to the upper and lower boundaries of the preview image and passing through the fourth coordinate. Determine the rectangular area corresponding to the first straight line, the second straight line, the third straight line, and the fourth straight line in the preview image as the target image area.
[0123] In one implementation, the processing unit 102 sets the current focusing mode according to the positional relationship between the target image area and the preset image area in the following manner, including: determining the area of the intersection area between the target image area and the preset image area, and determining the area of the union area between the target image area and the preset image area.
[0124] Determine the current ratio between the area of the intersection area and the area of the union area, and set the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio. The preset ratio is a critical ratio for focusing mode adjustment.
[0125] In one implementation, the processing unit 102 sets the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio in the following manner: In response to the current ratio being greater than or equal to the preset ratio, set the current focusing mode to the first mode. In the first mode, the focusing area of the camera corresponds to the target image area. In response to the current ratio being less than the preset ratio, set the current focusing mode to the second mode. In the second mode, the focusing area of the camera corresponds to the face image area in the preview image or the central area of the preview image.
[0126] In one implementation, in the first mode, the processing unit 102 makes the focusing area of the camera correspond to the target image area in the following manner: Obtain the coordinate set corresponding to the boundary of the target image area, and control the camera lens module according to the coordinate set so that the camera focus falls on the target object corresponding to the target image area.
[0127] In one implementation, the processing unit 102 is further configured to: in response to the current ratio being greater than or equal to a preset ratio, set label information for the target object image enclosed by the target image region, where the label information is used to identify the correspondence between the target object image and the target object.
[0128] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0129] Figure 22 FIG. 200 is a block diagram of a focusing device 200 according to an exemplary embodiment. The device 200 may be provided as a terminal. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0130] Referring to Figure 22 , the device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0131] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0132] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0133] The power component 206 provides power for various components of the device 200. The power component 206 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 200.
[0134] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0135] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0136] The I / O interface 212 provides an interface between the processing component 202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0137] The sensor assembly 214 includes one or more sensors for providing an assessment of the status of the device 200 in various aspects. For example, the sensor assembly 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0138] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0139] In an exemplary embodiment, the device 200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0140] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, and the above instructions can be executed by the processor 220 of the device 200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0141] It can be understood that in this disclosure, "a plurality of" means two or more, and other quantifiers are similar thereto. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The singular forms of "a", "the", and "said" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0142] It can be further understood that the terms "first", "second", etc. are used to describe various information, but this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other and do not represent a specific order or degree of importance. In fact, the expressions such as "first" and "second" can be used interchangeably. For example, without departing from the scope of this disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0143] It can be further understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this embodiment and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.
[0144] It can be further understood that unless otherwise specified, "connection" includes direct connection without other components between the two, and also includes indirect connection with other elements between the two.
[0145] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of this disclosure, it should not be understood as requiring these operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be beneficial.
[0146] Those skilled in the art will readily think of other embodiments of this disclosure after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this solution, which follow the general principles of this disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in this disclosure.
[0147] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A focusing method, characterized in that, Including: Obtain the preview image currently captured by the camera in real time, and determine the target image area in the preview image, where the target image area is the image area corresponding to the target object in the preview image; Set the current focusing mode according to the positional relationship between the target image area and the preset image area, and the preset image area is located in the central area of the preview image.
2. The method according to claim 1, wherein The preview image includes a person image, a target object image, and a background image; The determining the target image area in the preview image includes: Obtain the set of moving pixels in the target image area according to the motion detection algorithm, and the set of moving pixels corresponds to the person image and the target object image; Obtain the set of target object pixels in the set of moving pixels according to the human semantic segmentation algorithm, and the set of target object pixels corresponds to the target object image; Determine the target image area according to the pixel coordinates corresponding to the set of target object pixels in the preview image.
3. The method according to claim 2, wherein The determining the target image area according to the pixel coordinates corresponding to the set of target object pixels in the preview image includes: Determine the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate in the set of target object pixels. The first coordinate is the pixel coordinate closest to the upper boundary of the preview image corresponding to the set of target object pixels, the second coordinate is the pixel coordinate closest to the lower boundary of the preview image corresponding to the set of target object pixels, the third coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the set of target object pixels, and the fourth coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the set of target object pixels; Determine a first straight line parallel to the upper and lower boundaries of the preview image and passing through the first coordinate in the preview image, determine a second straight line parallel to the upper and lower boundaries of the preview image and passing through the second coordinate, determine a third straight line parallel to the left and right boundaries of the preview image and passing through the third coordinate, and determine a fourth straight line parallel to the upper and lower boundaries of the preview image and passing through the fourth coordinate; Determine the rectangular area corresponding to the first straight line, the second straight line, the third straight line, and the fourth straight line in the preview image as the target image area.
4. The method according to claim 1, characterized in that The setting the current focusing mode according to the positional relationship between the target image area and the preset image area includes: Determine the area of the intersection region between the target image area and the preset image area, and determine the area of the union region between the target image area and the preset image area; Determine the current ratio between the area of the intersection region and the area of the union region, and set the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio, where the preset ratio is the critical ratio for focus mode adjustment.
5. The method according to claim 4, characterized in that, The setting the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio includes: In response to the current ratio being greater than or equal to the preset ratio, set the current focusing mode to the first mode. In the first mode, the focusing area of the camera corresponds to the target image area; In response to the current ratio being less than the preset ratio, set the current focusing mode to the second mode. In the second mode, the focusing area of the camera corresponds to the face image area in the preview image or the central area of the preview image.
6. The method according to claim 5, characterized in that, In the first mode, the following method is adopted to make the focusing area of the camera correspond to the target image area: Obtain the coordinate set corresponding to the boundary of the target image area, and control the camera lens module according to the coordinate set so that the camera focus falls on the target corresponding to the target image area.
7. The method according to claim 5, characterized in that The method further includes: In response to the current ratio being greater than or equal to the preset ratio, set label information for the target object image enclosed by the target image area, and the label information is used to identify the corresponding relationship between the target object image and the target object.
8. A focusing device, characterized in that, It includes: A determination unit, configured to obtain the preview image currently captured by the camera in real time and determine the target image area in the preview image, where the target image area is the image area corresponding to the target object in the preview image; A processing unit, configured to set the current focusing mode according to the positional relationship between the target image area and the preset image area, and the preset image area is located in the central area of the preview image.
9. The device according to claim 8, characterized in that, The preview image includes a person image, a target object image, and a background image; The determination unit determines the target image area in the preview image in the following manner: Obtain the set of moving pixels in the target image area according to the motion detection algorithm, and the set of moving pixels corresponds to the person image and the target object image; Obtain the set of target object pixels in the set of moving pixels according to the human body semantic segmentation algorithm, and the set of target object pixels corresponds to the target object image; Determine the target image area according to the pixel coordinates corresponding to the set of target object pixels in the preview image.
10. The device according to claim 9, characterized in that, The determination unit determines the target image area according to the pixel coordinates corresponding to the set of target object pixels in the preview image in the following manner: Determine the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate in the set of target object pixels. The first coordinate is the pixel coordinate closest to the upper boundary of the preview image corresponding to the set of target object pixels, the second coordinate is the pixel coordinate closest to the lower boundary of the preview image corresponding to the set of target object pixels, the third coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the set of target object pixels, and the fourth coordinate is the pixel coordinate closest to the left boundary of the preview image corresponding to the set of target object pixels; Determine a first straight line parallel to the upper and lower boundaries of the preview image and passing through the first coordinate in the preview image, determine a second straight line parallel to the upper and lower boundaries of the preview image and passing through the second coordinate, determine a third straight line parallel to the left and right boundaries of the preview image and passing through the third coordinate, and determine a fourth straight line parallel to the upper and lower boundaries of the preview image and passing through the fourth coordinate; Determine the rectangular area corresponding to the first straight line, the second straight line, the third straight line, and the fourth straight line in the preview image as the target image area.
11. The device according to claim 8, characterized in that, The processing unit sets the current focusing mode according to the positional relationship between the target image area and the preset image area in the following manner, including: Determine the area of the intersection area between the target image area and the preset image area, and determine the area of the union area between the target image area and the preset image area; Determine the current ratio between the area of the intersection area and the area of the union area, and set the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio, where the preset ratio is a critical ratio for focusing mode adjustment.
12. The device according to claim 11, wherein, The processing unit sets the current focusing mode according to the magnitude relationship between the current ratio and the preset ratio in the following manner: In response to the current ratio being greater than or equal to the preset ratio, set the current focusing mode to the first mode. In the first mode, the focusing area of the camera corresponds to the target image area; In response to the current ratio being less than the preset ratio, set the current focusing mode to the second mode. In the second mode, the focusing area of the camera corresponds to the face image area in the preview image or the central area of the preview image.
13. The device according to claim 12, characterized in that, In the first mode, the processing unit makes the focusing area of the camera correspond to the target image area in the following manner: Obtain the coordinate set corresponding to the boundary of the target image area, and control the camera lens module according to the coordinate set so that the camera focus falls on the target corresponding to the target image area.
14. The device according to claim 12, characterized in that, The processing unit is further configured to: In response to the current ratio being greater than or equal to the preset ratio, set label information for the target object image surrounded by the target image area, where the label information is used to identify the corresponding relationship between the target object image and the target object.
15. A focusing device, characterized in that, Including: A processor: A memory for storing instructions executable by the processor; Wherein, the processor is configured to: execute the focusing method according to any one of claims 1 to 7.
16. A storage medium, characterized in that, Instructions are stored in the storage medium, and when the instructions in the storage medium are executed by the processor, the processor is enabled to execute the focusing method according to any one of claims 1 to 7.