Improved image projection for augmented reality

The described method and device improve AR systems by accurately projecting light guidance and adapting to product changes using U-Net and MLP networks, addressing cumbersome operations and limited guidance in existing AR systems.

WO2025147881A1PCT designated stage expired Publication Date: 2025-07-17TELEFONAKTIEBOLAGET LM ERICSSON (PUBL) +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/071515
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing augmented reality (AR) systems for smart manufacturing require cumbersome operations to reposition light guidance for new product models, and quality inspection systems provide limited guidance, necessitating frequent operator gaze shifts between the work area and a separate PC screen.

Method used

An electronic device and method for improved image projection that locates key points on a target image, applies perspective transforms, and projects guidance images accurately, adapting to product translations and rotations using U-Net and MLP networks for precise alignment.

Benefits of technology

Enables accurate automatic light guidance for new product series without expertise, supports different product heights, and facilitates quality inspection, enhancing production efficiency and convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024071515_17072025_PF_FP_ABST
    Figure CN2024071515_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure is related to an electronic device and a method for improved image projection for augmented reality. A method for image projection comprises: locating one or more first points on a first image of a target; locating one or more second points on the first image based on at least the one or more first points; and projecting a second image onto the target based on at least the one or more second points.
Need to check novelty before this filing date? Find Prior Art

Description

IMPROVED IMAGE PROJECTION FOR AUGMENTED REALITYTechnical Field

[0001] The present disclosure is related to the field of Augmented Reality (AR) , and in particular, to an electronic device and a method for improved image projection for AR.Background

[0002] AR is the integration of digital information with the user′s environment in Real Time. Unlike virtual reality (VR) , which creates a totally artificial environment, AR users experience a real-world environment with generated perceptual information overlaid on top of it.

[0003] Nowadays, AR is widely used in smart manufacture. It can help operators to execute working process correctly and provide guidance for each step. Therefore, it can improve production efficiency and product quality dramatically. For example, in some modern factories, this technology is applied to guide the operators to locate the correct screw hole in a correct order when using a screw driver. Further, a quality inspection can be implemented for each step by using AR.Summary

[0004] Some companies have provided their own smart manufacture systems based on the AR technology. For example, one of the smart manufacture systems may be a projector-camera system which can help operators by providing guidance and quality assurance. For another example, another smart manufacture system may provide animated guidance on a display for operations. When each step on a product is finished, a camera may take a picture of the product and detect its assembly quality.

[0005] However, with these systems, a lot of cumbersome operations needs to be done when a new type / model / series of products is to be manufactured, for example, positioning screw holes of the products such that light guidance from a projector can be projected to the screw holes to guide the correct operations.

[0006] With some smart manufacture systems, an operator needs to move a circular light spot 320 to a correct position where a corresponding screw hole (e.g., a screw hole 310) is located by using a mouse for example, as shown in Fig. 3. To be specific, as shown in (a) of Fig. 3, the light spot 320 is initially projected onto a product 300, and  it is probably not projected at its correct position. As shown in (b) of Fig. 3, the operator has to move the light spot 320 to the location of the screw hole 310, for example, by using a mouse to manually drag and drop the light spot 320 onto the screw hole 310. Theoretically, once this screw hole locating is finished, the light spot 320 can be automatically projected at the correct position (that is, the screw hole 310) for other products to be manufactured, which have a same product specification and is located at a same location as the product 300, as shown in (c) of Fig. 3.

[0007] Although only one screw hole 310 is shown in Fig. 3, there is typically a lot more number of screw holes in a real-world product than that shown in Fig. 3, resulting in a lot of tedious work. Further, a product to be manufactured is not always located at the expected location (e.g., as shown in (c) of Fig. 3) . For example, in a real-world production line, a translation and rotation of a product to be manufactured can be expected and probably inevitable, leading to for example incorrect positioning of the screw holes (e.g., as shown in (a) of Fig. 3) .

[0008] For another example, some other smart manufacture systems mainly focus on the quality inspection functionality, and it provides even less help in guidance than the systems mentioned above. For example, in order to obtain guidance or instructions from such systems, an operator needs to look up to a Personal Computer (PC) screen installed separately from a working area where a product is being manufactured, resulting in a frequent switching of line of sight between the working area and the PC screen, which is not convenient and efficient during the manufacturing of the product.

[0009] In order to address or at least partially alleviate at least one of the above problems, some embodiments of the present disclosure provide an electronic device and a method for improved image projection for AR.

[0010] According to a first aspect of the present disclosure, a method for image projection is provided. The method comprises: locating one or more first points on a first image of a target; locating one or more second points on the first image based on at least the one or more first points; and projecting a second image onto the target based on at least the one or more second points. Based on the first points for target locating, the system can determine the translation and / or rotation of the target and therefore accurately locate the second points where one or more operations need to be performed, and then light guidance can be projected accurately onto the second points.

[0011] In some embodiments, the one or more first points correspond to one or more features of the target, respectively. In some embodiments, the one or more second points correspond to one or more locations on the target, respectively, at which one or more operations are to be performed by an operator. In some embodiments, at least one of the one or more first points is located by: searching for a first feature in a first Region of Interest (ROI) in the first image; and determining a point in the first feature as the first point. In some embodiments, the first ROI is centered at a location with coordinates same as those of a reference point in a first reference image associated with the target. In some embodiments, the reference point corresponds to the first point.

[0012] In some embodiments, the step of searching for a first feature comprises: searching for the first feature by using a first U-Net convolutional network. In some embodiments, the step of determining a point in the first feature as the first point comprises: determining the center of the first feature as the first point by using the first U-Net convolutional network. In some embodiments, the step of locating one or more second points comprises: determining a first perspective transform from a first reference image associated with the target to the first image based on one or more first reference points in the first reference image and the one or more first points at least. In some embodiments, the first reference points correspond to the one or more first points, respectively. In some embodiments, the first perspective transform is defined by a first homography matrix that is determined based on the one or more first reference points and the one or more first points at least. In some embodiments, the one or more first points comprise four or more first points that are not collinear.

[0013] In some embodiments, at least one of the one or more second points is located by: determining an expected location in the first image corresponding to a second point based on at least the first perspective transform and a second reference point in the first reference image, the second reference point corresponding to the second point; searching for a second feature corresponding to the second point in a second ROI in the first image, the second ROI being centered at the expected location; and determining a point in the second feature as the second point.

[0014] In some embodiments, the first ROI has a larger area than that of the second ROI. In some embodiments, the step of searching for a second feature comprises: searching for the second feature by using a second U-Net convolution network. In some  embodiments, the step of determining a point in the second feature as the second point comprises: determining the center of the second feature as the second point by using the second U-Net convolutional network.

[0015] In some embodiments, the second image comprises one or more patterns which, when projected on the target, highlight one or more locations on the target, respectively. In some embodiments, the one or more highlighted locations on the target are locations where one or more operations are to be performed by an operator. In some embodiments, at least one of the one or more patterns is generated by: determining a location in a second reference image, the location corresponding to one of the one or more second points; determining a location in the second image, the determined location in the second image corresponding to the determined location in the second reference image; and generating the pattern in the second image based on at least the determined location in the second image. In some embodiments, the first reference image associated with the target is captured for a reference target that is placed on a supporting surface. In some embodiments, the second reference image is captured for the supporting surface without the reference target placed thereon. In some embodiments, the reference target has a same product specification as the target.

[0016] In some embodiments, at least two of the first image, the first reference image, and the second reference image are captured from a same viewpoint. In some embodiments, the step of determining a location in a second reference image comprises: calculating the location in the second reference image based on at least the corresponding second point and a second perspective transform from the first reference image to the second reference image. In some embodiments, the second perspective transform is defined by a second homography matrix that is calculated based on at least the first reference image and the second reference image. In some embodiments, the second perspective transform is determined by: capturing the first reference image when one or more first markers are projected onto a reference target placed on a supporting surface; determining, in the first reference image, the locations of the one or more first markers projected on the reference target; capturing the second reference image when one or more first markers are projected onto the supporting surface without the reference target placed thereon; determining, in the second reference image, the locations of the one or more first markers projected on the supporting surface; calculating the second perspective transform based on at least the locations of  the one or more first markers projected on the reference target in the first reference image and the locations of the one or more first markers projected on the supporting surface in the second reference image.

[0017] In some embodiments, the one or more first markers are one or more cross patterns, and the locations of the one or more first markers are the centers of the one or more cross patterns. In some embodiments, the step of determining a location in the second image comprises: determining the location in the second image based on at least the determined location in the second reference image and a mapping from locations in the second reference image to locations in the second image. In some embodiments, the mapping from locations in the second reference image to locations in the second image is determined by: capturing a third reference image when a third image is projected on a supporting surface without any target placed thereon, the third image comprising one or more second markers; determining one or more locations in the third reference image corresponding to the one or more second markers; and determining a mapping based on at least the one or more determined locations in the third reference image and the locations of the one or more second markers in the third image, as the mapping from locations in the second reference image to locations in the second image. In some embodiments, the one or more second markers comprise an array of circles.

[0018] In some embodiments, the one or more second markers projected on the supporting surface cover an area where the target is to be placed. In some embodiments, at least one of the one or more locations in the third reference image is determined by: labeling a point in a projected second marker in the third reference image; determining a third ROI comprising the projected second marker based on at least the labeled point; and searching for the projected second marker in the third ROI and determining the center of the projected second marker by using a third U-net convolutional network.

[0019] In some embodiments, the step of determining a mapping comprises: training a Multi-Layer Perceptron (MLP) network by using the determined one or more locations in the third reference image as inputs and using the locations of the one or more second markers in the third image as expected outputs, such that the trained MLP network is able to be used as the mapping from locations in the second reference image to locations in the second image. In some embodiments, the step of generating the  pattern in the second image comprises at least one of: generating a circle pattern in the second image that is centered at the determined location in the second image; and generating an arrow pattern in the second image that points to the determined location in the second image.

[0020] In some embodiments, the method further comprises: calibrating a camera that is used for capturing at least one of the first image, the first reference image, the second reference image, and the third reference image. In some embodiments, the step of calibrating the camera comprises: capturing one or more images of a calibration target by using the camera when the calibration target is placed at one or more places on the supporting surface; calculating a distortion of the camera based on at least the one or more images of the calibration target and one or more known characteristics of the calibration target; and compensating one or more images captured by the camera for the calculated distortion.

[0021] In some embodiments, the second image, when projected onto the target, further indicates at least one of: information related to the target; information related to the operations to be performed on the target; information related to status of a camera used for capturing the first image and / or a projector used for projecting the second image; and one or more icons for interaction with an operator.

[0022] According to a second aspect of the present disclosure, an electronic device is provided. The electronic device comprises: a processor; a memory storing instructions which, when executed by the processor, cause the processor to: locate one or more first points on a first image of a target; locate one or more second points on the first image based on at least the one or more first points; and project a second image onto the target based on at least the one or more second points. In some embodiments, the instructions, when executed by the processor, cause the processor further to perform any of the methods of the first aspect.

[0023] According to a third aspect of the present disclosure, an electronic device is provided. The electronic device comprises: a first locating module configured to locate one or more first points on a first image of a target; a second locating module configured to locate one or more second points on the first image based on at least the one or more first points; and a projecting module configured to project a second image onto the target based on at least the one or more second points. In some embodiments,  the electronic device comprises one or more further modules, each of which may perform any of the steps of any of the methods of the first aspect.

[0024] According to a fourth aspect of the present disclosure, a computer program comprising instructions is provided. The instructions, when executed by at least one processor, cause the at least one processor to carry out any of the methods of the first aspect.

[0025] According to a fifth aspect of the present disclosure, a carrier containing the computer program of the fourth aspect is provided. In some embodiments, the carrier may be one of an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0026] According to a sixth aspect of the present disclosure, a system for image projection is provided. The system comprises: a supporting surface on which a target to be placed; a projector configured to project an image onto the supporting surface and / or the target; a camera configured to capture an image of at least one of the supporting surface, the target, and the projected image; a processor; a memory storing instructions which, when executed by the processor, cause the processor to: locate one or more first points on a first image of the target captured by the camera; locate one or more second points on the first image based on at least the one or more first points; and project, by the projector, a second image onto the target based on at least the one or more second points. In some embodiments, the instructions, when executed by the processor, cause the processor further to perform any of the methods of the first aspect.

[0027] With some embodiments of the present disclosure, automatic light guidance may be provided accurately even when translations and / or rotations of products to be manufactured occur. Further, it is easier and more convenient for an operator to introduce a new series of product, and no specific expertise knowledge is required. Further, the system is applicable to different series of products with different heights. Further, it is easier to determine a mapping relationship from a canvas image to a projected image, such that a more accurate projection of the image onto a target may be achieved. Furthermore, downstream tasks, such as quality inspection, can be implemented based on this solution.Brief Description of the Drawings

[0028] The foregoing and other features of the present disclosure will become more fully apparent from the following description and appended claims, taken in conjunction with the accompanying drawings. Understanding that these drawings depict only several embodiments in accordance with the disclosure and therefore are not to be considered limiting of its scope, the disclosure will be described with additional specificity and detail through use of the accompanying drawings.

[0029] Fig. 1A and Fig. 1B are diagrams illustrating different perspective views of an exemplary system for image projection according to an embodiment of the present disclosure.

[0030] Fig. 2 is a diagram illustrating an exemplary top view of a working area in a system for image projection according to an embodiment of the present disclosure.

[0031] Fig. 3 is a diagram illustrating how a light spot is moved to a screw hole in the related art.

[0032] Fig. 4A and Fig. 4B are diagrams illustrating an exemplary calibration target for camera calibration at different locations in a working area according to an embodiment of the present disclosure.

[0033] Fig. 5 is a diagram illustrating exemplary images of a calibration target before and after calibration according to an embodiment of the present disclosure.

[0034] Fig. 6 is a diagram illustrating an exemplary scenario where a canvas image for calibration is projected onto a supporting surface according to an embodiment of the present disclosure.

[0035] Fig. 7A through Fig. 8B are diagrams illustrating an exemplary procedure for determining centers of circles in an array according to an embodiment of the present disclosure.

[0036] Fig. 9 is a diagram illustrating an exemplary MLP network trained for mapping coordinates in a projected image to coordinates in a canvas image according to an embodiment of the present disclosure.

[0037] Fig. 10A through Fig. 10C are diagrams illustrating an exemplary procedure for determining a perspective transform from coordinates in a scenario where a product is placed in a working area to coordinates in a scenario where the product is not placed in the working area according to an embodiment of the present disclosure.

[0038] Fig. 11A and Fig. 11B are diagrams illustrating how the centers of key points and screw holes of a reference product are determined according to an embodiment of the present disclosure.

[0039] Fig. 12A through Fig. 12F are diagrams illustrating how the centers of screw holes of a product under manufacture are determined and how an image for AR is projected based thereon according to an embodiment of the present disclosure.

[0040] Fig. 13 is a diagram illustrating an exemplary working area with additional operations performed and additional information provided according to an embodiment of the present disclosure.

[0041] Fig. 14 is a diagram illustrating an exemplary overall procedure for image projection for AR according to an embodiment of the present disclosure.

[0042] Fig. 15 is a flow chart illustrating an exemplary method for image projection according to an embodiment of the present disclosure.

[0043] Fig. 16 schematically shows an embodiment of an arrangement which may be used in an electronic device according to an embodiment of the present disclosure.

[0044] Fig. 17 is a block diagram of an exemplary electronic device according to an embodiment of the present disclosure.Detailed Description

[0045] Hereinafter, the present disclosure is described with reference to embodiments shown in the attached drawings. However, it is to be understood that those descriptions are just provided for illustrative purpose, rather than limiting the present disclosure. Further, in the following, descriptions of known structures and techniques are omitted so as not to unnecessarily obscure the concept of the present disclosure.

[0046] Those skilled in the art will appreciate that the term "exemplary" is used herein to mean "illustrative, " or "serving as an example, " and is not intended to imply that a particular embodiment is preferred over another or that a particular feature is essential. Likewise, the terms "first" , "second" , "third" , "fourth, " and similar terms, are used simply to distinguish one particular instance of an item or feature from another, and do not indicate a particular order or arrangement, unless the context clearly indicates otherwise. Further, the term "step, " as used herein, is meant to be synonymous with "operation" or "action. " Any description herein of a sequence of steps does not imply that these operations must be carried out in a particular order, or even that these  operations are carried out in any order at all, unless the context or the details of the described operation clearly indicates otherwise.

[0047] Conditional language used herein, such as "can, " "might, " "may, " "e.g., " and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. Also, the term "or" is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Further, the term "each, " as used herein, in addition to having its ordinary meaning, can mean any subset of a set of elements to which the term "each" is applied.

[0048] The term "based on" is to be read as "based at least in part on. " The term "one embodiment" and "an embodiment" are to be read as "at least one embodiment. " The term "another embodiment" is to be read as "at least one other embodiment. " Other definitions, explicit and implicit, may be included below. In addition, language such as the phrase "at least one of X, Y and Z, " unless specifically stated otherwise, is to be understood with the context as used in general to convey that an item, term, etc. may be either X, Y, or Z, or a combination thereof.

[0049] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limitation of example embodiments. As used herein, the singular forms "a" , "an" , and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" , "comprising" , "has" , "having" , "includes" and / or "including" , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. It will be also understood that the terms "connect (s) , " "connecting" , "connected" , etc. when used herein, just mean that there is an electrical or communicative connection between two  elements and they can be connected either directly or indirectly, unless explicitly stated to the contrary.

[0050] Of course, the present disclosure may be carried out in other specific ways than those set forth herein without departing from the scope and essential characteristics of the disclosure. One or more of the specific processes discussed below may be carried out in any electronic device comprising one or more appropriately configured processing circuits, which may in some embodiments be embodied in one or more application-specific integrated circuits (ASICs) . In some embodiments, these processing circuits may comprise one or more microprocessors, microcontrollers, and / or digital signal processors programmed with appropriate software and / or firmware to carry out one or more of the operations described above, or variants thereof. In some embodiments, these processing circuits may comprise customized hardware to carry out one or more of the functions described above. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.

[0051] Although multiple embodiments of the present disclosure will be illustrated in the accompanying Drawings and described in the following Detailed Description, it should be understood that the disclosure is not limited to the disclosed embodiments, but instead is also capable of numerous rearrangements, modifications, and substitutions without departing from the present disclosure that as will be set forth and defined within the claims.

[0052] The term "electronic device" used herein may refer to (but not limited to) a PC, a laptop, a workstation, a mainframe, a handheld device, a vehicle-mounted device, a drone, a User Equipment (UE) , a terminal device, a mobile device, a mobile terminal, a mobile station, a user device, a user terminal, a wireless device, a wireless terminal, a vehicle-mounted-device, a network node, a base station, a base transceiver station, an access point, a hot spot, a NodeB, an Evolved NodeB, a gNB, a network element, a network function, or any other equivalents.

[0053] Further, please note that although the following description of some embodiments of the present disclosure is given in the context of smart manufacture (e.g., projecting light spots onto screw holes for guidance) , the present disclosure is not limited thereto. Furthermore, the terms, "target" , "product" , "item" , and / or "object" can be used interchangeably herein. Furthermore, the terms "perspective transform" ,  "coordinate transform" , "coordinate system transform" , or the like may be used interchangeably herein.

[0054] Furthermore, relative terms, such as "lower" , "bottom" , "upper" , "top" , "left" , or "right, " may be used herein to describe one element′s relationship to another element as illustrated in the Figures. It will be understood that relative terms are intended to encompass different orientations of the object in addition to the orientation depicted in the Figures. For example, if the object in one of the figures is turned over, elements described as being on the "lower" side of other elements would then be oriented on "upper" sides of the other elements. The exemplary term "lower" , can therefore, encompasses both an orientation of "lower" and "upper, " depending on the particular orientation of the figure. Similarly, if the object in one of the figures is turned over, elements described as "below" or "beneath" other elements would then be oriented "above" the other elements. The exemplary terms "below" or "beneath" can, therefore, encompass both an orientation of above and below.

[0055] Exemplary embodiments of the present disclosure are described herein with reference to illustrations that are schematic illustrations of idealized embodiments (and intermediate structures) of the present disclosure. As such, variations from the shapes of the illustrations as a result, for example, of manufacturing techniques and / or tolerances, may be expected. Thus, the disclosed example embodiments of the present disclosure should not be construed as limited to the particular shapes of regions illustrated herein unless expressly so defined herein, but are to include deviations in shapes that result, for example, from manufacturing. Thus, the regions illustrated in the figures are schematic in nature and their shapes are not intended to illustrate the actual shape of a region of a device and are not intended to limit the scope of the invention, unless expressly so defined herein.

[0056] Further, the following paper is incorporated herein by reference in its entirety:

[0057] - [1] Olaf et al., "U-Net: Convolutional Networks for Biomedical Image Segmentation" , May 18, 2015.

[0058] As mentioned above, when a new product (e.g., a new model, a new series, or a new variant) is introduced and to be manufactured, positioning (e.g., positioning of screw holes) needs to be done so that lights from a projector can be projected to screw holes. Further, in a real-world production line, a translation and rotation of a product to  be manufactured can be expected and probably inevitable, leading to for example incorrect positioning of the screw holes.

[0059] To address or at least partially alleviate at least one of problems mentioned above, a solution is provided in some embodiments of the present disclosure.

[0060] In some embodiments, the solution may be divided into 3 parts as follows.

[0061] Part 1: Calibration

[0062] Part 1 aims at obtaining a corresponding mapping relationship between a projected image and its canvas where the image is generated and buffered. By the term "canvas" , it actually refers to a virtual canvas or a buffer in a computing system, such that an image generated and buffered in the canvas can be projected by a projector. In some embodiments, the term "canvas image" or "image in the canvas" may refer to an image buffered and to be projected.

[0063] Step 1: Camera calibration

[0064] a) Capture chessboard images (e.g., as shown in Fig. 4A and Fig. 4B) ;

[0065] b) Use OpenCV Application Programming Interface (API) to calculate distortion related matrix.

[0066] Step 2: Find the mapping relationship between the coordinates in a projected image and coordinates in a corresponding canvas image.

[0067] a) Project a canvas image including a 15-by-15 array of circles (or another image, such as a structured pattern) on the desk (e.g., as shown in 7B) ;

[0068] b) Use UNet [1] to segment each circle and obtain the coordinates of centers (e.g., as shown in Fig. 8A and Fig. 8B) . In some embodiments, this uNet may be denoted as UNet_projection_circle;

[0069] c) Build an MLP network (e.g., as shown in Fig. 9) to map coordinates between the centers of the circles in the projected image and the centers of the circles in the corresponding canvas image. In some embodiments, this MLP network may be denoted as MLP_Projection_Canvas_Relationship.

[0070] Part 2: New product introduction

[0071] Part 2 aims at finding the relationship between the product (e.g., the target 120 shown in Fig. 1A) and the desk (e.g., the supporting surface 105 shown in Fig. 1A) for the projection light.

[0072] Step 1: Find the perspective transform between 2 scenarios with and without product for projection light.

[0073] a) Project 4 crosses on the desk, use a program called "labelme" (https:  / / github. com / tzutalin / labelImg) to record the coordinates of the intersections of the crosses. In some embodiments, the coordinates may be denoted as P_intersection_without_product.

[0074] b) Place the product on the desk and project the 4 crosses on the product and use labelme to record the coordinates of the intersections of the crosses. In some embodiments, the coordinates may be denoted as P_intersection_with_product.

[0075] c) Obtain the perspective transform from P_intersection_with_product to P_intersection_without_product. In some embodiments, this homography matrix may be denoted as PT_product.

[0076] Step 2: Use labelme to label key points of the product (e.g., K0 through K3 shown in Fig. 2) and the centers of the screw holes (e.g., S0 through S7 shown in Fig. 2) . In some embodiments, they may be denoted as P_key_points_reference_product and P_screw_hole_reference_product.

[0077] In some embodiments, the product images captured in this part may be denoted as "reference images" . For example, before the system goes online for real production, a product prototype, which has a same product specification as those to be manufactured, may be used for generating the reference images.

[0078] Part 3: On-line detection

[0079] Part 3 aims at implement light guidance on-line when the products (which have a same product specification as that of the product prototype) keep on going into the product line.

[0080] Step 1: Obtain the perspective transform between a real image and the reference image in part 2, for example, based on the 4 key points.

[0081] a) Use large ROIs to search for the key objects (e.g., the key points K0 through K3 shown in Fig. 2) . In some embodiments, a UNet may be used to detect the centers of the key objects. In some embodiments, the coordinates of the centers may be denoted as P_key_points_real_product. In some embodiments, this UNet may be denoted as UNet_key_point.

[0082] b) Obtain the perspective transform from P_key_points_reference_product to P_key_points_real_product. In some embodiments, this homography matrix may be denoted as PT_key_points.

[0083] Step 2: Determine the centers of the screw holes.

[0084] a) Position screw holes coarsely based on the matrix PT_key_points and P_screw_hole_reference_product,

[0085] b) Use small ROIs to search for the screw holes. In some embodiments, a UNet may be used to detect the centers of the screw holes. In some embodiments, the coordinates may be denoted as P_screw_hole_real_product and the UNet may be denoted as UNet_screw_hole.

[0086] Step 3: Implement light guidance for the screw holes.

[0087] a) Use the homography matrix PT_product and P_screw_hole_real_product to determine the coordinates P_screw_hole_without_product. It means that if the product is removed from the desk, the light will be projected on the desk at this coordinates.

[0088] b) Use the MLP named MLP_Projection_Canvas_Relationship in step 2 of part 1 to obtain the coordinates in the canvas. In some embodiments, the coordinates in the canvas may be denoted as P_screw_hole_canvas_real_product. Then color circles can be drawn on the canvas and projected onto the product to realize the light guidance.

[0089] With some embodiments of the present disclosure, calibration of camera and / or projector in the system can be done in an easy way, which leads to an easier deployment of the production line. Further, the system can be automatically adapted to the translations and / or rotations of the product by using a large ROI and perspective transformation for product location. Therefore, it can be used widely in smart manufacture. Furthermore, it enables a very fast introduction of a new product. Also, this system can be automatically adapted to variant kinds of products having different heights. The production efficiency can be improved dramatically.

[0090] With some embodiments of the present disclosure, automatic light guidance may be provided accurately even when translations and / or rotations of products to be manufactured occur. Further, it is easier and more convenient for an operator to introduce a new series of product, and no specific expertise knowledge is required. Further, the system is applicable to different series of products with different heights. Further, it is easier to determine a mapping relationship from a canvas, where an image to be projected is generated, to a projected image, such that a more accurate projection of the image onto a target may be achieved. Furthermore, downstream tasks, such as quality inspection, can be implemented based on this solution.

[0091] Fig. 1A and Fig. 1B are diagrams illustrating different perspective views of an exemplary system 10 for image projection according to an embodiment of the present  disclosure. As shown in Fig. 1A and Fig. 1B, the system 10 may comprise a production station (or a working table / desk or the like) 100 which has its top surface used as a supporting surface (or working surface) 105. The supporting surface 105 may be used to provide a working area 115 and support a target 120 (e.g., a product to be manufactured) . Further, a frame 110 may be provided on the supporting surface 105, to which or in which a camera 111 and a projector 113 may be attached and fixed, such that images can be projected by the projector 113 onto the working area 115 and an image of anything in the working area 115 can be captured by the camera 111. Furthermore, a computing system (e.g., a PC, a laptop, a workstation, or the like) may be provided in the production station 100 or remote from and connected to the production station 100 (though not shown in Fig. 1A or Fig. 1B) , such that the images captured by the camera 111 and the images to be projected by the projector 113 can be processed by the computing system. In some embodiments, the computing system may be accommodated in the base compartment of the production station 100. In some embodiments, the computing system may be remotely located from the production station 100 and communicatively connected to the camera 111 and / or the projector 113 via a cable or wirelessly.

[0092] However, the present disclosure is not limited thereto, and the production station 100 and / or the system 10 may have more, less, or different components than those shown in Fig. 1A and Fig. 2B, and / or a different configuration of the components. For example, the camera 111 and / or the projector 113 may be located at different places other than those shown in Fig. 1B. For another example, the supporting surface 105 and / or the working area 115 may have different sizes and / or geometries than those shown in Fig. 1A.

[0093] Referring back to Fig. 1A and Fig. 1B, when the target 120 is placed in the working area 115, the camera 111 may capture an image of the target 120. After that, the computing system may detect the screw holes of the target 120 in the captured image and instruct the projector 113 to project an image onto the target 120. In some embodiments, the image projected onto the target 120 may provide light guidance (e.g., the circle patterns on the target 120 shown in Fig. 1A) for the detected screw holes.

[0094] Fig. 2 is a diagram illustrating an exemplary top view of a working area (e.g., the working area 115 shown in Fig. 1A) of a system (e.g., the system 10 shown in Fig. 1A and Fig. 1B) for image projection according to an embodiment of the present disclosure.  As shown in Fig. 2, a target (e.g., the target 120 shown in Fig. 1A) is placed in the working area 115, and it may comprise one or more features 121 for target locating and one or more features 123 to which image projection is to be performed.

[0095] In some embodiments, each of the features 121 may have a specific size and / or geometry that are relatively easy for the computing system to identify. For example, the features 121 may be circle patterns as shown in Fig. 2, or a more complex structure or even a printed label (e.g., a bar code, a QR code, or the like) . In some embodiments, a location (or coordinates) of a feature 121 may be determined as (coordinates of) any point in the feature (e.g., the center of the feature 121) , and the point may be used as the location of the feature 121 throughout the whole procedure. In some embodiments, if the feature 121 is a centrally symmetrical feature (e.g., a circle pattern) , then its center may be used as the location of the feature 121.

[0096] In some embodiments, each of the features 123 may also have a specific size and / or geometry that can be identified by the computing system. In some embodiments, the features 123 may correspond to the locations where one or more operations are to be performed. For example, a feature 123 may correspond to a screw hole into which a screw is to be driven. In some embodiments, a location (or coordinates) of a feature 123 may be determined as (coordinates of) any point in the feature (e.g., the center of the feature 123) , and the point may be used as the location of the feature 123 throughout the whole procedure. In some embodiments, if the feature 123 is a centrally symmetrical feature (e.g., a circle pattern) , then its center may be used as the location of the feature 123.

[0097] As shown in Fig. 2, a total of four features 121 and a total of eight features 123 are provided on the target 120. However, the present disclosure is not limited thereto. Based on the actual needs, a different number of features 121 and / or a different number of features 123 with potentially different geometries and / or sizes may be provided at potentially different locations on the target 120.

[0098] As also shown in Fig. 2, a projector (e.g., the projector 113 shown in Fig. 1B) may project information for assisting an operator in manufacturing the target 120. For example, on the left side of the working area 115, progresses for various steps in the production procedure may be listed. For example, as shown in Fig. 2, when the target 120 is an electronic device, the information projected on the left side of the working area 115 may comprise:

[0099] -The manufacture of its main board is completed.

[0100] -The screws for the main board are already driven into the main board.

[0101] -An Electromagnetic Compatibility (EMC) cover is attached to the main board.

[0102] -Now, the current step is to drive screws into the target 120 such that the EMC cover can be fixed to the main board.

[0103] -Further, the operator is also prompted with which tool he / she should use (e.g., a Phillips screwdriver) when driving the screws.

[0104] As also shown in Fig. 2, on the right side of the working area 115, information related to the production station (e.g., the production station 100) and / or the target 120 is provided, for example, the production line to which the production station belongs, the product model to be manufactured, the status of the projector 113, the status of the camera 111, and the industrial inspection result. Further, an icon for user interaction (e.g., "the next step" icon shown in Fig. 2) may be provided for the operator.

[0105] Although some exemplary information are shown in Fig. 2 and described above, the present disclosure is not limited thereto. In some other embodiments, additional information and / or different information may be projected into the working area 115, such that the operator may be provided with useful information during his / her manufacture of the target 120.

[0106] As mentioned earlier, an exemplary overall procedure for projecting light guidance onto a product in an accurate and robust manner may be divided into 3 parts: Part 1 for Calibration, Part 2 for new product introduction, and Part 3 for real production. Next, a detailed description of the overall procedure will be provided below with reference to Fig. 4A through Fig. 12F.

[0107] Part 1: Calibration

[0108] Step 1: Camera calibration

[0109] In some embodiments, this step may be optional since the camera (e.g., the camera 111) may already be well calibrated before its use. In some embodiments, following operations may be performed to calibrate the camera 111.

[0110] First, a target for calibration or a "calibration target" 400 may be placed at different places in a working area (e.g., the working area 115) , and one or more images thereof may be captured by using the camera 111 to be calibrated. For example, as shown in Fig. 4A and Fig. 4B, the calibration target 400 may be a chessboard like target or another type of target for calibration. The calibration target 400 is placed in the  upper-left corner of the working area 115 as shown in Fig. 4A and in the lower-right corner of the working area 115 as shown in Fig. 4B. Additionally or alternatively, the calibration target 400 may be placed at other locations.

[0111] With the images and the known characteristics of the calibration target 400, the computing system may calculate distortions in the images, which are caused by the camera 111, and compensate the images for the calculated distortion. For example, when the calibration target 400 is placed as shown in Fig. 4B, a barrel distortion may be observed on the right side of the image captured by the camera 111 (for example, a part of the image captured by the camera 111 is shown in (a) of Fig. 5, in which a barrel distortion can be observed) , while it is known that the calibration target 400 has straight edges (for example, as shown in (b) of Fig. 5) . The computing system may detect such a distortion in the image and compensate the image for the detected distortion, for example, by using a tool called OpenCV API. In some embodiments, the OpenCV API may be used to calculate a distortion related matrix to compensate the camera distortion. Since the OpenCV API is well known in the art, a detailed description thereof is omitted for simplicity. After the calibration, the part of the image may look like the original calibration target 400 shown in (b) of Fig. 5.

[0112] Step 2: Find the mapping relationship between the coordinates in the projected image and the coordinates in the corresponding canvas image.

[0113] In some embodiments, in addition to the camera distortion, a distortion in an image projected by the projector 113 may also be observed, for example, as shown in Fig. 6. In Fig. 6, when a canvas image 600 that has multiple circles 610 provided along its edges is projected onto the supporting surface 105, a part of the projected image including some of the projected circles 610′ cannot be fully observed, for example, those in the lower-left corner of the working area 115. In other words, the edge of the projected image is not parallel to the edge of the desk and / or the edge of the image captured by the camera 111.

[0114] Please note that although an image projected on the supporting surface 105 may cover a same area (i.e. same size and same geometry) as that of the working area 115 as shown in Fig. 6, the present disclosure is not limited thereto. Likewise, please also note that although an image captured by the camera 111 may cover a same area (i.e., same size and geometry) as that of the working area 115, the present disclosure is not limited thereto. In some other embodiments, the projected image and / or the captured  image may cover an area (e.g., a different size and / or a different geometry) different from that of the working area 115 and / or different from each other. In some embodiments, the projected image and / or the captured image may have a larger area than that of the working area 115, such that the working area 115 can be fully covered by the captured image and / or the projected image.

[0115] Although it is difficult to calibrate the projector-camera system, a simple method is designed to solve the problem as follows.

[0116] First, a canvas image comprising one or more markers (e.g., an array of circles as shown in Fig. 7A) 700 may be generated in the canvas and projected onto the working area 115. In some embodiments, the one or more markers 700 may be a 15-by-15 array of circles as shown in Fig. 7A. In some embodiments, a number may be generated and projected together with each circle 700, and the numbers may be placed in or next to the circles, respectively. However, the present disclosure is not limited thereto. In some other embodiments, the one or more markers 700 may have a different geometry and / or a different size. In some other embodiments, the canvas image may have a different arrangement of the one or more markers 700 than the circle array shown in Fig. 7A.

[0117] Fig. 7B shows a projected image corresponding to the canvas image shown in Fig. 7A, and a distortion may be observed in the projected markers 700′ due to one or more factors, such as projector lens distortion, an unexpected angle between the projector 113, the camera 111, and the supporting surface 105, or the like.

[0118] As shown in Fig. 7C, the program "labelme" as mentioned above may be used to label the centers of the circles in a rough / coarse manner. After that, a UNet as mentioned above may be used to determine, in the projected image, multiple ROIs based on at least the roughly labelled centers. In some embodiments, each of the multiple ROIs may be a square region and include a circle, but the present disclosure is not limited thereto. In some other embodiments, the ROIs may have a different geometry than square. After the multiple ROIs are determined, the precise coordinates of the centers of the circles may be determined by the UNet in a highly accurate manner as will be described with reference to Fig. 8A and Fig. 8B below.

[0119] Referring back to Fig. 7C, in some embodiments, the operator may use "labelme" to label the centers of the circles, for example, by manually drawing lines and traversing all the circles in the array. However, the labelled centers are not accurate enough, and  therefore multiple ROIs may be determined in the projected image with each ROI centered at one of the labelled centers as shown in (a) of Fig. 8A. In some embodiments, an ROI may have 120 pixels in both of its width and height directions. However, the present disclosure is not limited thereto. In some other embodiments, an ROI may have a different size and / or geometry as also mentioned earlier.

[0120] As also mentioned earlier, a UNet may be used to segment the circles and determine its center. In some embodiments, a UNet named UNet_projection_circle may be trained to do so. In order to train this network, a number of pairs of original image and its mask may be collected, and some examples of the pairs are shown in (b) of Fig. 8A.

[0121] After finishing the training of UNet_projection_circle, the OpenCV API may be used to track the contours of the circles and determine the centers of the circles, as can be seen from (c) of Fig. 8A. Fig. 8B shows an overall result for an exemplary center detection, which shows that the centers 715 of all the circles are identified successfully.

[0122] In some embodiments, an MLP network may be built to map the coordinates of the centers of the circles in the projected image (e.g., that shown in Fig. 8B) to the coordinates of the centers of the circles in the corresponding canvas image (e.g., that shown in Fig. 7A) . In some embodiments, this MLP may be denoted as MLP_Projection_Canvas_Relationship. Fig. 9 shows an exemplary MLP network that can be used to do such a mapping.

[0123] To train this MLP network, the pairs of the coordinates of the centers of the projected circles and the coordinates of the centers of the circles in the canvas image can be used as inputs and expected outputs, respectively. With the trained MLP network, the inputs (or the coordinates in the projected image) can be mapped to the outputs (or the coordinates in the canvas image) as shown in Fig. 9. Through this neural network, the mapping relation can be determined.

[0124] Part 2: New product introduction

[0125] Step 1: Find the perspective transform between a scenario in which the product (e.g., the target 120 shown in Fig. 2) is placed in the working area (e.g., the working area 115) and a scenario in which the product is not placed in the working area, such that coordinates in these two scenarios can be converted to each other for later use for the projection light. This step will be described in detail with reference to Fig. 10A through Fig. 10C.

[0126] In some embodiments, a canvas image with one or more markers may be generated, for example, as shown in Fig. 10A. In some embodiments, the markers may be crosses. In some embodiments, the number of the markers may be greater than or equal to 4 to make sure the perspective transform determined based thereon is more accurate, robust, and reliable.

[0127] As shown in Fig. 10A, four crosses P0 through P3 are generated in the canvas image. Then the canvas image may be projected onto the working area 115 as shown in Fig. 10B. By comparing Fig. 10A and Fig. 10B, it is clear that the projected markers P0′ through P3′ may have a slightly different space relationship than that of the original markers P0 through P3 in the canvas image. This difference may be caused by the projector lens distortion, an unexpected angle between the projector 113, the camera 111, and the working area 105 as mentioned earlier, and it will be eliminated or compensated at a later stage, for example, by using the MLP network mentioned above.

[0128] In some embodiments, an image of projected markers P0′ through P3′ may be captured by the camera 111, and the centers of the projected markers P0′ through P3′ (e.g., the intersection points of the crosses) may be determined in the captured image, for example, by using labelme as indicated by the dashed line, and the coordinates of the intersection points may be recorded as, for example, P_intersection_without_product.

[0129] In some embodiments, the target 120 may be placed in the working area 115 and the canvas image including the 4 crosses may be projected onto the target 120 and also onto the working area 115. Similarly, an image of projected markers P0" through P3" may be captured by the camera 111, and the centers of the projected markers P0" through P3" (e.g., the intersection points of the crosses) may be determined in the captured image, for example, by using labelme as indicated by the dashed line, and the coordinates of the intersection points may be recorded as, for example, P_intersection_with_product.

[0130] As can be observed from Fig. 10B and Fig. 10C, there is a perspective transform between the different projections of the same markers P0 through P3 due to the existence / absence of the target 120. In some embodiments, the perspective transform may be calculated or determined based on at least P_intersection_with_product and P_intersection_without_product. In some embodiments, this holography matrix may be denoted as PT_product. Please note that how to calculate a perspective transform  based on coordinates before and after the perspective transform, and therefore a detailed description thereof is omitted for simplicity.

[0131] Step 2: Use labelme to label key points (e.g., K0 through K3) of the product and the centers (e.g., S0 through S7) of the screw holes. In some embodiments, they may be denoted as P_key_points_reference_product and P_screw_hole_reference_product, respectively.

[0132] In some embodiments, the key points cannot be located in a same line, and each key point is far away from other screw holes.

[0133] Note: the captured product image in this part may be denoted as the reference image.

[0134] As shown in Fig. 11A, the key points, P_key_points_reference_product, may consist of K0, K1, K2, and K3, for example, which may be identified by using labelme. As shown in Fig. 11B, the centers of the screw holes, P_screw_hole_reference_product, may consist of S0, S1, S2, S3, S4, S5, S6, and S7, for example, which may be identified by using labelme.

[0135] Part 3: On-line detection

[0136] Step 1: Find the perspective transform between a real image in the production line and the reference image in part 2 based on the (4) key points.

[0137] As mentioned earlier, a problem that the screw holes cannot be correctly highlighted by light spots will arise when the product is placed on the desk and a random translation and / or rotation may happen. If small ROIs are used to search for the centers of the screw holes, it is easy to miss the real screw holes since they may be located outside of the ROIs caused by the translation and / or rotation. On the other hand, if large ROIs are used, more than 1 screw hole may be located inside a same ROI, and therefore the program may fail to identify the correct center of the specific screw hole. This is a contradiction.

[0138] To solve or at least alleviate this problem, a solution with 4 (or more) key points selected is proposed. In some embodiments, these 4 points (e.g., centers of 4 holes and they need not to be screw holes) are not in the same line and they are located far away from other screw holes. With such a configuration, large ROIs may be used on them. From these 4 key points and algorithms, the screw holes can be positioned even when the translation and / or rotation happen.

[0139] In some embodiments, large ROIs may be used to search for the key objects (e.g., the key points K0 through K3) . In some embodiments, a UNet may be used to detect the centers of the key objects. In some embodiments, the coordinates of the centers may be denoted as P_key_points_real_product. In some embodiments, this UNet may be denoted as UNet_key_point. In some embodiments, the training method may be performed in a same or similar way as that described in step 2 in part 1 above.

[0140] Fig. 12A shows an exemplary result of the key point detection. As shown in Fig. 12A, dotted bounding boxes may be used as the large ROIs 1210 for the key point detection. In some embodiments, the centers of the ROIs 1210 may be determined as the coordinates named P_key_points_reference_product as that in Step 2 in part 2 above. In some embodiments, the size of the ROIs 1210 may be 320 *320 pixels.

[0141] In some embodiments, the perspective transform from P_key_points_reference_product (as shown in Fig. 11A) to P_key_points_real_product (as shown in Fig. 12A) may be calculated. In some embodiments, this homography matrix may be denoted as PT_key_points. In some embodiments, this homography matrix can be used to calculate the coordinates in the real image based on the coordinates in the reference image.

[0142] Step 2: Determine the centers of the screw holes

[0143] In some embodiments, screw holes may be located coarsely by using the matrix PT_key_points (which is calculated in Step 1, Part 3) and P_screw_hole_reference_product (the coordinates determined in Step 2, Part 2) as shown in Fig. 12B. That is, the expected locations of the screw holes in the real image may be determined based on the perspective transform and the coordinates of the screw holes in the reference image.

[0144] In some embodiments, small ROIs 1220 may be used to search for the screw holes. In some embodiments, a UNet may be used to detect the centers of the screw holes in the real image. In some embodiments, this UNet may be denoted as UNet_screw_hole. In some embodiments, the training method may be same as that described in step 2 in part 1. In some embodiments, the coordinates may be denoted as P_screw_hole_real_product.

[0145] As shown in Fig. 12C, the dotted bounding box may be used as the ROIs 1220 for screw hole detection. As mentioned earlier, the centers of the ROIs 1220 may be the coordinates of the expected locations of the screw holes calculated above. In some  embodiments, the size of the ROIs 1220 may be 108 *108 pixels, which is smaller than the ROIs 1210 for locating the key points.

[0146] Step 3: Implement light guidance for the screw holes

[0147] In some embodiments, the homography matrix PT_product (which is determined in Step 1, Part 2) and P_screw_hole_real_product (which is determined in step 2, Part 3) may be used to determine the coordinates P_screw_hole_without_product. It means that if the product 120 is removed from the working area 115, the light will be projected onto the supporting surface 105 at the location with the determined coordinates. In this way, if the product 120 is placed on the working area 115, the light will be projected on the expected location (e.g., the locations of the screw holes) .

[0148] In some embodiments, the MLP network determined in Step 2, Part 1 named MLP_Projection_Canvas_Relationship may be used to obtain the coordinates of the screw holes (or the expected locations where light guidance for the screw holes is to be provided) in the canvas image. In some embodiments, the coordinates in the canvas image may be denoted as P_screw_hole_canvas_real_product. In some embodiments, circles centered at the coordinates can be drawn in the canvas image as shown in Fig. 12D and therefore light spots 1230 corresponding to the circles may be projected onto the product 120 to realize the light guidance, as shown in Fig. 12E. In some embodiments, the circles and thus the projected light spots 1230 may have a color other than white.

[0149] Although circles are generated and projected to highlight the screw holes, the present disclosure is not limited thereto. For example, arrow patterns 1240 may be generated and projected onto the product 120 as shown in Fig. 12F.

[0150] Note: the height of a product of one kind may be different from that of a different kind. However, this difference in heights may be automatically handled by the method described above, for example, by the perspective transform PT_Product. In other words, this method can convert all different heights into the same height of the desk (i.e., the supporting surface 105 or the working area 115) . In this way, the system can work even if the product height is different for the new product introduction.

[0151] Further, as shown by Fig. 12A through Fig. 12F, even when the product 120 is placed in the working area 115 with a translation and rotation, the screw holes of the product 120 can still be accurately located and highlighted by the projected light spots 1230 / arrows 1240. In other words, the system can work well in various situations.

[0152] Fig. 13 is a diagram illustrating an exemplary working area with additional operations performed and additional information provided according to an embodiment of the present disclosure. In some embodiments, some subsequent tasks can be implemented after the screw hole detection and / or the light guidance projection. For example, screws may be driven into some of the screw holes S0 through S7 as shown in Fig. 13, and the system may detect / inspect the existence of the screws in the screw holes.

[0153] In some embodiments, the system may detect whether there are screws in the screw holes based on at least the images captured by the camera 111. For example, the system may determine that there are screws in the screw holes S0, S1, S6, and S7 with a very high confidence (1 or close to 1) , and determine that the confidences that there are screws in the screw holes S2 through S5 are very low (close to 0) . In such a case, the system may provide the operator with the related information, for example, the information projected on the right side of the product 120 in the working area 115 as shown in Fig. 13.

[0154] Fig. 14 is a diagram illustrating an exemplary overall procedure 1400 for image projection for AR according to an embodiment of the present disclosure. As shown in Fig. 14, the procedure may begin with step S1410 where an image of a product (e.g., the target 120) to be manufactured may be captured by a camera (e.g., the camera 111) . In some embodiments, this step may be performed as described above in Step 1, Part 3.

[0155] At step S1420, the captured image may be calibrated for camera distortion. The information used for calibration (e.g., one or more distortion coefficients) may be obtained at step S1415 where the distortion coefficients may be obtained from images of a calibration target (e.g., the chessboard like target 400) captured by the camera. In some embodiments, this step may be performed as described above in Step 1, Part 1.

[0156] At step S1430, one or more (e.g., 4) key points (e.g., S0 through S3) on the product in the captured image may be determined by using large ROIs (compared to the ROIs for locating screw holes) . In some embodiments, this step may be performed as described above in Step 1, Part 3.

[0157] At step S1440, the coordinates of the screw holes (e.g., S0 through S7) may be determined in small ROIs (compared to the ROIs for locating the key points) that are centered at expected locations of the screw holes. In some embodiments, the expected  locations of the screw holes may be calculated from the locations of the key points determined at step S1430 and a perspective transform (e.g., PT_key_points mentioned earlier) . In some embodiments, this step may be performed as described above in Step 2, Part 3.

[0158] At step S1450, the coordinates of the screw holes in the image captured when the product is placed on the table may be converted into the coordinates in the image captured when the product is not placed on the table, such that the coordinates of the screw holes on different products with different heights, which means these coordinates are provided in different coordinate systems, may be converted into a same coordinate system. In other words, coordinates of screw hole projection on the desk may be obtained assuming the product is removed from the desk. In some embodiments, this step may be performed as described above in Step 3, Part 3.

[0159] To perform such a conversion, a perspective transform may be pre-calculated at step S1445 where the perspective transform may be obtained by using images with and without a reference product (e.g., a product prototype) . In some embodiments, this step may be performed as described above in Step 1, Part 2.

[0160] At step S1460, the coordinates of the screw holes in the same coordinate system may be mapped to the coordinates in a canvas image, for example, by using an MLP network, such that an image for light guidance may be generated. In some embodiments, this step may be performed as described above in Step 3, Part 3.

[0161] In some embodiments, the MLP network may be obtained (e.g., generated and trained) , for example, by using a 15-by-15 circle array and its projected image at step S1455 as described above. In some embodiments, this step may be performed as described above in Step 2, Part 1.

[0162] At step S1470, the image for light guidance may be projected onto the product, such that the screw holes (or other places where operations are needed) may be highlighted. In some embodiments, this step may be performed as described above in Step 3, Part 3.

[0163] With the embodiments described above, automatic light guidance may be provided accurately even when translations and / or rotations of products to be manufactured occur. Further, it is easier and more convenient for an operator to introduce a new series of product, and no specific expertise knowledge is required. Further, the system is applicable to different series of products with different heights.  Further, it is easier to determine a mapping relationship from a canvas, where an image to be projected is generated, to a projected image, such that a more accurate projection of the image onto a target may be achieved. Furthermore, downstream tasks, such as quality inspection, can be implemented based on this solution.

[0164] Fig. 15 is a flow chart of an exemplary method 1500 for image projection according to an embodiment of the present disclosure. The method 1500 may be performed at an electronic device. The method 1500 may comprise step S1510, S1520, and step S1530. However, the present disclosure is not limited thereto. In some other embodiments, the method 1500 may comprise more steps, less steps, different steps or any combination thereof. Further the steps of the method 1500 may be performed in a different order than that described herein. Further, in some embodiments, a step in the method 1500 may be split into multiple sub-steps and performed by different entities, and / or multiple steps in the method 1500 may be combined into a single step.

[0165] The method 1500 may begin at step S1510, where one or more first points on a first image of a target may be located.

[0166] At step S1520, one or more second points on the first image may be located based on at least the one or more first points.

[0167] At step S1530, a second image may be projected onto the target based on at least the one or more second points.

[0168] In some embodiments, the one or more first points may correspond to one or more features of the target, respectively. In some embodiments, the one or more second points may correspond to one or more locations on the target, respectively, at which one or more operations may be performed by an operator. In some embodiments, at least one of the one or more first points may be located by: searching for a first feature in a first ROI in the first image; and determining a point in the first feature as the first point. In some embodiments, the first ROI may be centered at a location with coordinates same as those of a reference point in a first reference image associated with the target. In some embodiments, the reference point may correspond to the first point.

[0169] In some embodiments, the step of searching for a first feature may comprise: searching for the first feature by using a first U-Net convolutional network. In some embodiments, the step of determining a point in the first feature as the first point may comprise: determining the center of the first feature as the first point by using the first  U-Net convolutional network. In some embodiments, the step of locating one or more second points may comprise: determining a first perspective transform from a first reference image associated with the target to the first image based on one or more first reference points in the first reference image and the one or more first points at least. In some embodiments, the first reference points may correspond to the one or more first points, respectively. In some embodiments, the first perspective transform may be defined by a first homography matrix that is determined based on the one or more first reference points and the one or more first points at least. In some embodiments, the one or more first points may comprise four or more first points that are not collinear.

[0170] In some embodiments, at least one of the one or more second points may be located by: determining an expected location in the first image corresponding to a second point based on at least the first perspective transform and a second reference point in the first reference image, the second reference point corresponding to the second point; searching for a second feature corresponding to the second point in a second ROI in the first image, the second ROI being centered at the expected location; and determining a point in the second feature as the second point.

[0171] In some embodiments, the first ROI may have a larger area than that of the second ROI. In some embodiments, the step of searching for a second feature may comprise: searching for the second feature by using a second U-Net convolution network. In some embodiments, the step of determining a point in the second feature as the second point may comprise: determining the center of the second feature as the second point by using the second U-Net convolutional network.

[0172] In some embodiments, the second image may comprise one or more patterns which, when projected on the target, highlight one or more locations on the target, respectively. In some embodiments, the one or more highlighted locations on the target may be locations where one or more operations are to be performed by an operator. In some embodiments, at least one of the one or more patterns may be generated by: determining a location in a second reference image, the location corresponding to one of the one or more second points; determining a location in the second image, the determined location in the second image corresponding to the determined location in the second reference image; and generating the pattern in the second image based on at least the determined location in the second image. In some embodiments, the first reference image associated with the target may be captured for a reference target that  is placed on a supporting surface. In some embodiments, the second reference image may be captured for the supporting surface without the reference target placed thereon. In some embodiments, the reference target may have a same product specification as the target.

[0173] In some embodiments, at least two of the first image, the first reference image, and the second reference image may be captured from a same viewpoint. In some embodiments, the step of determining a location in a second reference image may comprise: calculating the location in the second reference image based on at least the corresponding second point and a second perspective transform from the first reference image to the second reference image. In some embodiments, the second perspective transform may be defined by a second homography matrix that is calculated based on at least the first reference image and the second reference image. In some embodiments, the second perspective transform may be determined by: capturing the first reference image when one or more first markers are projected onto a reference target placed on a supporting surface; determining, in the first reference image, the locations of the one or more first markers projected on the reference target; capturing the second reference image when one or more first markers are projected onto the supporting surface without the reference target placed thereon; determining, in the second reference image, the locations of the one or more first markers projected on the supporting surface; calculating the second perspective transform based on at least the locations of the one or more first markers projected on the reference target in the first reference image and the locations of the one or more first markers projected on the supporting surface in the second reference image.

[0174] In some embodiments, the one or more first markers may be one or more cross patterns, and the locations of the one or more first markers may be the centers of the one or more cross patterns. In some embodiments, the step of determining a location in the second image may comprise: determining the location in the second image based on at least the determined location in the second reference image and a mapping from locations in the second reference image to locations in the second image. In some embodiments, the mapping from locations in the second reference image to locations in the second image may be determined by: capturing a third reference image when a third image is projected on a supporting surface without any target placed thereon, the third image comprising one or more second markers; determining one or more locations  in the third reference image corresponding to the one or more second markers; and determining a mapping based on at least the one or more determined locations in the third reference image and the locations of the one or more second markers in the third image, as the mapping from locations in the second reference image to locations in the second image. In some embodiments, the one or more second markers may comprise an array of circles.

[0175] In some embodiments, the one or more second markers projected on the supporting surface may cover an area where the target is to be placed. In some embodiments, at least one of the one or more locations in the third reference image may be determined by: labeling a point in a projected second marker in the third reference image; determining a third ROI comprising the projected second marker based on at least the labeled point; and searching for the projected second marker in the third ROI and determining the center of the projected second marker by using a third U-net convolutional network.

[0176] In some embodiments, the step of determining a mapping may comprise: training an MLP network by using the determined one or more locations in the third reference image as inputs and using the locations of the one or more second markers in the third image as expected outputs, such that the trained MLP network is able to be used as the mapping from locations in the second reference image to locations in the second image. In some embodiments, the step of generating the pattern in the second image may comprise at least one of: generating a circle pattern in the second image that is centered at the determined location in the second image; and generating an arrow pattern in the second image that points to the determined location in the second image.

[0177] In some embodiments, the method 1500 may further comprise: calibrating a camera that is used for capturing at least one of the first image, the first reference image, the second reference image, and the third reference image. In some embodiments, the step of calibrating the camera may comprise: capturing one or more images of a calibration target by using the camera when the calibration target is placed at one or more places on the supporting surface; calculating a distortion of the camera based on at least the one or more images of the calibration target and one or more known characteristics of the calibration target; and compensating one or more images captured by the camera for the calculated distortion.

[0178] In some embodiments, the second image, when projected onto the target, may further indicate at least one of: information related to the target; information related to the operations to be performed on the target; information related to status of a camera used for capturing the first image and / or a projector used for projecting the second image; and one or more icons for interaction with an operator.

[0179] Fig. 16 schematically shows an embodiment of an arrangement 1600 which may be used in an electronic device according to an embodiment of the present disclosure. Comprised in the arrangement 1600 are a processing unit 1606, e.g., with a Digital Signal Processor (DSP) or a Central Processing Unit (CPU) . The processing unit 1606 may be a single unit or a plurality of units to perform different actions of procedures described herein. The arrangement 1600 may also comprise an input unit 1602 for receiving signals from other entities, and an output unit 1604 for providing signal (s) to other entities. The input unit 1602 and the output unit 1604 may be arranged as an integrated entity or as separate entities.

[0180] Furthermore, the arrangement 1600 may comprise at least one computer program product 1608 in the form of a non-volatile or volatile memory, e.g., an Electrically Erasable Programmable Read-Only Memory (EEPROM) , a flash memory and / or a hard drive. The computer program product 1608 comprises a computer program 1610, which comprises code / computer readable instructions, which when executed by the processing unit 1606 in the arrangement 1600 causes the arrangement 1600 and / or the electronic device in which it is comprised to perform the actions, e.g., of the procedure described earlier in conjunction with Fig. 4A through Fig. 15 or any other variant.

[0181] The computer program 1610 may be configured as a computer program code structured in computer program modules 1610A, 1610B, and 1610C. Hence, in an exemplifying embodiment when the arrangement 1600 is used in an electronic device, the code in the computer program of the arrangement 1600 includes: a module 1610A configured to locate one or more first points on a first image of a target; a module 1610B configured to locate one or more second points on the first image based on at least the one or more first points; and a module 1610C configured to project a second image onto the target based on at least the one or more second points.

[0182] The computer program modules could essentially perform the actions of the flow illustrated in Fig. 4A through Fig. 15, to emulate the electronic device. In other words,  when the different computer program modules are executed in the processing unit 1606, they may correspond to different modules in the electronic device.

[0183] Although the code means in the embodiments disclosed above in conjunction with Fig. 16 are implemented as computer program modules which when executed in the processing unit causes the arrangement to perform the actions described above in conjunction with the figures mentioned above, at least one of the code means may in alternative embodiments be implemented at least partly as hardware circuits.

[0184] The processor may be a single CPU (Central processing unit) , but could also comprise two or more processing units. For example, the processor may include general purpose microprocessors; instruction set processors and / or related chips sets and / or special purpose microprocessors such as Application Specific Integrated Circuit (ASICs) . The processor may also comprise board memory for caching purposes. The computer program may be carried by a computer program product connected to the processor. The computer program product may comprise a computer readable medium on which the computer program is stored. For example, the computer program product may be a flash memory, a Random-access memory (RAM) , a Read-Only Memory (ROM) , or an EEPROM, and the computer program modules described above could in alternative embodiments be distributed on different computer program products in the form of memories within the electronic device.

[0185] Correspondingly to the method 1500 as described above, an electronic device 1700 is provided. Fig. 17 is a block diagram of an exemplary electronic device 1700 according to an embodiment of the present disclosure. The electronic device 1700 can comprise e.g., the arrangement 1600.

[0186] The electronic device 1700 can be configured to perform the method 1500 as described above in connection with Fig. 15. As shown in Fig. 17, the electronic device 1700 may comprise a first locating module 1710 configured to locate one or more first points on a first image of a target; a second locating module 1720 configured to locate one or more second points on the first image based on at least the one or more first points; and a projecting module 1730 configured to project a second image onto the target based on at least the one or more second points.

[0187] The above modules 1710, 1720, and / or 1730 can be implemented as a pure hardware solution or as a combination of software and hardware, e.g., by one or more of: a processor or a micro-processor and adequate software and memory for storing of  the software, a Programmable Logic Device (PLD) or other electronic component (s) or processing circuitry configured to perform the actions described above, and illustrated, e.g., in Fig. 15. Further, the electronic device 1700 may comprise one or more further modules, each of which may perform any of the steps of the method 1500 described with reference to Fig. 15.

[0188] The present disclosure is described above with reference to the embodiments thereof. However, those embodiments are provided just for illustrative purpose, rather than limiting the present disclosure. The scope of the disclosure is defined by the attached claims as well as equivalents thereof. Those skilled in the art can make various alternations and modifications without departing from the scope of the disclosure, which all fall into the scope of the disclosure. Abbreviation          Explanation AR                    Augmented Reality ROI                   Region of Interest MLP                   Multi-Layer Perceptron

Claims

1.A method (1500) for image projection, the method (1500) comprising:locating (S1430, S1510) one or more first points (K0-K3) on a first image of a target (120) ;locating (S1440, S1520) one or more second points (S0-S7) on the first image based on at least the one or more first points (K0-K3) ; andprojecting (S1470, S1530) a second image onto the target (120) based on at least the one or more second points (S0-S7) .2.The method (1500) of claim 1, wherein the one or more first points (K0-K3) correspond to one or more features (121) of the target (120) , respectively.3.The method (1500) of claim 1 or 2, wherein the one or more second points (S0-S7) correspond to one or more locations on the target (120) , respectively, at which one or more operations are to be performed by an operator.4.The method (1500) of any of claims 1 to 3, wherein at least one of the one or more first points (K0-K3) is located by:searching for a first feature (121) in a first Region of Interest (ROI) (1210) in the first image; anddetermining a point in the first feature (121) as the first point (K0, K1, K2, K3) .5.The method (1500) of claim 4, wherein the first ROI (1210) is centered at a location with coordinates same as those of a reference point in a first reference image associated with the target (120) ,wherein the reference point corresponds to the first point (K0, K1, K2, K3) .6.The method (1500) of claim 4 or 5, wherein the step of searching for a first feature (121) comprises: searching for the first feature (121) by using a first U-Net convolutional network, and / orwherein the step of determining a point in the first feature (121) as the first point (K0, K1, K2, K3) comprises: determining the center of the first feature (121) as the first point (K0, K1, K2, K3) by using the first U-Net convolutional network.7.The method (1500) of any of claims 1 to 6, wherein the step of locating (S1440, S1520) one or more second points (S0-S7) comprises:determining a first perspective transform from a first reference image associated with the target (120) to the first image based on one or more first reference points in the first reference image and the one or more first points (K0-K3) at least,wherein the first reference points correspond to the one or more first points (K0-K3) , respectively.8.The method (1500) of claim 7, wherein the first perspective transform is defined by a first homography matrix that is determined based on the one or more first reference points and the one or more first points (K0-K3) at least.9.The method (1500) of claim 7 or 8, wherein the one or more first points (K0-K3) comprise four or more first points (K0-K3) that are not collinear.10.The method (1500) of any of claims 7 to 9, wherein at least one of the one or more second points (S0-S7) is located by:determining an expected location in the first image corresponding to a second point (S0, S1, S2, S3, S4, S5, S6, S7) based on at least the first perspective transform and a second reference point in the first reference image, the second reference point corresponding to the second point (S0, S1, S2, S3, S4, S5, S6, S7) ;searching for a second feature (123) corresponding to the second point (S0, S1, S2, S3, S4, S5, S6, S7) in a second ROI (1220) in the first image, the second ROI (1220) being centered at the expected location; anddetermining a point in the second feature (123) as the second point (S0, S1, S2, S3, S4, S5, S6, S7) .11.The method (1500) of any of claims 4 to 10, wherein the first ROI (1210) has a larger area than that of the second ROI (1220) .12.The method (1500) of claim 10 or 11, wherein the step of searching for a second feature (123) comprises: searching for the second feature (123) by using a second U-Net convolution network, and / orwherein the step of determining a point in the second feature (123) as the second point (S0, S1, S2, S3, S4, S5, S6, S7) comprises: determining the center of the second feature (123) as the second point (S0, S1, S2, S3, S4, S5, S6, S7) by using the second U-Net convolutional network.13.The method (1500) of any of claims 1 to 12, wherein the second image comprises one or more patterns (S0′-S7′) which, when projected on the target (120) , highlight one or more locations (S0-S7) on the target (120) , respectively,wherein the one or more highlighted locations (S0-S7) on the target (120) are locations (S0-S7) where one or more operations are to be performed by an operator.14.The method (1500) of claim 13, wherein at least one of the one or more patterns (S0′-S7′) is generated by:determining a location in a second reference image, the location corresponding to one of the one or more second points (S0-S7) ;determining a location in the second image, the determined location in the second image corresponding to the determined location in the second reference image; andgenerating the pattern in the second image based on at least the determined location in the second image.15.The method (1500) of claim 14, wherein the first reference image associated with the target (120) is captured for a reference target that is placed on a supporting surface (105) ,wherein the second reference image is captured for the supporting surface (105) without the reference target placed thereon.16.The method (1500) of claim 15 wherein the reference target has a same product specification as the target (120) .17.The method (1500) of any of claims 14 to 16, wherein at least two of the first image, the first reference image, and the second reference image are captured from a same viewpoint.18.The method (1500) of any of claims 14 to 17, wherein the step of determining a location in a second reference image comprises:calculating the location in the second reference image based on at least the corresponding second point and a second perspective transform from the first reference image to the second reference image.19.The method (1500) of claim 18, wherein the second perspective transform is defined by a second homography matrix that is calculated based on at least the first reference image and the second reference image.20.The method (1500) of claim 18 or 19, wherein the second perspective transform is determined by:capturing the first reference image when one or more first markers (P0-P3) are projected onto a reference target placed on a supporting surface (105) ;determining, in the first reference image, the locations of the one or more first markers (P0″-P3″) projected on the reference target;capturing the second reference image when one or more first markers (P0-P3) are projected onto the supporting surface (105) without the reference target placed thereon;determining, in the second reference image, the locations of the one or more first markers (P0′-P3′) projected on the supporting surface (105) ;calculating the second perspective transform based on at least the locations of the one or more first markers (P0″-P3″) projected on the reference target in the first reference image and the locations of the one or more first markers (P0′-P3′) projected on the supporting surface (105) in the second reference image.21.The method (1500) of any of claims 19 to 20, wherein the one or more first markers (P0-P3) are one or more cross patterns, and the locations of the one or more first markers (P0-P3) are the centers of the one or more cross patterns.22.The method (1500) of any of claims 14 to 21, wherein the step of determining a location in the second image comprises:determining the location in the second image based on at least the determined location in the second reference image and a mapping from locations in the second reference image to locations in the second image.23.The method (1500) of claim 22, wherein the mapping from locations in the second reference image to locations in the second image is determined by:capturing a third reference image when a third image is projected on a supporting surface (105) without any target (120) placed thereon, the third image comprising one or more second markers (700) ;determining one or more locations (715) in the third reference image corresponding to the one or more second markers (700) ; anddetermining a mapping based on at least the one or more determined locations (715) in the third reference image and the locations of the one or more second markers (700) in the third image, as the mapping from locations in the second reference image to locations in the second image.24.The method (1500) of claim 23, wherein the one or more second markers comprise an array of circles.25.The method (1500) of claim 23 or 24, wherein the one or more second markers projected on the supporting surface cover an area where the target (120) is to be placed.26.The method (1500) of any of claims 23 to 25, wherein at least one of the one or more locations in the third reference image is determined by:labeling a point in a projected second marker (700′) in the third reference image;determining a third ROI comprising the projected second marker (700′) based on at least the labeled point; andsearching for the projected second marker (700′) in the third ROI and determining the center (715) of the projected second marker (700′) by using a third U-net convolutional network.27.The method (1500) of any of claims 23 to 26, wherein the step of determining a mapping comprises:training a Multi-Layer Perceptron (MLP) network by using the determined one or more locations in the third reference image (700′) as inputs and using the locations of the one or more second markers (700) in the third image as expected outputs, such that the trained MLP network is able to be used as the mapping from locations in the second reference image to locations in the second image.28.The method (1500) of any of claims 14 to 27, wherein the step of generating the pattern in the second image comprises at least one of:generating a circle pattern (S0′-S7′) in the second image that is centered at the determined location in the second image; andgenerating an arrow pattern in the second image that points to the determined location in the second image.29.The method (1500) of any of claims 1 to 28, further comprising:calibrating a camera (111) that is used for capturing at least one of the first image, the first reference image, the second reference image, and the third reference image.30.The method (1500) of claim 29, wherein the step of calibrating the camera (111) comprises:capturing one or more images of a calibration target (400) by using the camera (111) when the calibration target (400) is placed at one or more places on the supporting surface (105) ;calculating a distortion of the camera (111) based on at least the one or more images of the calibration target (400) and one or more known characteristics of the calibration target (400) ; andcompensating one or more images captured by the camera (111) for the calculated distortion.31.The method (1500) of any of claims 1 to 31, wherein the second image, when projected onto the target (120) , further indicates at least one of:- information related to the target (120) ;- information related to the operations to be performed on the target (120) ;- information related to status of a camera (111) used for capturing the first image and / or a projector (113) used for projecting the second image; and- one or more icons for interaction with an operator.32.An electronic device (1600, 1700) , comprising:a processor (1606) ;a memory (1608) storing instructions which, when executed by the processor (1606) , cause the processor (1606) to:locate one or more first points (K0-K3) on a first image of a target (120) ;locate one or more second points (S0-S7) on the first image based on at least the one or more first points (K0-K3) ; andproject a second image onto the target (120) based on at least the one or more second points (S0-S7) .33.The electronic device (1600, 1700) of claim 32, wherein the instructions, when executed by the processor (1606) , cause the processor (1606) further to perform the method (1500) of any of claims 2 to 31.34.A computer program (1610) comprising instructions which, when executed by at least one processor (1606) , cause the at least one processor (1606) to carry out the method (1500) of any of claims 1 to 31.35.A carrier (1608) containing the computer program (1610) of claim 34, wherein the carrier (1608) is one of an electronic signal, optical signal, radio signal, or computer readable storage medium.36.A system (10) for image projection, the system (10) comprising:a supporting surface (105) on which a target (120) is to be placed;a projector (113) configured to project an image onto the supporting surface (105) and / or the target (120) ;a camera (111) configured to capture an image of at least one of the supporting surface (105) , the target (120) , and the projected image;a processor (1106) ;a memory (1108) storing instructions which, when executed by the processor (1106) , cause the processor (1106) to:locate one or more first points (K0-K3) on a first image of the target (120) captured by the camera (111) ;locate one or more second points (S0-S7) on the first image based on at least the one or more first points (K0-K3) ; andproject, by the projector (113) , a second image onto the target (120) based on at least the one or more second points (S0-S7) .37.The system (10) of claim 36, wherein the instructions, when executed by the processor (1606) , cause the processor (1606) further to perform the method (1500) of any of claims 2 to 31.

Citation Information

Patent Citations

  • Superficial venous image augmented reality method and device based on perspective projection

    CN103226817A

  • Hierarchical multi-vision positioning method, system and device

    CN108074264A

  • Mounting positioning method and device and optical projection device

    CN110596997A

  • Screw hole positioning method and device, computer equipment and storage medium

    CN116051928A

  • Hole position detecting device

    JP1997042915A