An image automatic acquisition and labeling method, system, device and storage medium

By displaying sample images on an LED display screen and using an industrial camera for image acquisition and automatic annotation, the problems of high cost, low efficiency, and insufficient accuracy in image acquisition and annotation in LED display defect detection models are solved, achieving efficient automatic annotation.

CN122369007APending Publication Date: 2026-07-10BEIJING TRICOLOR TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING TRICOLOR TECH
Filing Date
2026-04-10
Publication Date
2026-07-10

Smart Images

  • Figure CN122369007A_ABST
    Figure CN122369007A_ABST
Patent Text Reader

Abstract

The application provides an image automatic acquisition and labeling method, system, device and storage medium. The method can control the acquisition end to automatically acquire the sample image displayed on the display end, and convert the first position information of the to-be-recognized object in the sample image according to the coordinate conversion relationship between the sample image and the target image automatically acquired, to obtain the second position information of the to-be-recognized object in the target image, so as to automatically label the to-be-recognized object contained in the target image based on the obtained second position information, and effectively solve the problems of high cost, low efficiency and insufficient precision of image acquisition and labeling in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and more specifically, to an automatic image acquisition and annotation method, system, device, and storage medium. Background Technology

[0002] When training a defect detection model for LED (Light Emitting Diode) displays, a large amount of "image-annotation" pairing data is required. The images in the pairing data refer to the images of the LED display taken by a camera, while the annotations in the pairing data refer to the location markers of defects on the LED display in the images taken by the camera.

[0003] Currently, existing technologies mainly rely on manual annotation to acquire this data (i.e., annotators manually outline defect areas in camera-captured images). However, manual annotation suffers from high costs, long processing times, and poor consistency in annotation standards, making it difficult to meet the demands of industrial production for large-scale, high-quality datasets. Summary of the Invention

[0004] In view of this, this application provides an automatic image acquisition and annotation method, system, device, and storage medium, which can control the acquisition end to automatically acquire sample images displayed on the display end, and convert the first position information of the object to be identified in the sample image according to the coordinate transformation relationship between the sample image and the automatically acquired target image to obtain the second position information of the object to be identified in the target image. Based on the obtained second position information, the object to be identified contained in the target image is automatically annotated, which effectively solves the problems of high cost, low efficiency, and insufficient accuracy of image acquisition and annotation in the prior art.

[0005] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.

[0006] In a first aspect, embodiments of this application provide an automatic image acquisition and annotation method, the automatic image acquisition and annotation method comprising: The display terminal shows a sample image containing the object to be identified; The target image corresponding to the sample image is obtained by acquiring the sample image through the acquisition terminal; Based on the first position information of the object to be identified in the sample image, a target geometric unit matching the first position information is determined from a plurality of geometric units; wherein, the plurality of geometric units represent a plurality of geometric units that make up the display area of ​​the display terminal; Based on the local mapping relationship between the target geometric unit and the first coordinate system and the second coordinate system, the first position information is transformed to obtain the second position information of the object to be identified in the second coordinate system; wherein, the first coordinate system represents the coordinate system corresponding to the sample image, and the second coordinate system represents the coordinate system corresponding to the target image; Based on the second location information, the object to be identified contained in the target image is automatically labeled to generate a labeled sample image corresponding to the sample image.

[0007] Secondly, embodiments of this application provide an automatic image acquisition and annotation system, which includes: a display terminal, an acquisition terminal, and a control terminal; wherein the control terminal is used for: The display terminal displays a sample image containing the object to be identified. The target image corresponding to the sample image is obtained by acquiring the sample image through the acquisition terminal; Based on the first position information of the object to be identified in the sample image, a target geometric unit matching the first position information is determined from a plurality of geometric units; wherein, the plurality of geometric units represent a plurality of geometric units that make up the display area of ​​the display terminal; Based on the local mapping relationship between the target geometric unit and the first coordinate system and the second coordinate system, the first position information is transformed to obtain the second position information of the object to be identified in the second coordinate system; wherein, the first coordinate system represents the coordinate system corresponding to the sample image, and the second coordinate system represents the coordinate system corresponding to the target image; Based on the second location information, the object to be identified contained in the target image is automatically labeled to generate a labeled sample image corresponding to the sample image.

[0008] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described automatic image acquisition and annotation method.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described automatic image acquisition and annotation method.

[0010] The technical solutions provided by the embodiments of this application may include the following beneficial effects: This application provides an automatic image acquisition and annotation method, system, device, and storage medium. It can control the acquisition end to automatically acquire sample images displayed on the display end, and convert the first position information of the object to be identified in the sample image according to the coordinate transformation relationship between the sample image and the automatically acquired target image to obtain the second position information of the object to be identified in the target image. Based on the obtained second position information, the object to be identified contained in the target image is automatically annotated, which effectively solves the problems of high cost, low efficiency, and insufficient accuracy of image acquisition and annotation in the prior art. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This illustration shows a structural schematic diagram of an automatic image acquisition and annotation system provided in an embodiment of this application; Figure 2 A flowchart illustrating an automatic image acquisition and annotation method provided in an embodiment of this application is shown. Figure 3 This illustration shows a schematic diagram of a chessboard calibration diagram with direction indicator marks provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0014] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0015] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0016] One embodiment of the automatic image acquisition and annotation method in this application can be run in an automatic image acquisition and annotation system; wherein, Figure 1 This paper illustrates a schematic diagram of the structure of an automatic image acquisition and annotation system provided in an embodiment of this application, as shown below. Figure 1 As shown, the automatic image acquisition and annotation system includes a display terminal, an acquisition terminal, and a control terminal.

[0017] Specifically, the display end may include an LED display screen and a display control module, while the acquisition end may include an industrial camera (e.g., with a resolution of no less than 20 megapixels, supporting autofocus and exposure control) and a video stream interface. During the automatic image acquisition phase, the control end can control the display end to automatically generate and display sample images containing the object to be identified (e.g., bright spots, dark spots, dead pixels, missing lines, color deviations, and other display defects). The control end controls the industrial camera in the acquisition end to acquire images of the sample images displayed on the display end, and receives the target images acquired by the acquisition end through the video stream interface in the acquisition end. During the automatic image annotation phase, the control end can transform the first position information of the object to be identified in the sample image according to the coordinate transformation relationship between the sample image and the automatically acquired target image, obtaining the second position information of the object to be identified in the target image. Based on the obtained second position information, the control end automatically annotates the object to be identified in the target image, effectively solving the problems of high cost, low efficiency, and insufficient accuracy in image acquisition and annotation in existing technologies.

[0018] It should be noted that when setting up the above-mentioned automatic image acquisition and annotation system, the industrial camera on the acquisition end side can be fixedly mounted on an adjustable bracket. Depending on the size of the display screen on the display end side and the shooting requirements, multiple suitable camera fixing positions can be designed (for example, for a 60-inch display screen, five shooting positions can be designed: one in the center and one at each of the four corners to ensure coverage of the entire display screen area) as shooting positions for the industrial camera on the display screen. The specific number and installation position of the industrial camera on the acquisition end side are not subject to mandatory limitations in this application embodiment.

[0019] It should be noted that the display end can maintain a communication connection with the control end through interfaces such as HDMI (High Definition Multimedia Interface) and DP (DisplayPort), while the acquisition end can maintain a communication connection with the control end through the aforementioned video stream interfaces. The control end integrates the linkage function of image switching and camera shooting, which can control the display end to switch to display different sample images at any time according to user needs. It can also automatically calibrate the camera parameters (focal length, exposure time, white balance) in the industrial camera after the shooting position of the industrial camera on the acquisition end is adjusted, so as to avoid the parameter fluctuations affecting the image quality.

[0020] To facilitate understanding of the embodiments of this application, a detailed description of an automatic image acquisition and annotation method, system, device, and storage medium provided in the embodiments of this application is provided below.

[0021] Reference Figure 2 As shown, Figure 2 The diagram illustrates a flowchart of an automatic image acquisition and annotation method provided in an embodiment of this application, wherein the automatic image acquisition and annotation method includes steps S201-S205; specifically: S201, Displays a sample image containing the object to be identified via a display terminal.

[0022] Here, the object to be identified can refer to different types of display screen defects (equivalent to the image targets that the display screen defect detection model needs to identify from the input image when training the display screen defect detection model later). For example, the object to be identified can include, but is not limited to, display screen defects such as bright spots, dark spots, dead pixels, missing lines, and color deviations.

[0023] Specifically, as an optional embodiment, the control terminal can control the display terminal to display the above sample image through the following steps a1-a2: Step a1: The sample image is generated by combining various image parameters in different ways using an automated script running on the display terminal.

[0024] Here, the aforementioned image parameters include the type, quantity, and location of the object to be identified. That is, when generating sample images, the automated script can automatically determine the type of display defect (i.e., the type of object to be identified), the quantity of defects (i.e., the quantity of objects to be identified), and the location of each display defect in the sample image (i.e., the location of the object to be identified), thereby generating multiple sample images of different styles (i.e., the defect type, quantity, and location contained in different sample images are different).

[0025] It should be noted that the above-mentioned automated scripts can be pre-written using programming languages ​​such as Python and MATLAB. The resolution of the generated sample images can be consistent with the physical resolution of the display screen on one side of the display terminal.

[0026] Step a2: Display the sample image through the display terminal.

[0027] Specifically, when multiple sample images of different styles are generated through the aforementioned automated script, the control terminal can control the display terminal to automatically switch and display different sample images in a preset order (e.g., the order in which the sample images were generated).

[0028] It should be noted that during the process of generating sample images, the control terminal can also synchronously record key information of the object to be identified in the currently generated sample image through the aforementioned automated script on the display terminal (such as the position coordinates of multiple key coordinate points corresponding to the object to be identified in the sample image, the type and quantity of the object to be identified contained in the sample image, etc.).

[0029] S202, the sample image is acquired by the acquisition terminal to obtain the target image corresponding to the sample image.

[0030] Here, the control terminal can use the industrial camera in the acquisition terminal to capture the sample image displayed on the display terminal to obtain the target image corresponding to the sample image, and receive the target image captured by the industrial camera through the video stream interface in the acquisition terminal.

[0031] It should be noted that when the display terminal automatically switches to display different sample images in a preset order, whenever the display terminal switches to a sample image, the control terminal can synchronously trigger the acquisition terminal to acquire the image of the currently switched sample image, obtain the target image corresponding to the sample image, and automatically store the acquired target image in the control terminal according to the preset naming rules (e.g., shooting position + image number).

[0032] Specifically, as an optional embodiment, in order to reduce latency and improve the consistency between "sample image and target image", a buffered frame dropping and latency compensation strategy can be adopted: after switching the display content on the display end, multiple grab reads are performed to clear the cache, and then a retrieve operation is performed to read the captured image of the previous frame as the valid captured image (i.e., the target image corresponding to the sample image displayed after switching the display content).

[0033] S203, based on the first position information of the object to be identified in the sample image, determine the target geometric unit that matches the first position information from multiple geometric units.

[0034] Here, in this embodiment of the application, considering that there are pitch angles, yaw angles and lens distortions between the industrial camera and the display screen of the display terminal, the target image captured by the camera may contain complex perspective and nonlinear deformations, so that the mapping relationship between different local areas on the display screen is different between the two coordinate systems (i.e., the coordinate systems corresponding to the sample image and the target image respectively). Therefore, unlike the prior art which uses a single global homography matrix to represent the overall mapping relationship between the two coordinate systems, this embodiment of the application divides the display area on the display terminal (i.e., the effective display area on the display screen) into multiple geometric units (i.e., the multiple geometric units represent multiple geometric units that make up the display area of ​​the display terminal), and determines the local mapping relationship between each geometric unit and the two coordinate systems. Thus, in the automatic annotation stage, it is necessary to first determine the target geometric unit that matches the position of the object to be identified from the multiple geometric units.

[0035] It should be noted that the geometric unit can be a square grid unit or a triangular triangular unit. The specific shape of the geometric unit is not subject to any mandatory limitation in the embodiments of this application.

[0036] Specifically, the first location information mentioned above includes multiple key coordinate points of the object to be identified in the sample image. These multiple key coordinate points can be determined based on the bounding rectangle of the object to be identified in the sample image. For example, the multiple key coordinate points may include, but are not limited to, the center point, the upper left corner, and the lower right corner of the bounding rectangle.

[0037] In addition, the aforementioned first position information may also include: x_n, y_n, w_n, and h_n; where x_n represents the normalized value of the abscissa of the center point of the circumscribed rectangle, y_n represents the normalized value of the ordinate of the center point of the circumscribed rectangle, w_n represents the normalized value of the width of the circumscribed rectangle, and h_n represents the normalized value of the height of the circumscribed rectangle. Based on the aforementioned first position information, the position of the object to be identified in the sample image can also be determined.

[0038] Here, when performing step S203, since the first location information may contain multiple key coordinate points, as an optional embodiment, the control terminal can determine, for each key coordinate point, from the multiple geometric units, that the geometric unit matching the key coordinate point belongs to the target geometric unit. That is, the target geometric unit may be one or multiple, and this embodiment of the application does not limit it in any way.

[0039] It should be noted that since the above sample image is displayed on the display end, the position coordinates of the above key coordinate points in the first coordinate system (that is, the coordinate system on one side of the display end) are known, and the position coordinates of each geometric unit in the first coordinate system are also known (that is, the display area is divided into multiple geometric units). Thus, when the control end executes step S203, it can determine the target geometric unit to which the key coordinate point belongs from the above multiple geometric units based on the position coordinates of a key coordinate point in the first coordinate system.

[0040] S204, based on the local mapping relationship between the target geometric unit and the first coordinate system and the second coordinate system, the first position information is transformed to obtain the second position information of the object to be identified in the second coordinate system.

[0041] Here, the first coordinate system represents the coordinate system corresponding to the sample image (which is also equivalent to the coordinate system on the display side), and the second coordinate system represents the coordinate system corresponding to the target image (which is also equivalent to the coordinate system on the acquisition side).

[0042] In this embodiment of the application, as an optional embodiment, the control terminal can first determine the local mapping relationship between the first coordinate system and the second coordinate system for each geometric unit through the steps shown in b1-b4. Then, the target geometric unit determined in step S203 can be used as an index to determine the local mapping relationship corresponding to the target geometric unit from the local mapping relationships corresponding to all geometric units. Specifically: Step b1: Display a grid calibration diagram with direction indicator marks on the display terminal.

[0043] Here, the control terminal can pre-control the display terminal to display the above-mentioned grid calibration map before executing step S201 (equivalent to pre-determining the coordinate mapping relationship between the first coordinate system on the display terminal side and the second coordinate system on the acquisition terminal side before performing automatic image acquisition and annotation), or it can control the display terminal to display the grid calibration map with the same resolution as the sample image after executing step S202, and control the acquisition terminal to acquire the above-mentioned grid calibration map at the same shooting angle (i.e., the same shooting angle as the sample image).

[0044] Here, the grid calibration diagram includes multiple marker points. The position coordinates of the marker points in the grid calibration diagram can be used to represent the position coordinates of the marker points in the first coordinate system (that is, the screen standard coordinate system corresponding to the LED display screen). The types and number of marker points included in different types of grid calibration diagrams are different, and this application embodiment does not impose any limitations on this.

[0045] Specifically, when the grid calibration map is a checkerboard calibration map, the above-mentioned multiple marker points are the interior corner points of the checkerboard calibration map; where, if the checkerboard calibration map contains A×B (i.e., A rows and B columns) of black and white squares, then all the marker points included in the checkerboard calibration map can be represented as: an array of interior corner points of (A-1)×(B-1) (the interior corner points are the remaining intersection points in the checkerboard calibration map excluding the four edge corner points).

[0046] Specifically, when the grid calibration map belongs to the ArUco or AprilTag array, the above-mentioned multiple marker points belong to the visual reference markers in the ArUco or AprilTag array; among them, unlike the checkerboard calibration map, in the ArUco or AprilTag array, each visual reference marker carries a unique ID identifier for the camera to identify and locate.

[0047] Here, the direction indicator is used to indicate the sorting method of the above-mentioned multiple marker points in the grid calibration map. On the control end side, the control end can obtain the position coordinates of each marker point in the grid calibration map displayed on the display end through the display control module on the display end side according to the sorting method indicated by the direction indicator, thereby obtaining the position coordinates of each marker point in the first coordinate system.

[0048] It should be noted that, to more intuitively represent the position coordinates of the aforementioned multiple marker points in the first coordinate system (i.e., their position coordinates in the grid calibration diagram), as an optional embodiment, it is preferable to set the aforementioned direction indicator in the upper left corner of the grid calibration diagram (equivalent to indicating that the control terminal starts detecting marker points from the upper left corner of the grid calibration diagram). This allows the acquired multiple marker points to be arranged sequentially from left to right and from top to bottom, forming a positional relationship with the aforementioned grid calibration diagram. Figure 1 A corresponding matrix of marker points.

[0049] Specifically, when the aforementioned grid calibration map is a checkerboard calibration map, the aforementioned direction indicator is used to indicate the starting mark point located in the upper left corner of the checkerboard calibration map; wherein, as an optional embodiment, Figure 3 This illustration shows a schematic diagram of a checkerboard calibration diagram with direction indicator marks provided in an embodiment of this application, such as... Figure 3As shown, different asymmetrical directional markers can be set at the four corners of the chessboard calibration diagram: a triangle marker with the word "START" is set at the upper left corner of the chessboard calibration diagram (to indicate that the control end will start detecting the inner corner points of the chessboard calibration diagram from the upper left corner), a circle marker is set at the upper right corner (to indicate that the row edge of the chessboard calibration diagram has been reached), a square marker is set at the lower left corner (to indicate that the column edge of the chessboard calibration diagram has been reached), and a "+" marker is set at the lower right corner (to indicate that the detection of the inner corner points can be ended).

[0050] Specifically, when the above-mentioned grid calibration map belongs to the ArUco or AprilTag array, the above-mentioned direction indicator is used to indicate the starting naming direction corresponding to the ID of the visual reference mark; for example, in the ArUco or AprilTag array, the ID of the visual reference mark is sequentially increased from left to right and from top to bottom (e.g., the ID of the visual reference mark located in the upper left corner is 0), and at this time the above-mentioned direction indicator can be used to indicate the upper left corner direction in the ArUco or AprilTag array.

[0051] Step b2: Acquire an image of the grid calibration map through the acquisition terminal to obtain the captured image corresponding to the grid calibration map.

[0052] Here, the control terminal can use the industrial camera in the acquisition terminal to capture the grid calibration map displayed on the display terminal, obtain the captured image corresponding to the grid calibration map, and receive the captured image obtained by the industrial camera through the video stream interface in the acquisition terminal.

[0053] It should be noted that since the image content of the captured image is the same as that of the grid calibration map, the control terminal can also obtain the position coordinates of the above-mentioned marker points in the captured image as the position coordinates of the above-mentioned marker points in the second coordinate system (that is, the position coordinates in the coordinate system on the acquisition terminal side).

[0054] Step b3: Using the marked points as vertices of the geometric units, divide the display area on the display terminal into the multiple geometric units.

[0055] Specifically, since the position coordinates of the marker points in the two coordinate systems (i.e., the first coordinate system and the second coordinate system) are known, when dividing the geometric units, the marker points need to be used as the vertices of the geometric units, so as to ensure that each geometric unit contains multiple marker points that can be used to solve the above local mapping relationship.

[0056] Step b4: For each geometric unit, based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, determine the local mapping relationship between the geometric unit and the second coordinate system.

[0057] Here, in step b4, as an optional embodiment, if the geometric unit is a square mesh unit, the local mapping relationship corresponding to the geometric unit can be solved in the manner shown in step c1 below: Step c1: Based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, solve for the perspective transformation matrix corresponding to the geometric unit as the local mapping relationship.

[0058] Specifically, when the geometric unit is a square grid unit, the marker points contained in the geometric unit are the four vertices of the grid unit; that is, when the geometric unit is a square grid unit, the control terminal can obtain four sets of corresponding points between the first coordinate system and the second coordinate system (a set of corresponding points is used to represent the position coordinates of a marker point in the first coordinate system and the second coordinate system respectively); among them, the perspective transformation matrix corresponding to the geometric unit can be obtained by using the get Perspective Transform function (a function provided by computer vision libraries such as OpenCV, whose core function is to solve a perspective transformation matrix through four sets of corresponding points).

[0059] Specifically, when the control terminal determines the aforementioned local mapping relationship according to the method shown in step c1 above, the perspective transformation matrix corresponding to the target geometric unit is denoted as H_{u,v}. Then, for any key coordinate point s in the target geometric unit, given that the position coordinates of the key coordinate point s in the first coordinate system are (x_s, y_s), the second position information (x_s, y_s) of the key coordinate point s in the second coordinate system can be solved according to the method shown in formulas 1-3 below. y_ ): Formula 1; x_ = Formula 2; y_ = Formula 3; Where x_s is the x-coordinate of the key coordinate point s in the first coordinate system, and y_s is the y-coordinate of the key coordinate point s in the first coordinate system; This is the scaling factor in homogeneous coordinates. The difference between homogeneous and ordinary coordinates is that ordinary coordinates use the tuple (x_s, y_s) to represent the key coordinate point s, while homogeneous coordinates use... This triple represents the key coordinate point s, and w can be determined based on the depth of the key coordinate point s from the display screen. H_{u,v} is the perspective transformation matrix corresponding to the target geometric unit to which the key coordinate point s belongs; x_ The x-coordinate of the key coordinate point s in the second coordinate system is y_. It is the ordinate of the key coordinate point s in the second coordinate system; It is x_ A corresponding intermediate calculation result, It is y_ A corresponding intermediate calculation result; It is a vector Transpose of; It is a vector The transpose of .

[0060] Here, in step b4, as an optional embodiment, if the geometric element is a triangular element of a triangle, the above-mentioned local mapping relationship corresponding to the geometric element can be solved in the manner shown in step d1 below: Step d1: Based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the target image, solve for the affine transformation model corresponding to the geometric unit as the local mapping relationship.

[0061] Here, the affine transformation model consists of parameters A and t, where parameter A represents the linear transformation part of the affine transformation model and parameter t represents the translation part of the affine transformation model.

[0062] Specifically, when a geometric unit belongs to a triangular unit of a triangle, the marked points contained in the geometric unit are the three vertices of the triangular unit; where the three vertices of the triangular unit are denoted as s1, s2, and s3 respectively, the coordinates of vertex s1 in the first coordinate system are ( , The position coordinates of vertex s2 in the first coordinate system are ( , The position coordinates of vertex s3 in the first coordinate system are ( , In the case of ), the linear transformation parameters A and translation parameters t that constitute the above affine transformation model can be obtained by solving according to the following formulas 4-6: =A× Formula 4; =A× Formula 5; =A× Formula 6; The position coordinates of vertex s1 in the second coordinate system are ( , ); The position coordinates of vertex s2 in the second coordinate system are ( , ); The position coordinates of vertex s3 in the second coordinate system are ( , ).

[0063] Specifically, when the control terminal determines the aforementioned local mapping relationship according to the method shown in step d1 above, the linear transformation parameter A and translation parameter t constituting the aforementioned affine transformation model can be obtained by solving the linear equation system composed of formulas 4-6 above; where, after knowing the aforementioned linear transformation parameter A and translation parameter t, for the key coordinate point s in the target geometric unit, the known position coordinates of the key coordinate point s in the first coordinate system are ( , In the case of ), the position coordinates of the key coordinate point s in the second coordinate system can be obtained by following the method shown in Formula 7 below. , ): =A× Formula 7.

[0064] Here, when performing step b4, as an optional embodiment, if the shape of the geometric unit is not limited in any way, the above-mentioned local mapping relationship corresponding to the geometric unit can be solved in the general manner shown in steps e1-e3 below: Step e1: Use the position coordinates of the marker points contained in the geometric unit in the grid calibration map as the input of the first target interpolation function, use the horizontal coordinates of the marker points contained in the geometric unit in the captured image as the output of the first target interpolation function, solve for the unknown parameters in the first target interpolation function, and obtain the solved first target interpolation function.

[0065] Specifically, the first target interpolation function can be the TPS (Thin Plate Spline) interpolation function; where, for each marker point contained in the geometric unit, the input of the first target interpolation function (equivalent to the independent variable on the right side of the first target interpolation function equation) is the position coordinate of the marker point in the first coordinate system (that is, the position coordinate in the mesh calibration diagram), and the output of the first target interpolation function (equivalent to the dependent variable on the left side of the first target interpolation function equation) is the abscissa of the marker point in the second coordinate system.

[0066] At this point, if the geometric unit contains n marked points, then n sets of first target interpolation functions can be obtained. By performing linear fitting on the n sets of first target interpolation functions, the specific parameter values ​​of the unknown parameters originally included in the right side of the first target interpolation function equation can be obtained. Substituting the specific parameter values ​​of the unknown parameters into the blank first target interpolation function, the first target interpolation function that no longer contains unknown parameters can be obtained as the solved first target interpolation function.

[0067] Step e2: Use the position coordinates of the marker points contained in the geometric unit in the grid calibration map as the input of the second target interpolation function, use the ordinate of the marker points contained in the geometric unit in the captured image as the output of the second target interpolation function, solve for the unknown parameters in the second target interpolation function, and obtain the solved second target interpolation function.

[0068] Specifically, the second objective interpolation function can also be a TPS interpolation function; where, for each marker point contained in the geometric unit, the input of the second objective interpolation function (equivalent to the independent variable on the right side of the second objective interpolation function equation) is the position coordinate of the marker point in the first coordinate system (that is, the position coordinate in the mesh calibration diagram), while the output of the second objective interpolation function (equivalent to the dependent variable on the left side of the second objective interpolation function equation) is the ordinate of the marker point in the second coordinate system.

[0069] At this point, in step e2, the specific method for solving the unknown parameters in the second target interpolation function is the same as the method for solving the unknown parameters in the first target interpolation function in step e1 above, and the repetition will not be repeated here.

[0070] Step e3: Use the solved first target interpolation function and the solved second target interpolation function as the local mapping relationship.

[0071] Here, combining the content of steps e1-e2 above, we know that the input of the solved first target interpolation function and the solved second target interpolation function only requires the position coordinates of a single point (i.e., a point on the display screen) in the first coordinate system. The outputs of the solved first target interpolation function and the solved second target interpolation function are the x-coordinate and y-coordinate of the point in the second coordinate system, respectively. Therefore, the solved first target interpolation function and the solved second target interpolation function can be used as the above-mentioned local mapping relationship corresponding to the geometric unit. This ensures that after knowing the position coordinates of any point in the first coordinate system within the geometric unit, the position coordinates of that point in the second coordinate system can be obtained by coordinate transformation based on this local mapping relationship.

[0072] Specifically, when the control terminal determines the aforementioned local mapping relationship according to the method shown in steps e1-e3 above, based on the content of step e1 above, it can be known that in the solved first target interpolation function, the independent variable is the position coordinate of the key coordinate point s (equivalent to a position point of the object to be identified in the sample image) in the first coordinate system, and the dependent variable in the solved first target interpolation function is the abscissa of the key coordinate point s in the second coordinate system. At this time, it is only necessary to substitute the position coordinate of the key coordinate point s in the first coordinate system into the solved first target interpolation function to obtain the abscissa of the object to be calibrated s in the second coordinate system.

[0073] Based on the content of step e2 above, the independent variable in the solved second target interpolation function is the position coordinate of the key coordinate point s in the first coordinate system, and the dependent variable in the solved second target interpolation function is the ordinate of the key coordinate point s in the second coordinate system. At this time, we only need to substitute the position coordinate of the key coordinate point s in the first coordinate system into the solved second target interpolation function to obtain the ordinate of the key coordinate point s in the second coordinate system.

[0074] S205, based on the second location information, automatically label the object to be identified contained in the target image to generate a labeled sample image corresponding to the sample image.

[0075] Here, as an optional embodiment, the control terminal can automatically label the objects to be identified in the target image according to the method shown in steps f1-f3 below, specifically: Step f1: Based on the second location information, determine the target location information corresponding to the outer rectangle of the object to be identified from the target image.

[0076] Here, taking the key coordinate points contained in the first position information as: the upper left corner A and the lower right corner B of the bounding rectangle where the object to be identified is located in the sample image as an example, based on the aforementioned step S204, the position coordinates a of corner point A in the second coordinate system can be determined as (x1', y1') and the position coordinates b of corner point B in the second coordinate system can be determined as (x2', y2') from the second position information.

[0077] At this point, based on position coordinates a and b, the x-coordinate of the center point of the circumscribed rectangle in the target image can be determined as (x1'+x2') / 2, and the y-coordinate as (y1'+y2') / 2; the width of the circumscribed rectangle is |x2'-x1'|, and the height is |y2'-y1'|. Thus, the coordinate information of the center point, the position coordinates of corner points A and B in the second coordinate system, and the width and height are used as the target position information.

[0078] Step f2: Based on the target location information and the type of the object to be identified, generate the annotation information corresponding to the object to be identified in the target image.

[0079] Here, based on the relevant description in step S201 above, it can be seen that during the process of generating the sample image, the control terminal can also synchronously record the type of the object to be identified contained in the sample image through the above-mentioned automated script on the display terminal (e.g., the label for the defect type "bright spot" is recorded as 0, the label for the defect type "dark spot" is recorded as 1, etc.). Based on this, when executing step f2, the control terminal can use the target location information determined in step f1 above and the type of the object to be identified previously recorded as the annotation information corresponding to the object to be identified in the target image (similar to the annotation required for image target detection: the image detection box where the target is located is consistent with the category to which the target belongs).

[0080] Step f3: Add the annotation information to the target image to generate the annotated sample image.

[0081] Here, by adding the above annotation information to the target image, an automatically annotated sample image can be obtained, which replaces the method of manually annotating the target image in the existing technology. This effectively solves the problems of high cost, low efficiency and insufficient accuracy of image acquisition and annotation in the existing technology.

[0082] It should be noted that, in order to facilitate the use of the generated labeled sample images as training data for the image detection model, the above labeled sample images can also be saved in a specific format that conforms to the image detection model (e.g., conforming to the YOLO object detection format).

[0083] Regarding the calibration method shown in steps b1-b4, it should be noted that since the interior corner points in the checkerboard calibration map do not have unique ID identifiers, for a regular checkerboard calibration map (i.e., a checkerboard calibration map without direction indicator labels), due to the rotation and mirror symmetry of the regular checkerboard calibration map, it is easy to make mistakes in the detection order of interior corner points when directly detecting them (e.g., mistakenly identifying the interior corner point in the lower right direction of the checkerboard calibration map as the starting mark point), which in turn causes errors in the subsequently established mapping model (i.e., the algorithm model used to represent the mapping relationship between the first coordinate system and the second coordinate system).

[0084] Based on this, to avoid the aforementioned error in the detection order of interior corner points, in this embodiment of the application, as an optional embodiment, after performing step b1, the detection order of interior corner points (which is also equivalent to the order of the marker points in the aforementioned marker point matrix) can be corrected according to the method shown in steps g1-g5 below, using the direction indicator carried in the chessboard calibration map, to ensure that the marker point represented by index (0,0) in the aforementioned marker point matrix is ​​always consistent with the starting marker point indicated by the direction indicator. Specifically: Step g1: According to the sorting method indicated by the direction indicator, detect the position coordinates of each inner corner point in the first coordinate system from the chessboard calibration map to obtain the set of inner corner point position coordinates, and rearrange each inner corner point into the target matrix according to the detection order.

[0085] Here, the kth detected interior corner point is denoted as c_k. The position coordinates of the interior corner point c_k in the first coordinate system (which is also equivalent to the position coordinates of the interior corner point c_k in the chessboard calibration diagram) can be expressed as (x_k, y_k); where x_k represents the abscissa of the interior corner point c_k in the first coordinate system, and y_k represents the ordinate of the interior corner point c_k in the first coordinate system.

[0086] Specifically, if the chessboard calibration map contains M×N interior corner points, then after rearranging all the detected interior corner points according to the detection order, an M-row N-column target matrix p(i,j) can be obtained; where the value range of i is [0, M-1] and the value range of j is [0, N-1]. The target matrix obtained at this time is also equivalent to the aforementioned marker point matrix.

[0087] Step g2: Take the interior corner points located on the edge vertex direction in the target matrix as target interior corner points, and obtain the position coordinates of the multiple target interior corner points in the first coordinate system from the set of interior corner point position coordinates.

[0088] Here, the interior corner points located at the edges of the target matrix (i.e., the target interior corner points) are also equivalent to the interior corner points located at the four corners of the target matrix. Taking the target matrix as p(i,j) as an example, the target interior corner points can be represented as: p00=P(0,0); p0N=P(0,N-1); pM0=P(M-1,0); pMN=P(M-1,N-1). In this case, p00 is the interior corner point located at the top left corner of the target matrix, p0N is the interior corner point located at the top right corner of the target matrix, pM0 is the interior corner point located at the bottom left corner of the target matrix, and pMN is the interior corner point located at the bottom right corner of the target matrix.

[0089] Step g3: For each target interior corner point, calculate the sum of the x-coordinate and y-coordinate of the target interior corner point in the first coordinate system as the candidate score corresponding to the target interior corner point.

[0090] Specifically, taking the target interior corners as p00, p0N, pM0, and pMN as examples, for each target interior corner, the position coordinates of each target interior corner in the first coordinate system can be obtained from the aforementioned set of interior corner position coordinates. Then, the candidate score for each target interior corner is calculated according to the method shown in Formula 8-11 below: s00 = x(p00) + y(p00) (Formula 8) s0N=x(p0N)+y(p0N) Formula 9; sM0=x(pM0)+y(pM0) Formula 10; sMN = x(pMN) + y(pMN) (Formula 11) Where s00 represents the candidate score corresponding to the inner corner point p00 of the target, x(p00) represents the abscissa of the inner corner point p00 of the target in the first coordinate system, and y(p00) represents the ordinate of the inner corner point p00 in the first coordinate system. s0N represents the candidate score corresponding to the inner corner point p0N of the target, x(p0N) represents the x-coordinate of the inner corner point p00 of the target in the first coordinate system, and y(p0N) represents the y-coordinate of the inner corner point p00 in the first coordinate system. sM0 represents the candidate score corresponding to the inner corner point pM0 of the target, x(pM0) represents the x-coordinate of the inner corner point pM0 of the target in the first coordinate system, and y(pM0) represents the y-coordinate of the inner corner point pM0 in the first coordinate system. sMN represents the candidate score corresponding to the inner corner point pMN of the target, x(pMN) represents the x-coordinate of the inner corner point pMN of the target in the first coordinate system, and y(pMN) represents the y-coordinate of the inner corner point pMN in the first coordinate system.

[0091] Step g4: From the multiple target interior corner points, determine the target interior corner point with the smallest candidate score as the candidate starting mark point.

[0092] Here, since the direction indicator is used to indicate the starting point located in the upper left corner of the chessboard calibration diagram, the specific values ​​of the horizontal and vertical coordinates of the starting point in the first coordinate system are relatively small. That is, in the chessboard calibration diagram, the sum of the horizontal and vertical coordinates of the starting point in the first coordinate system is the smallest.

[0093] Based on this, by calculating the candidate scores as described above, the target interior corner point with the smallest candidate score can be determined from the multiple target interior corner points as the candidate starting mark point. At this time, by comparing whether the candidate starting mark point is consistent with the starting mark point indicated by the direction indicator, it can be determined whether the detection order of the interior corner points has been incorrect.

[0094] Step g5: In response to the mismatch between the candidate starting marker and the starting marker indicated by the direction indicator, the target matrix is ​​rotated until a candidate starting marker that matches the starting marker is obtained.

[0095] Here, when the candidate starting marker point is inconsistent with the starting marker point indicated by the direction indicator (i.e., mismatch), it can be determined that the inner corner point detection order is incorrect. Therefore, the target matrix obtained by rearranging the inner corner point detection order needs to be rotated until a candidate starting marker point that matches the starting marker point is obtained (equivalent to ensuring that the marker point represented by p(0,0) in the target matrix is ​​consistent with the starting marker point indicated by the direction indicator).

[0096] It should be noted that when rotating the target matrix, the specific rotation method can be determined based on the specific candidate starting point that is inconsistent with the starting point at this time.

[0097] Specifically, if the candidate starting point is the target inner corner point p0N located in the upper right direction of the target matrix, the target matrix can be horizontally flipped; if the candidate starting point is the target inner corner point pM0 located in the lower left direction of the target matrix, the target matrix can be vertically flipped; if the candidate starting point is the target inner corner point pMN located in the lower right direction of the target matrix, all elements of the target matrix can be rotated 180° around the center point.

[0098] Based on the above-described automatic image acquisition and annotation method provided in the embodiments of this application, the acquisition end can automatically acquire sample images displayed on the display end, and according to the coordinate transformation relationship between the sample image and the automatically acquired target image, the first position information of the object to be identified in the sample image is transformed to obtain the second position information of the object to be identified in the target image. Based on the obtained second position information, the object to be identified contained in the target image is automatically annotated, which effectively solves the problems of high cost, low efficiency and insufficient accuracy of image acquisition and annotation in the prior art.

[0099] Based on the same inventive concept, this application also provides an automatic image acquisition and annotation system corresponding to the above-mentioned automatic image acquisition and annotation method. Since the principle of solving the problem by the automatic image acquisition and annotation system in the embodiments of this application is similar to that of the above-mentioned automatic image acquisition and annotation method in the embodiments of this application, the implementation of the automatic image acquisition and annotation system can refer to the implementation of the above-mentioned automatic image acquisition and annotation method, and the repeated parts will not be described again.

[0100] like Figure 1 As shown, the automatic image acquisition and annotation system includes: a display terminal, an acquisition terminal, and a control terminal; wherein, the control terminal is used for: The display terminal displays a sample image containing the object to be identified. The target image corresponding to the sample image is obtained by acquiring the sample image through the acquisition terminal; Based on the first position information of the object to be identified in the sample image, a target geometric unit matching the first position information is determined from a plurality of geometric units; wherein, the plurality of geometric units represent a plurality of geometric units that make up the display area of ​​the display terminal; Based on the local mapping relationship between the target geometric unit and the first coordinate system and the second coordinate system, the first position information is transformed to obtain the second position information of the object to be identified in the second coordinate system; wherein, the first coordinate system represents the coordinate system corresponding to the sample image, and the second coordinate system represents the coordinate system corresponding to the target image; Based on the second location information, the object to be identified contained in the target image is automatically labeled to generate a labeled sample image corresponding to the sample image.

[0101] In one optional implementation, when the sample image containing the object to be identified is displayed on the display terminal, the control terminal is used to: The sample image is generated by combining various image parameters in different ways using an automated script running on the display terminal; wherein, the various image parameters include: the type, quantity, and location of the object to be identified; The sample image is displayed on the display terminal.

[0102] In an optional implementation, the control terminal is used to determine the local mapping relationship between the first coordinate system and the second coordinate system for each geometric unit by the following method: The display terminal displays a grid calibration map with direction indicators; wherein, the grid calibration map includes multiple marker points, and the direction indicators are used to indicate the sorting order of the multiple marker points in the grid calibration map; The image of the grid calibration map is acquired by the acquisition terminal to obtain the captured image corresponding to the grid calibration map; Using the marked points as vertices of geometric units, the display area on the display terminal is divided into the plurality of geometric units; For each geometric unit, the local mapping relationship between the geometric unit and the second coordinate system is determined based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, respectively.

[0103] In an optional implementation, when the grid calibration map belongs to a checkerboard calibration map, the plurality of marker points belong to the interior corner points of the checkerboard calibration map, and the direction indicator is used to indicate the starting marker point located in the upper left corner direction of the checkerboard calibration map; wherein, after the grid calibration map with the direction indicator is displayed through the display terminal, the control terminal is further used to: According to the sorting method indicated by the direction indicator, the position coordinates of each inner corner point in the first coordinate system are detected sequentially from the chessboard calibration map to obtain the set of inner corner point position coordinates, and each inner corner point is rearranged into a target matrix according to the detection order; The interior corner points located at the edge vertices of the target matrix are taken as target interior corner points, and the position coordinates of the multiple target interior corner points in the first coordinate system are obtained from the set of interior corner point position coordinates. For each target interior corner point, the sum of the x-coordinate and y-coordinate corresponding to the target interior corner point in the first coordinate system is calculated as a candidate score for the target interior corner point. From the plurality of target interior corner points, the target interior corner point with the smallest candidate score is determined as the candidate starting mark point; In response to a mismatch between the candidate starting marker and the starting marker indicated by the direction indicator, the target matrix is ​​rotated until a candidate starting marker that matches the starting marker is obtained.

[0104] In an optional implementation, when determining the local mapping relationship between the geometric unit and the second coordinate system based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, the control terminal is used to: Based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, the perspective transformation matrix corresponding to the geometric unit is obtained as the local mapping relationship; or, Based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, the affine transformation model corresponding to the geometric unit is obtained as the local mapping relationship.

[0105] In an optional implementation, the first location information includes multiple key coordinate points of the object to be identified in the sample image, wherein, when determining a target geometric unit matching the first location information from multiple geometric units based on the first location information of the object to be identified in the sample image, the control terminal is used to: For each key coordinate point, from the plurality of geometric units, determine the geometric unit that matches the key coordinate point and belongs to the target geometric unit.

[0106] In an optional implementation, when automatically labeling the object to be identified contained in the target image based on the second location information to generate a labeled sample image corresponding to the sample image, the control terminal is used to: Based on the second location information, the target location information corresponding to the bounding rectangle of the object to be identified is determined from the target image; Based on the target location information and the type of the object to be identified, generate annotation information corresponding to the object to be identified in the target image; The annotation information is added to the target image to generate the annotated sample image.

[0107] like Figure 4 As shown, this application provides an electronic device 400 for executing the automatic image acquisition and annotation method of this application. The device includes a memory 401, a processor 402, and a computer program stored in the memory 401 and executable on the processor 402. The memory 401 and the processor 402 are connected via a bus for communication. When the processor 402 executes the computer program, it implements the steps of the automatic image acquisition and annotation method described above.

[0108] Specifically, the memory 401 and processor 402 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 402 runs the computer program stored in the memory 401, it can execute the above-mentioned automatic image acquisition and annotation method.

[0109] Corresponding to the automatic image acquisition and annotation method in this application, this application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the steps of the above-described automatic image acquisition and annotation method.

[0110] Specifically, the storage medium can be a general-purpose storage medium, such as a portable disk or hard disk. When the computer program on the storage medium is run, it can execute the above-mentioned automatic image acquisition and annotation method.

[0111] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0112] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0113] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0114] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0115] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0116] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. An automatic image acquisition and annotation method, characterized in that, The automatic image acquisition and annotation method includes: The display terminal shows a sample image containing the object to be identified; The target image corresponding to the sample image is obtained by acquiring the sample image through the acquisition terminal; Based on the first position information of the object to be identified in the sample image, a target geometric unit matching the first position information is determined from a plurality of geometric units; wherein, the plurality of geometric units represent a plurality of geometric units that make up the display area of ​​the display terminal; Based on the local mapping relationship between the target geometric unit and the first coordinate system and the second coordinate system, the first position information is transformed to obtain the second position information of the object to be identified in the second coordinate system; wherein, the first coordinate system represents the coordinate system corresponding to the sample image, and the second coordinate system represents the coordinate system corresponding to the target image; Based on the second location information, the object to be identified contained in the target image is automatically labeled to generate a labeled sample image corresponding to the sample image.

2. The automatic image acquisition and annotation method according to claim 1, characterized in that, The step of displaying a sample image containing the object to be identified via a display terminal includes: The sample image is generated by combining various image parameters in different ways using an automated script running on the display terminal; wherein, the various image parameters include: the type, quantity, and location of the object to be identified; The sample image is displayed on the display terminal.

3. The automatic image acquisition and annotation method according to claim 1, characterized in that, The local mapping relationship between the first coordinate system and the second coordinate system for each geometric element is determined by the following method: The display terminal displays a grid calibration map with direction indicators; wherein, the grid calibration map includes multiple marker points, and the direction indicators are used to indicate the sorting order of the multiple marker points in the grid calibration map; The image of the grid calibration map is acquired by the acquisition terminal to obtain the captured image corresponding to the grid calibration map; Using the marked points as vertices of geometric units, the display area on the display terminal is divided into the plurality of geometric units; For each geometric unit, the local mapping relationship between the geometric unit and the second coordinate system is determined based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, respectively.

4. The automatic image acquisition and annotation method according to claim 3, characterized in that, When the grid calibration map belongs to a checkerboard calibration map, the plurality of marker points belong to the interior corner points of the checkerboard calibration map, and the direction indicator is used to indicate the starting marker point located in the upper left corner of the checkerboard calibration map; wherein, after the grid calibration map with the direction indicator is displayed on the display terminal, the automatic image acquisition and annotation method further includes: According to the sorting method indicated by the direction indicator, the position coordinates of each inner corner point in the first coordinate system are detected sequentially from the chessboard calibration map to obtain the set of inner corner point position coordinates, and each inner corner point is rearranged into a target matrix according to the detection order; The interior corner points located at the edge vertices of the target matrix are taken as target interior corner points, and the position coordinates of the multiple target interior corner points in the first coordinate system are obtained from the set of interior corner point position coordinates. For each target interior corner point, the sum of the x-coordinate and y-coordinate corresponding to the target interior corner point in the first coordinate system is calculated as a candidate score for the target interior corner point. From the plurality of target interior corner points, the target interior corner point with the smallest candidate score is determined as the candidate starting mark point; In response to a mismatch between the candidate starting marker and the starting marker indicated by the direction indicator, the target matrix is ​​rotated until a candidate starting marker that matches the starting marker is obtained.

5. The automatic image acquisition and annotation method according to claim 3, characterized in that, The step of determining the local mapping relationship between the geometric unit and the second coordinate system based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image includes: Based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, the perspective transformation matrix corresponding to the geometric unit is obtained as the local mapping relationship; or, Based on the position coordinates of the marker points contained in the geometric unit in the grid calibration map and the captured image, the affine transformation model corresponding to the geometric unit is obtained as the local mapping relationship.

6. The automatic image acquisition and annotation method according to claim 1, characterized in that, The first location information includes multiple key coordinate points of the object to be identified in the sample image, wherein determining the target geometric unit matching the first location information from multiple geometric units based on the first location information of the object to be identified in the sample image includes: For each key coordinate point, from the plurality of geometric units, determine the geometric unit that matches the key coordinate point and belongs to the target geometric unit.

7. The automatic image acquisition and annotation method according to claim 1, characterized in that, The step of automatically labeling the object to be identified contained in the target image based on the second location information, and generating a labeled sample image corresponding to the sample image, includes: Based on the second location information, the target location information corresponding to the bounding rectangle of the object to be identified is determined from the target image; Based on the target location information and the type of the object to be identified, generate annotation information corresponding to the object to be identified in the target image; The annotation information is added to the target image to generate the annotated sample image.

8. An automatic image acquisition and annotation system, characterized in that, The automatic image acquisition and annotation system includes: a display terminal, an acquisition terminal, and a control terminal; wherein, the control terminal is used for: The display terminal displays a sample image containing the object to be identified. The target image corresponding to the sample image is obtained by acquiring the sample image through the acquisition terminal; Based on the first position information of the object to be identified in the sample image, a target geometric unit matching the first position information is determined from a plurality of geometric units; wherein, the plurality of geometric units represent a plurality of geometric units that make up the display area of ​​the display terminal; Based on the local mapping relationship between the target geometric unit and the first coordinate system and the second coordinate system, the first position information is transformed to obtain the second position information of the object to be identified in the second coordinate system; wherein, the first coordinate system represents the coordinate system corresponding to the sample image, and the second coordinate system represents the coordinate system corresponding to the target image; Based on the second location information, the object to be identified contained in the target image is automatically labeled to generate a labeled sample image corresponding to the sample image.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the automatic image acquisition and annotation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the automatic image acquisition and annotation method as described in any one of claims 1 to 7.