Multi-modal data collaborative data annotation method, system and equipment
By generating 3D point cloud models and optical reflectance maps, and collaboratively analyzing the microscopic morphology and optical properties of product surfaces, the problem of inaccurate product image labeling in existing technologies is solved, enabling accurate identification and comprehensive labeling of material and process details.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies cannot accurately identify the microscopic physical details of product materials and processes in the data annotation of product images, resulting in inaccurate and incomplete annotation.
By acquiring color images, surface attribute text data, and deformable grating image sequences, a 3D point cloud model is generated and its micromorphology is quantified. Combined with optical reflectance maps, multimodal data collaborative annotation is performed.
It achieves accurate recognition and comprehensive annotation of product images, and can identify specific materials and process details, thus improving the accuracy and comprehensiveness of data annotation.
Smart Images

Figure CN121725480A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data annotation technology, and in particular relates to a data annotation method, system and device for multimodal data collaboration. Background Technology
[0002] Data labeling, a core component of artificial intelligence, involves processing raw data and adding labels to provide training samples for machine learning models. With the rapid development of computer vision and natural language processing technologies, high-quality labeled data plays a crucial role in numerous scenarios such as autonomous driving, intelligent manufacturing, and e-commerce, and its application prospects are vast.
[0003] In the field of e-commerce data annotation, existing multimodal data collaborative annotation methods typically combine product images and user review text to automatically annotate product attributes. These methods usually utilize image recognition technology to analyze visual features in images, while employing natural language processing to parse descriptive words in the text. A fusion model then links the two to achieve a preliminary judgment and annotation of product features.
[0004] However, existing technologies have significant shortcomings when it comes to detailed material and detail annotation of product images. Because their analysis is limited to a two-dimensional visual level, they cannot identify details such as product materials and manufacturing processes that depend on microscopic physical structures, resulting in inaccurate and incomplete data annotation of product images. Therefore, existing technologies suffer from insufficient accuracy and comprehensiveness in product image data annotation. Summary of the Invention
[0005] The purpose of this application is to provide a multimodal data collaborative data annotation method, system, electronic device, and storage medium to solve the problem of insufficient accuracy and comprehensiveness of data annotation for product images in the prior art.
[0006] To address the aforementioned technical problems, in a first aspect, this application provides a multimodal data collaborative data annotation method, comprising:
[0007] The system acquires color images of the target product, surface attribute text data, and a sequence of deformable grating images. The surface attribute text data includes text describing the material and process of the target product. The sequence of deformable grating images is obtained by projecting an coded grating pattern onto the physical surface of the target product and capturing deformable grating images modulated by the contours of the physical surface from different angles.
[0008] Semantic analysis is performed on surface attribute text data to extract keywords corresponding to materials and processes, resulting in a surface feature vocabulary set;
[0009] By decoding the encoded information of the deformed grating image in the deformed grating image sequence, the three-dimensional spatial coordinates of multiple points on the surface of the target product are calculated, and the three-dimensional spatial coordinates are aggregated to generate a three-dimensional point cloud model representing the surface contour of the target product.
[0010] By calculating the geometric feature parameters of local regions in a 3D point cloud model, the microscopic morphology of the target product surface is quantified, and a set of geometric feature descriptors is generated.
[0011] By spatially registering the 3D point cloud model with the color image, a spatial position mapping relationship between the 3D point cloud model and the color image is established;
[0012] Based on spatial location mapping, geometric feature descriptors are associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image.
[0013] In one feasible implementation, the method further includes:
[0014] Based on the average brightness of the grating stripes at the same pixel location in multiple deformed grating images in a deformed grating image sequence, the reflection intensity value representing the surface reflectivity of each pixel location on the physical surface of the target product is calculated, and a surface reflectivity map is generated.
[0015] The reflection intensity values in the surface reflectivity map are associated with the corresponding spatial coordinate points in the 3D point cloud model.
[0016] Based on spatial location mapping, geometric feature descriptors are associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image, including:
[0017] Based on spatial location mapping, the reflection intensity values corresponding to spatial coordinate points in the geometric feature descriptor and the 3D point cloud model are associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and to generate annotation information including attribute labels on the target image region in the color image.
[0018] In one feasible implementation, based on spatial location mapping, the reflection intensity values corresponding to spatial coordinate points in the 3D point cloud model are associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image, including:
[0019] Based on the spatial location mapping relationship, the corresponding set of three-dimensional spatial coordinate points is determined for the target image region in the color image;
[0020] Extract the geometric feature descriptors and reflection intensity values corresponding to the set of three-dimensional spatial coordinate points, and combine the geometric feature descriptors and reflection intensity values to generate a combined feature vector;
[0021] Calculate the similarity between the combined feature vector and the preset feature vector corresponding to each keyword in the surface feature vocabulary set;
[0022] Select at least one keyword with a similarity exceeding a preset similarity threshold as the attribute label corresponding to the target image region, and generate annotation information including the attribute label in the target image region.
[0023] In one feasible implementation, semantic analysis is performed on the surface attribute text data to extract keywords corresponding to the material and process, resulting in a surface feature vocabulary set, including:
[0024] Syntactic structure analysis is performed on surface attribute text data to identify and extract noun phrases and adjective phrases describing the physical attributes of goods, resulting in a candidate phrase set;
[0025] Each phrase in the candidate phrase set is matched with a preset material terminology library and a preset process terminology library, and the successfully matched words are classified to generate a material keyword subset and a process keyword subset.
[0026] By combining subsets of material keywords and subsets of process keywords, a set of surface feature terms is generated.
[0027] In one feasible implementation, by decoding the encoded information of the deformed grating image in the deformed grating image sequence, the three-dimensional spatial coordinates of multiple points on the surface of the target product are calculated, and the three-dimensional spatial coordinates are aggregated to generate a three-dimensional point cloud model representing the contour of the target product surface, including:
[0028] By decoding the binary coded pattern used to determine the fringe period in the deformed grating image sequence, a unique fringe period number is determined for each pixel position.
[0029] Based on the brightness change at the same pixel position presented by the sinusoidal fringe pattern used for localization in the deformed grating image sequence, a wrapping phase value that cycles within a single fringe period is calculated.
[0030] The fringe period number is used to determine the integer multiple of the phase period to which the wrapped phase value belongs, and the wrapped phase value is superimposed on the integer multiple of the phase period to calculate an absolute phase value without periodic jumps.
[0031] Based on the absolute phase value and the system parameters of the geometric relationship between the projection of the representation-encoded grating pattern and the capture of the deformed grating image, the three-dimensional spatial coordinates are calculated by performing triangulation calculation on the absolute phase value and its corresponding pixel position.
[0032] A three-dimensional point cloud model is constructed based on all three-dimensional spatial coordinates in a unified coordinate system.
[0033] In one feasible implementation, the geometric feature parameters include local curvature and the rate of change of the normal vector;
[0034] By calculating the geometric feature parameters of local regions in a 3D point cloud model, the microscopic morphology of the target product surface is quantified, generating a set of geometric feature descriptors, including:
[0035] Using each three-dimensional spatial coordinate point in the three-dimensional point cloud model as the center coordinate point, select neighboring three-dimensional spatial coordinate points to form a local region point set corresponding to each three-dimensional spatial coordinate point;
[0036] For each local region point set, an approximate surface of the local region point set is obtained through surface fitting, and based on the approximate surface, the local curvature of the center coordinate point is calculated through differential geometry.
[0037] Based on the approximate surface, the surface normal vectors of all three-dimensional spatial coordinate points in the local region point set are analyzed, and the rate of change of the normal vector of the center coordinate point is obtained by calculating the angular dispersion of the surface normal vector of each three-dimensional spatial coordinate point in the local region point set.
[0038] By combining the local curvature value and the rate of change of the normal vector, the geometric feature descriptor of the center coordinate point is obtained, and a set of geometric feature descriptors is formed based on the geometric feature descriptors of all three-dimensional space coordinate points.
[0039] In one feasible implementation, a spatial position mapping relationship between the 3D point cloud model and the color image is established by spatially registering the 3D point cloud model with the color image, including:
[0040] Using a pre-defined camera parameter model that represents the transformation relationship between the coordinate system of a color image and a 3D point cloud model, each 3D spatial coordinate point in the 3D point cloud model is projected onto the 2D image coordinate system of the color image to obtain the 2D projected coordinates corresponding to each 3D spatial coordinate point.
[0041] For each pixel in a color image, find at least one three-dimensional spatial coordinate point within the neighborhood of the pixel's corresponding two-dimensional projected coordinates.
[0042] Each pixel in the color image is associated with at least one found 3D spatial coordinate point to establish a spatial location mapping relationship.
[0043] Secondly, this application provides a multimodal data collaborative data annotation system, comprising:
[0044] The acquisition module is used to acquire color images, surface attribute text data, and deformable grating image sequences of the target product. The surface attribute text data includes text describing the material and process of the target product. The deformable grating image sequence is obtained by projecting coded grating patterns onto the physical surface of the target product and capturing deformable grating images modulated by the contours of the physical surface from different angles.
[0045] The extraction module is used to perform semantic analysis on surface attribute text data to extract keywords corresponding to materials and processes, thereby obtaining a set of surface feature vocabularies.
[0046] The generation module is used to calculate the three-dimensional spatial coordinates of multiple points on the surface of the target product by decoding the encoded information of the deformed grating image in the deformed grating image sequence, and to aggregate the three-dimensional spatial coordinates to generate a three-dimensional point cloud model representing the contour of the target product surface.
[0047] The generation module is also used to quantify the micro-morphology of the target product surface by calculating the geometric feature parameters of local regions in the 3D point cloud model, and generate a set of geometric feature descriptors.
[0048] The module is used to establish a spatial mapping relationship between a 3D point cloud model and a color image by spatially registering the 3D point cloud model with the color image.
[0049] The matching module is used to associate and match geometric feature descriptors with keywords in the surface feature vocabulary set based on spatial location mapping relationships, so as to identify the material and process corresponding to the target image region in the color image, and generate annotation information including attribute labels on the target image region in the color image.
[0050] Thirdly, this application provides an electronic device, comprising:
[0051] Memory, used to store computer programs;
[0052] A processor is used to execute computer programs to implement the steps of the multimodal data collaborative data annotation method as described in the first aspect above.
[0053] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the multimodal data collaborative data annotation method described in the first aspect above.
[0054] The multimodal data collaborative data annotation method provided in this application introduces deformable grating image sequences and decodes them to generate a 3D point cloud model representing the surface contour of a product, thereby quantifying the microscopic morphology of the product surface. This solves the problem of existing technologies that rely solely on 2D color images and cannot perceive physical structural details. By associating the quantified microscopic morphology features with text keywords, this application enables the annotation process to go beyond macroscopic visual appearance and delve into details such as materials and processes determined by physical structure. Compared with existing technologies, this application can identify and annotate specific material and process details, improving the accuracy and comprehensiveness of product image data annotation.
[0055] Furthermore, by analyzing the brightness information of the deformable grating image sequence, a surface reflectance map that can quantitatively characterize the optical reflective properties of the product surface is generated. This solves the problem that relying solely on three-dimensional geometry cannot distinguish materials or processes with similar microstructures but different optical properties, such as distinguishing between matte and glossy surfaces. Therefore, this application, by introducing the physical dimension of surface reflectance and coordinating it with geometric features for correlation matching, further enhances the ability to identify product material and process details, making the data annotation results more accurate and detailed. Attached Figure Description
[0056] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 A flowchart illustrating a multimodal data collaborative data annotation method provided in this application embodiment;
[0058] Figure 2 A schematic diagram illustrating a specific implementation of a method for determining annotation information of a color image, provided in an embodiment of this application;
[0059] Figure 3 This is a schematic diagram illustrating a specific implementation of a method for generating geometric feature descriptors provided in an embodiment of this application;
[0060] Figure 4 This application provides a schematic diagram of the structure of a multimodal data collaborative data annotation system.
[0061] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0062] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0063] To address the problems of existing technologies, embodiments of this application provide a data annotation method, apparatus, device, computer storage medium, and computer program product for multimodal data collaboration. The data annotation method for multimodal data collaboration provided in this application embodiment will be described first below.
[0064] Figure 1 This illustration shows a flowchart of a multimodal data collaborative data annotation method provided in one embodiment of this application. Figure 1 As shown, the method includes:
[0065] S110: Acquire a color image of the target product, surface attribute text data, and a sequence of deformed grating images. The surface attribute text data includes text describing the material and process of the target product. The sequence of deformed grating images is obtained by projecting an coded grating pattern onto the physical surface of the target product and capturing deformed grating images modulated by the contours of the physical surface from different angles.
[0066] Surface attribute text data refers to textual information associated with a target product that describes its various potential physical attributes, but the textual information does not specify the correspondence between each attribute and a specific part of the product. A morphing grating image sequence refers to a set of digital images captured by an image acquisition device, recording the shape of an coded grating pattern modulated by the surface contour of the target product. This set of images typically contains stripe patterns of various frequencies or encoding methods.
[0067] For a target product to be initially labeled, three sets of heterogeneous source data are first acquired. These three sets of data include high-resolution color images for standard product display, surface attribute text data for semantic analysis, and a sequence of deformable grating images that record the modulation information of the product surface to the coded grating pattern. For example, a piece of surface attribute text data could be "This product uses sandblasting, anodizing, and powder coating processes." The acquisition of the deformable grating image sequence is achieved by placing the physical target product in a pre-built and calibrated 3D scanning environment, which includes a digital grating projection device and one or more industrial cameras. During the acquisition process, the digital grating projection device projects a series of coded grating patterns onto the surface of the product in a preset sequence. The industrial cameras synchronize with the projection device, capturing a deformable grating image modulated by the product surface contour for each pattern projection, ultimately forming this deformable grating image sequence.
[0068] For example, taking a high-end computer mouse as an example, the data to be labeled for this mouse includes a JPEG color image showing its gray metallic shell, a general text data on surface properties stating "high-end computer mouse, using sandblasting, anodizing, and powder coating processes," and a morphing image sequence. This morphing image sequence is obtained by fixing the actual mouse on a 3D scanning stage and projecting a set of coded grating patterns onto its metal shell using a digital grating projection device. The main body of the shell has a fine sandblasting treatment, the left and right button areas have a matte anodizing treatment, and the decorative strip in the middle has a thin powder coating treatment. These three processes are all presented as a matte gray that is almost indistinguishable to the naked eye in the color image. An industrial camera simultaneously captures multiple morphing images. These images record stripe patterns that are similar in macroscopic vision but have significant morphological differences in microscopic scale, modulated by the sandblasted surface of the mouse body, the anodized surface of the button area, and the powder-coated surface of the central decorative strip, respectively, forming the morphing image sequence of this mouse.
[0069] S120: Perform semantic analysis on the surface attribute text data to extract keywords corresponding to the material and process, and obtain a set of surface feature vocabularies.
[0070] A surface feature vocabulary set refers to a structured list of keywords extracted from surface attribute text data that are directly related to the material or process of a product.
[0071] First, syntactic structure analysis is performed on the input surface attribute text data. This analysis identifies noun phrases and adjective phrases in the text, as these phrases typically contain core information describing physical properties. Next, the identified phrases are matched against a pre-defined knowledge base containing a large number of material and process-related word roots. This matching process determines whether each phrase belongs to the material or process category and categorizes the successfully matched words. Finally, all categorized material keywords and process keywords are combined to form a structured set of surface feature vocabularies.
[0072] S130: By decoding the encoded information of the deformed grating image in the deformed grating image sequence, the three-dimensional spatial coordinates of multiple points on the surface of the target product are calculated, and the three-dimensional spatial coordinates are aggregated to generate a three-dimensional point cloud model representing the surface contour of the target product.
[0073] Encoded information refers to the phase or spatial encoded data carried by the projection pattern embedded in the deformable grating image, which changes due to the modulation of the product surface contour. A 3D point cloud model is a set of a large number of 3D spatial coordinate points that can accurately represent the physical contour of the target product surface in a unified coordinate system.
[0074] First, the encoded information in the deformable grating image sequence is decoded. This process typically involves combining multiple patterns. For example, decoding the binary encoded pattern roughly determines the fringe period of each pixel, while analyzing the brightness changes of the sinusoidal fringe pattern precisely calculates its phase value within a single period. Finally, these two methods are combined to obtain a unique, continuous absolute phase value. Next, based on the calculated absolute phase value and pre-calibrated system geometric parameters, triangulation is used to calculate the coordinates of each pixel in three-dimensional space. Finally, all calculated three-dimensional coordinates are aggregated to construct a complete three-dimensional point cloud model.
[0075] S140: By calculating the geometric feature parameters of local regions in the 3D point cloud model, the microscopic morphology of the target product surface is quantified, and a set of geometric feature descriptors is generated.
[0076] Geometric feature parameters refer to a series of statistical values used to describe the local geometric properties of a point cloud, which may include local curvature, rate of change of normal vector, etc. A geometric feature descriptor is a vector or data set composed of multiple geometric feature parameters, which assigns a physical attribute signature to each point in the point cloud that describes whether the region it is in is smooth, rough, regularly undulating, or anisotropic.
[0077] First, the system iterates through every 3D spatial coordinate point in the 3D point cloud model. Centering each point, a local region set consisting of its neighboring points is determined by setting a radius or a number of neighboring points. Next, this local region set is analyzed, for example through surface fitting or covariance analysis, to calculate a series of geometric feature parameters. This calculation process yields a local curvature value quantifying the curvature of the region around the center point, and a normal vector change rate value quantifying the stability of the normal direction. Finally, the calculated local curvature, normal vector change rate, and other geometric feature parameters are combined to form a geometric feature descriptor corresponding one-to-one with the center point. Ultimately, this generates individual geometric feature descriptors for all points in the entire 3D point cloud model.
[0078] S150: By spatially registering the 3D point cloud model with the color image, a spatial position mapping relationship between the 3D point cloud model and the color image is established.
[0079] Spatial location mapping refers to a data structure that describes the correspondence between three-dimensional spatial coordinates in a three-dimensional point cloud model and two-dimensional pixels in a color image. Through this mapping, the position of any pixel in three-dimensional space or the projection position of any three-dimensional point on a two-dimensional image can be queried.
[0080] First, a pre-calibrated camera parameter model is used. This model includes the camera's intrinsic and extrinsic parameters, as well as distortion coefficients, and accurately describes the projection transformation relationship from 3D spatial points to 2D image pixels. Using this camera parameter model, each 3D spatial coordinate point in the 3D point cloud model is projected one by one into the 2D image coordinate system of the color image, thus calculating a corresponding 2D projected coordinate for each 3D point. Next, to address potential holes or occlusions after projection, a search is performed within the neighborhood of each pixel's corresponding 2D projected coordinate to find one or more nearest 3D spatial coordinate points. Finally, each pixel in the color image is associated with its found 3D spatial coordinate points, establishing a complete spatial mapping relationship covering the entire image.
[0081] S160: Based on spatial location mapping, geometric feature descriptors are associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image.
[0082] Labeling information refers to a combination of visual and textual information generated on the original color image to indicate specific material or process areas and their attributes. It may include bounding boxes or pixel-level segmentation masks for precisely delineating feature areas, and attribute labels attached to these areas describing their specific properties.
[0083] First, based on the established spatial mapping relationship, a set of 3D spatial coordinate points corresponding to a target image region in the color image is determined in the 3D point cloud model. Next, all geometric feature descriptors corresponding to this set of 3D spatial coordinate points are extracted, and these descriptors are statistically or aggregated, for example, by calculating the average value of each geometric feature parameter in the descriptor, to generate a comprehensive feature vector that can represent the macroscopic physical properties of the entire target image region. Then, this comprehensive feature vector is compared with a pre-defined feature library, which stores a corresponding pre-defined feature vector for each keyword in the surface feature vocabulary set. By calculating the similarity between the comprehensive feature vector and each pre-defined feature vector in the feature library, the keyword with the highest similarity is selected as the most likely material or process attribute of the target image region. Finally, based on the identified keyword, a bounding box or segmentation mask is generated around the target image region on the original color image, and the keyword is attached as an attribute label to form the final annotation information.
[0084] For example, firstly, the gray area of the mouse's main body is selected as the target image region on the color image. Based on the established spatial mapping relationship, tens of thousands of three-dimensional spatial coordinate points corresponding to this region are determined. Geometric feature descriptors for each of these points are extracted, and the average value of each geometric feature parameter in all descriptors is calculated to obtain a comprehensive feature vector representing the physical properties of the mouse's main body region. For example, this vector might show a low local curvature value and a moderate rate of change of the normal vector. Simultaneously, a preset feature library stores preset feature vectors corresponding to the keywords sandblasting, anodizing, and powder coating. By calculating the Euclidean distance between the comprehensive feature vector of the mouse's main body region and these three preset feature vectors, it is found that the distance to the preset feature vector corresponding to sandblasting is the smallest. Therefore, the system automatically generates an attribute label with the content "Process: Sandblasting" on the mouse's main body region selected on the color image, forming a complete annotation information.
[0085] In one feasible implementation, the method further includes:
[0086] Based on the average brightness of the grating stripes at the same pixel location in multiple deformed grating image sequences, the reflection intensity value representing the surface reflectivity of each pixel location on the actual surface of the target product is calculated, and a surface reflectivity map is generated.
[0087] A surface reflectance map is a grayscale image of the same size as the original image. The grayscale value of each pixel represents the reflectivity of a corresponding point on the surface of the product to a specific wavelength of light. The reflectance intensity value is used to quantify this reflectivity; a higher value indicates stronger reflectivity.
[0088] Multiple phase-shifted sinusoidal fringe patterns are extracted from a deformed grating image sequence. For each pixel in the image, its brightness value sequence under these patterns is obtained, and the arithmetic mean of this sequence is calculated. This average value physically corresponds to the background light intensity at that point, unaffected by changes in fringe brightness, and can well approximate the surface reflectivity of that point, thus being used as the reflectivity value of that pixel. The reflectivity values of all pixels are normalized and mapped to grayscale values to generate a complete surface reflectivity map. Taking a high-end computer mouse as an example, for four phase-shifted sinusoidal fringe patterns in its deformed grating image sequence, the brightness value sequence of a certain pixel in the image might be 210, 150, 90, 150, and the average brightness value obtained is a reflectivity value of 150.
[0089] The reflection intensity values in the surface reflectivity map are correlated with the corresponding spatial coordinate points in the 3D point cloud model.
[0090] Since both the surface reflectance map and the 3D point cloud model originate from the same set of deformable grating image sequences, each pixel in the image has a one-to-one correspondence with a 3D spatial coordinate point in the 3D point cloud model. By iterating through each pixel in the surface reflectance map and storing its corresponding reflection intensity value as a new attribute of its corresponding 3D spatial coordinate point, the 3D point cloud model gains not only 3D coordinates and geometric feature descriptors but also the optical attribute of reflection intensity. Continuing with the example of the pixel on a computer mouse, its corresponding 3D spatial coordinate point (10.5, 30.2, 55.8) now has a reflection intensity value of 150 added to its original geometric feature descriptor.
[0091] Step S160, based on spatial location mapping, associates and matches geometric feature descriptors with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and generates annotation information including attribute labels on the target image region in the color image, including:
[0092] Based on spatial location mapping, the reflection intensity values corresponding to spatial coordinate points in the geometric feature descriptor and the 3D point cloud model are associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and to generate annotation information including attribute labels on the target image region in the color image.
[0093] First, based on the established spatial mapping relationship, a set of three-dimensional spatial coordinate points is determined for a target image region in the color image. Next, the geometric feature descriptors and reflection intensity values corresponding to all points in this set are extracted, and these two physical attributes are combined to form a combined feature vector that represents the comprehensive physical characteristics of the region. Then, this combined feature vector is compared with a pre-defined feature library, which stores a pre-defined feature vector with matching dimension and structure for each keyword in the surface feature vocabulary. By calculating the similarity between the combined feature vector and each pre-defined feature vector, the keyword with the highest similarity is selected as the most likely attribute of the target image region. Finally, attribute labels are generated based on the identified keywords, and these labels, along with bounding boxes or segmentation masks, constitute the final annotation information.
[0094] Figure 2 A flowchart illustrating a method for determining annotation information of a color image according to an embodiment of this application is shown.
[0095] In one feasible implementation, based on spatial location mapping, the reflection intensity values corresponding to spatial coordinate points in the 3D point cloud model are associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image, such as... Figure 2 As shown, the method includes:
[0096] S210: Based on the spatial location mapping relationship, determine the corresponding set of three-dimensional spatial coordinate points for the target image region in the color image.
[0097] First, a target image region is defined on the color image. This region can be manually selected by the user or automatically determined by an image segmentation algorithm. Next, every pixel within this target image region is traversed. For each pixel, one or more associated 3D spatial coordinate points are found by querying the established spatial location mapping relationship. Finally, the 3D spatial coordinate points corresponding to all pixels within the target image region are aggregated to form a complete set of 3D spatial coordinate points. Taking a high-end computer mouse as an example, the gray area of its left and right buttons is selected on the color image as the target image region. By querying the spatial location mapping relationship, it can be determined that this region corresponds to tens of thousands of 3D spatial coordinate points in a 3D point cloud model.
[0098] S220: Extract the geometric feature descriptors and reflection intensity values corresponding to the set of three-dimensional spatial coordinate points, and combine the geometric feature descriptors and reflection intensity values to generate a combined feature vector.
[0099] A combined feature vector is a high-dimensional vector composed of multiple numerical values that can simultaneously describe the macroscopic geometric shape and macroscopic optical reflection characteristics of a region. The core of this step lies in fusing the two different physical attributes extracted from the 3D point cloud model—geometric shape and optical reflection—to form a more discriminative and unified feature representation.
[0100] First, for the set of 3D spatial coordinate points determined in the previous step, extract the geometric feature descriptor and reflection intensity value associated with each point. Next, statistically aggregate all geometric feature descriptors within this set, for example, by calculating the average of each geometric feature parameter, to obtain a geometric mean vector representing the macroscopic geometry of the region. Similarly, statistically aggregate all reflection intensity values, for example, by calculating their average, to obtain a mean reflectance value representing the macroscopic optical reflection characteristics of the region. Finally, concatenate this geometric mean vector and this mean reflectance value to form a combined feature vector.
[0101] A geometric feature descriptor can be defined as a two-dimensional vector containing the average local curvature and the rate of change of the average normal vector, while the reflection intensity value is a scalar. The resulting combined feature vector... It can be represented as a three-dimensional vector: .in, The average local curvature of the target image region. This represents the rate of change of the average normal vector. This represents the average reflection intensity value. Continuing with the example of a computer mouse button area, we extract the geometric feature descriptors and reflection intensity values of all points in the corresponding three-dimensional spatial coordinate set. By calculating the average value, we can obtain that the average local curvature of this area is 0.01, the average normal vector change rate is 0.08, and the average reflection intensity value is 125. Combining these three values generates a combined feature vector representing the physical properties of the button area, with values [0.01, 0.08, 125].
[0102] S230: Calculate the similarity between the combined feature vector and the preset feature vector corresponding to each keyword in the surface feature vocabulary set.
[0103] A predefined feature vector refers to a high-dimensional vector defined in a pre-built feature library for each keyword in the surface feature vocabulary set, such as sandblasting or anodizing, that represents the standard physical properties of that keyword. The dimension and structure of this vector are consistent with the combined feature vector. For example, in the predefined feature library, the predefined feature vectors corresponding to the keywords "sandblasting," "anodizing," and "powder coating" can be represented as follows: , and Preset feature vectors for sandblasting process: Preset feature vector for anodizing process: Preset feature vector for powder coating process: The components of each vector. These represent the standard or typical local curvature, rate of change of normal vector, and reflection intensity values for the corresponding process, respectively.
[0104] Similarity is a numerical value used to measure how close two vectors are, and it can be obtained by calculating the distance or angle between the vectors. The core of this step is to quantitatively compare the physical properties of the area under test with the standard physical properties of known processes or materials.
[0105] First, a pre-defined feature vector, corresponding one-to-one with all keywords in the surface feature vocabulary set, is loaded from a pre-defined feature library. Next, the combined feature vector generated in the previous step is compared with each pre-defined feature vector in the feature library for similarity calculation. This calculation can employ various mathematical methods, such as calculating the cosine similarity between two vectors. Through this process, the matching score between the physical properties of the area under test and each known process or material can be obtained. Taking the aforementioned button area of a computer mouse as an example, its combined feature vector [0.01, 0.08, 125] is compared with three pre-defined feature vectors representing "sandblasting," "anodizing," and "powder coating" in the pre-defined feature library. Assume the pre-defined feature vector for "sandblasting" is [0.05, 0.20, 150], the pre-defined feature vector for "anodizing" is [0.01, 0.10, 130], and the pre-defined feature vector for "powder coating" is [0.15, 0.30, 180]. By calculating the cosine similarity, a set of similarity values is obtained. For example, the similarity with "anodizing" is 0.98, with "sandblasting" is 0.75, and with "powder coating" is 0.60. In this example, the combined feature vector of the button area is [0.01, 0.08, 125]. It shows the highest similarity with the preset feature vector of anodizing [0.01, 0.10, 130] because the local curvature is highly consistent and other parameters are close.
[0106] S240: Select at least one keyword with a similarity exceeding a preset similarity threshold as the attribute label corresponding to the target image region, and generate annotation information including the attribute label in the target image region.
[0107] The preset similarity threshold is a pre-defined value used to determine whether the similarity calculation result is high enough to confirm a matching relationship. First, a series of similarity values calculated in the previous step are compared, and the highest value is found. Then, this highest similarity value is compared with a preset similarity threshold. If the highest similarity value exceeds this threshold, the keyword corresponding to that similarity value is determined as the final attribute label for the target image region. Finally, a bounding box or pixel-level segmentation mask is generated around this target image region on the original color image, and an attribute label containing the keyword is attached, forming a complete annotation information. Continuing with the example of the button area of a computer mouse, a series of similarity values are calculated: 0.98 for "anodizing," 0.75 for "sandblasting," and 0.60 for "powder coating," and compared with a preset similarity threshold of 0.9. Since 0.98 is the highest value and exceeds the threshold, the corresponding keyword "anodizing" is selected as the final attribute label. Therefore, an attribute label will be automatically generated on the mouse button area selected in the color image, with the content "Process: Anodizing", completing the accurate labeling of this area.
[0108] In one feasible implementation, the construction process of the preset feature vector library is as follows: Prepare multiple sets of physical samples with known single materials or processes, such as a standard sandblasted aluminum plate and a standard anodized aluminum plate. For each physical sample, process it using steps S110 to S140 of this application. The processing also includes analyzing the brightness information of the deformable grating image sequence to calculate the reflection intensity value, thereby calculating the corresponding combined feature vector for each physical sample. After calculation, associate the known material or process keywords of each physical sample with its calculated combined feature vector, and store this association to collectively constitute the preset feature vector library.
[0109] This embodiment generates a surface reflectance map that can quantitatively characterize the optical reflectance properties of a product surface by analyzing the brightness information of a deformable grating image sequence. This solves the problem that relying solely on three-dimensional geometry cannot distinguish materials or processes with similar microstructures but different optical properties, such as differentiating between matte and glossy surfaces. Therefore, this application, by introducing the physical dimension of surface reflectance and coordinating it with geometric features for correlation matching, further enhances the ability to identify product material and process details, resulting in more accurate and detailed data annotation results.
[0110] In one feasible implementation, step S120 performs semantic analysis on the surface attribute text data to extract keywords corresponding to the material and process, obtaining a surface feature vocabulary set, including:
[0111] Syntactic structure analysis is performed on surface attribute text data to identify and extract noun phrases and adjective phrases describing the physical attributes of goods, resulting in a candidate phrase set.
[0112] A candidate phrase set is a temporary list of words or phrases that may contain descriptions of material or process, selected from surface attribute text data through preliminary syntactic analysis.
[0113] First, the input surface attribute text data is part-of-speech tagging, assigning each word a grammatical role, such as a noun or adjective. Next, a rule-based phrase extractor is applied, which searches for and combines words that conform to a preset grammatical pattern. For example, a regular expression-based phrase extractor can be used, where the rule can be defined as "one or more adjectives followed by a noun," to extract adjectival phrases like "frosted glass," or directly extract individual nouns as noun phrases. Taking the surface attribute text data of a high-end computer mouse, "high-end computer mouse, using sandblasting, anodizing, and powder coating processes," as an example, part-of-speech tagging and phrase extraction might yield a candidate phrase set containing phrases such as "high-end computer mouse," "sandblasting," "anodizing," "powder coating," and "process."
[0114] Each phrase in the candidate phrase set is matched with a preset material terminology library and a preset process terminology library. The successfully matched words are then categorized to generate a material keyword subset and a process keyword subset.
[0115] The material terminology library and the process terminology library refer to two pre-built knowledge bases that respectively store a large number of basic material terms and basic process terms. For example, the material terminology library may include terms such as metal, plastic, and wood, while the process terminology library may include terms such as sandblasting, polishing, and wire drawing. These two terminology libraries can be built and expanded by organizing publicly available technical materials, standards, and commonly used descriptive terms from e-commerce platforms in relevant industries. The material keyword subset and the process keyword subset are two independent keyword lists obtained by matching candidate phrases with these two knowledge bases and classifying them separately.
[0116] First, iterate through each phrase in the candidate phrase set generated in the previous step. For each phrase, perform string matching or semantic similarity comparison with all roots in the material and process root word libraries. For example, a word vector-based similarity calculation method, such as Word2Vec or GloVe models, can be used to convert both phrases and roots into vectors and then calculate their cosine similarity. If a phrase's similarity to a root word in the process root word library exceeds a preset threshold, the phrase is included in the process keyword subset. Continuing with the candidate phrase set for computer mice as an example, each phrase is matched with the root word library. "Sandblasting," "anodizing," and "powder coating" all highly match words in the process root word library and are therefore included in the process keyword subset; while phrases such as "high-end computer mouse" and "process" are ignored due to low matching. At this point, the material keyword subset is empty.
[0117] By combining subsets of material keywords and subsets of process keywords, a set of surface feature terms is generated.
[0118] The generated subsets of material keywords and process keywords are merged, and any duplicates are removed to form a single, unique keyword list. This final list serves as the surface feature vocabulary for subsequent steps. Using the computer mouse example, combining the empty subset of material keywords with the subset of process keywords containing "sandblasting," "anodizing," and "powder coating" results in a final surface feature vocabulary list containing these three keywords.
[0119] In one feasible implementation, step S130 calculates the three-dimensional spatial coordinates of multiple points on the surface of the target product by decoding the encoded information of the deformed grating image in the deformed grating image sequence, and aggregates the three-dimensional spatial coordinates to generate a three-dimensional point cloud model representing the contour of the target product surface, including:
[0120] By decoding the binary coded pattern used to determine the fringe period in the deformed grating image sequence, a unique fringe period number is determined for each pixel position.
[0121] Deformed grating image sequences typically contain multiple coded patterns, such as binary coded patterns for coarse positioning and phase-shifted sinusoidal fringe patterns for precise positioning. A binary coded pattern is a set of alternating black and white stripes whose stripe widths vary according to a binary rule, used for coarse spatial encoding of the measurement field of view. The fringe period number is an integer value assigned to each pixel after decoding this binary coded pattern; this integer value uniquely identifies the period interval of the sinusoidal fringe corresponding to that pixel.
[0122] Each binary-coded pattern in the deformable grating image sequence is processed sequentially. For each pixel in the image, a binary code string is determined based on whether its brightness value under different binary-coded patterns is higher or lower than a brightness threshold. For example, taking a high-end computer mouse as an example, if the brightness sequence of a pixel in its image is "bright, dark, bright, bright" under four binary-coded patterns, then its code string is determined to be "1011". Then, this binary code string is converted into a decimal integer, i.e., 11, and this integer 11 is the stripe period number of that pixel position.
[0123] Based on the brightness variation at the same pixel location presented by the sinusoidal fringe pattern used for localization in the deformable grating image sequence, a wrapping phase value that cycles within a single fringe period is calculated.
[0124] A sinusoidal fringe pattern refers to a set of fringe images whose brightness varies according to a sinusoidal function and is regularly shifted in a sequence, used to achieve sub-pixel-level precise positioning. The wrapping phase value is a phase value between 0 and 2π radians calculated by analyzing the brightness changes of a pixel under multiple phase-shifted sinusoidal fringe patterns, representing the relative position of that pixel within a single fringe period.
[0125] Multiple phase-shifted sinusoidal fringe patterns are extracted from a deformable grating image sequence. For each pixel in the image, its brightness value sequence under these patterns is obtained. Then, this brightness value sequence is substituted into a standard phase-shift calculation model for solution. For example, a four-step phase-shift algorithm, which includes phase values... It can be calculated using formula (1).
[0126] (1)
[0127] in, , , and These represent the image coordinates respectively. The brightness value of a pixel under four sinusoidal stripe patterns with phase shifts of 0, π / 2, π, and 3π / 2 respectively.
[0128] For example, by analyzing its brightness changes under four phase-shift patterns, its wrapping phase value can be calculated to be 1.2π radians.
[0129] The fringe period number is used to determine the integer multiple of the phase period to which the wrapped phase value belongs, and the wrapped phase value is superimposed on the integer multiple of the phase period to calculate an absolute phase value without periodic jumps.
[0130] The absolute phase value is a continuous phase value that monotonically increases and has no periodic jumps throughout the entire measurement field of view. It uniquely identifies the position of each pixel in global space.
[0131] For each pixel in the image, its fringe period number obtained in the first step is multiplied by a constant (i.e., 2π) representing the phase range of a single period, resulting in an integer multiple of the phase period. This integer multiple of the phase period is then added to the wrapping phase value obtained for that pixel in the second step to calculate the absolute phase value of the pixel. For example, if its fringe period number is 11 and its wrapping phase value is 1.2π, then its absolute phase value is calculated as 11 multiplied by 2π plus 1.2π, which is 23.2π.
[0132] Based on the absolute phase value and the system parameters of the geometric relationship between the projection of the coded grating pattern and the capture of the deformed grating image, the three-dimensional spatial coordinates are calculated by triangulating the absolute phase value and its corresponding pixel position.
[0133] System parameters refer to a set of values obtained through a pre-calibrated system calibration process that precisely describes the relative position, orientation, and internal optical parameters between the projection device and the image acquisition device.
[0134] For each pixel in the image, its calculated absolute phase value, along with its two-dimensional pixel coordinates and preset system parameters, are substituted into a pre-established triangulation geometric model for solution. This model calculates the spatial intersection between the observation ray determined by the pixel coordinates and the projection plane determined by the absolute phase, thereby calculating the three-dimensional spatial coordinates of each pixel. Taking the aforementioned pixel on the computer mouse as an example, substituting its absolute phase value of 23.2π, pixel coordinates, and system parameters into the model, its three-dimensional spatial coordinates can be calculated as (10.5, 30.2, 55.8).
[0135] A three-dimensional point cloud model is constructed based on all three-dimensional spatial coordinates in a unified coordinate system.
[0136] The 3D spatial coordinates calculated for all pixels in the previous step are then aggregated within a unified world coordinate system. By organizing all these 3D spatial coordinates containing spatial location information as a set of points, a 3D point cloud model capable of representing the complete surface contour of the target product is ultimately constructed. Repeating the above process for all pixels on the computer mouse image yields hundreds of thousands of 3D spatial coordinate points. Aggregating these points together constructs the complete 3D point cloud model of the mouse.
[0137] Figure 3A flowchart illustrating a method for generating geometric feature descriptors according to an embodiment of this application is shown.
[0138] In one feasible implementation, the geometric feature parameters include local curvature and the rate of change of the normal vector. Step S140 quantifies the microscopic morphology of the target product surface by calculating the geometric feature parameters of local regions in the 3D point cloud model, generating a set of geometric feature descriptors, such as... Figure 3 As shown, the method includes:
[0139] S310: Using each three-dimensional spatial coordinate point in the three-dimensional point cloud model as the center coordinate point, select neighboring three-dimensional spatial coordinate points to form a local region point set corresponding to each three-dimensional spatial coordinate point.
[0140] A local point set refers to a subset of points within a certain radius of a given point in a 3D point cloud model. The process involves iterating through each 3D spatial coordinate point in the point cloud model and using it sequentially as the center point. For each center point, a neighborhood search algorithm is used to determine its nearest 3D spatial coordinate points. For example, the K-nearest neighbor search algorithm can be used to find the few points with the closest spatial distance to the center point. Taking a 3D spatial coordinate point located in the sandblasting area of a high-end computer mouse's 3D point cloud model as an example, K-nearest neighbor search can find 50 nearest neighbor points. These 51 points together constitute a local point set corresponding one-to-one with the center point.
[0141] S320: For each local region point set, an approximate surface of the local region point set is obtained through surface fitting, and based on the approximate surface, the local curvature of the center coordinate point is obtained through differential geometry calculation.
[0142] An approximate surface is a continuous, smooth surface constructed mathematically that best approximates the spatial distribution of a set of points in a local region. Local curvature is a numerical value that quantifies the degree of curvature of a surface at a certain point; a larger curvature value indicates that the surface is more severely curved near that point.
[0143] For each local point set generated at the central coordinate point, a surface fitting algorithm, such as the moving least squares method, is used to construct a parameterized approximate surface that represents the geometry of the point set. Then, based on the mathematical equations of this approximate surface, the first and second partial derivatives of the surface at the central coordinate point are solved using differential geometry methods, and the principal curvature or Gaussian curvature of the central coordinate point is calculated based on these derivative values.
[0144] In another possible implementation, when using covariance analysis, local curvature This can be obtained by performing eigenvalue decomposition on the covariance matrix of the local point set. If the eigenvalues after decomposition are... And satisfy Then the local curvature can be represented by the ratio of the minimum eigenvalue to the sum of all eigenvalues, as shown in formula (2).
[0145] (2)
[0146] in, These are the three eigenvalues of the local point set covariance matrix, which represent the degree of dispersion of the point set in the three principal directions. For example, after performing surface fitting on the local point set, the calculated local curvature may be a small value, such as 0.01, which indicates that the sandblasted surface where the point is located is relatively flat macroscopically.
[0147] S330: Based on an approximate surface, the surface normal vectors of all three-dimensional spatial coordinate points within the local region point set are analyzed, and the rate of change of the normal vector of the center coordinate point is obtained by calculating the angular dispersion of the surface normal vector of each three-dimensional spatial coordinate point in the local region point set.
[0148] A surface normal vector is a unit vector perpendicular to the tangent plane of a surface at a point, used to represent the surface orientation at that point. The rate of change of the normal vector is a value that quantifies the consistency of the normal direction within a region; a larger rate of change indicates that the surface of that region is more uneven or contains more details.
[0149] Based on the approximate surface constructed for each local region point set in the previous step, its mathematical equations are used to analyze the corresponding surface normal vectors for each of the three-dimensional spatial coordinate points within that local region point set. Then, the angular dispersion among these surface normal vectors is calculated; for example, the standard deviation of the angle between all normal vectors and the average normal vector can be calculated. This calculated angular dispersion value is used as the rate of change of the normal vector at the center coordinate point of that local region point set. For example, due to the microscopic irregularities of the sandblasted surface, the surface normal vector directions of the 51 points within its local region point set will have some random fluctuations. After calculating its angular dispersion, a moderate rate of change of the normal vector may be obtained, such as 0.08.
[0150] S340: Combine the local curvature value and the rate of change of the normal vector to obtain the geometric feature descriptor of the center coordinate point, and form a set of geometric feature descriptors based on the geometric feature descriptors of all three-dimensional space coordinate points.
[0151] For each center coordinate point, its calculated local curvature value and calculated normal vector change rate value are combined. This combination can be a simple two-dimensional vector, where the first element is the local curvature value and the second element is the normal vector change rate value. This vector is the geometric feature descriptor for that center coordinate point. Taking the center coordinate point on the mouse as an example, its local curvature value of 0.05 and normal vector change rate value of 0.2 are combined to form a two-dimensional vector [0.05, 0.2], which serves as the geometric feature descriptor for that point. After repeating the entire process from S310 to S340 for all three-dimensional spatial coordinate points in the three-dimensional point cloud model, a corresponding geometric feature descriptor will be generated for each point. All these descriptors together constitute the final set of geometric feature descriptors.
[0152] In one feasible implementation, step S150 establishes a spatial position mapping relationship between the 3D point cloud model and the color image by spatially registering the 3D point cloud model with the color image, including:
[0153] Using a pre-defined camera parameter model that represents the transformation relationship between the coordinate system of a color image and a 3D point cloud model, each 3D spatial coordinate point in the 3D point cloud model is projected onto the 2D image coordinate system of the color image to obtain the 2D projected coordinates corresponding to each 3D spatial coordinate point.
[0154] A camera parameter model is a mathematical model that includes the camera's intrinsic and extrinsic parameters and distortion coefficients. It describes how a point in three-dimensional space is captured by the camera lens and imaged onto a two-dimensional image sensor. Two-dimensional projected coordinates refer to the pixel position on a two-dimensional color image calculated from the camera parameter model of a point in three-dimensional space.
[0155] First, a camera parameter model obtained beforehand through camera calibration is loaded. Then, each 3D spatial coordinate point in the 3D point cloud model is traversed. For each 3D spatial coordinate point, its 3D coordinates are substituted into the projection transformation formula of the camera parameter model for calculation. This calculation process simulates the imaging optical path of the camera, ultimately solving for the 2D projection coordinates of each 3D spatial coordinate point on the color image. Taking a point in the 3D point cloud model of a high-end computer mouse as an example, its 3D spatial coordinates are (10.5, 30.2, 55.8). After calculation by the camera parameter model, its 2D projection coordinates on the color image may be (850, 620).
[0156] For each pixel in a color image, find at least one three-dimensional spatial coordinate point within the neighborhood of the pixel's corresponding two-dimensional projected coordinates.
[0157] The algorithm iterates through every pixel in the color image. For each pixel, it uses the 2D projection coordinates projected onto the pixel location or its vicinity from the previous step as a basis to search within a preset neighborhood. For example, a fast nearest neighbor search algorithm based on a 2D KD tree can be used to find one or more 2D projection coordinates closest to the current pixel. Taking a color image of a computer mouse as an example, for the pixel with coordinates (851, 621), since there may not be a 3D point projected exactly to this precise location, the system will search within a 5x5 pixel window around it and find that the 2D projection coordinates projected onto (850, 620) in the previous step are the closest.
[0158] Each pixel in the color image is associated with at least one found 3D spatial coordinate point to establish a spatial location mapping relationship.
[0159] Create a data structure, such as a lookup table or hash map, to store the associations between 2D pixels and 3D spatial coordinates. Iterate through each pixel in the color image and associate it with the index or coordinate value of one or more 3D spatial coordinates found in the previous step. For example, associate pixel (851, 621) with the 3D point at coordinates (10.5, 30.2, 55.8). After all pixels have been processed, a complete spatial mapping relationship is established. Through this mapping relationship, you can quickly find the corresponding 3D point for any pixel, or find the position of any 3D point in the image.
[0160] The multimodal data collaborative data annotation method provided in this application introduces deformable grating image sequences and decodes them to generate a 3D point cloud model representing the surface contour of a product, thereby quantifying the microscopic morphology of the product surface. This solves the problem of existing technologies that rely solely on 2D color images and cannot perceive physical structural details. By associating the quantified microscopic morphology features with text keywords, this application enables the annotation process to go beyond macroscopic visual appearance and delve into details such as materials and processes determined by physical structure. Compared with existing technologies, this application can identify and annotate specific material and process details, improving the accuracy and comprehensiveness of product image data annotation.
[0161] Figure 4 This is a schematic diagram illustrating a specific implementation of a multimodal data collaborative data annotation system provided in this application. (Refer to...) Figure 4 The system may include:
[0162] The acquisition module 410 is used to acquire color images of the target product, surface attribute text data, and deformable grating image sequences. The surface attribute text data includes text describing the material and process of the target product. The deformable grating image sequence is obtained by projecting an coded grating pattern onto the physical surface of the target product and capturing deformable grating images modulated by the contour of the physical surface from different angles.
[0163] Extraction module 420 is used to perform semantic analysis on surface attribute text data to extract keywords corresponding to material and process, and obtain a surface feature vocabulary set;
[0164] The generation module 430 is used to calculate the three-dimensional spatial coordinates of multiple points on the surface of the target product by decoding the encoded information of the deformed grating image in the deformed grating image sequence, and to aggregate the three-dimensional spatial coordinates to generate a three-dimensional point cloud model representing the contour of the target product surface.
[0165] The generation module 430 is also used to quantify the micro-morphology of the target commodity surface by calculating the geometric feature parameters of local regions in the three-dimensional point cloud model, and to generate a set of geometric feature descriptors.
[0166] Module 440 is used to establish a spatial position mapping relationship between the 3D point cloud model and the color image by spatially registering the 3D point cloud model with the color image.
[0167] The matching module 450 is used to associate and match geometric feature descriptors with keywords in the surface feature vocabulary set based on spatial location mapping relationship, so as to identify the material and process corresponding to the target image region in the color image, and generate annotation information including attribute labels on the target image region in the color image.
[0168] The multimodal data collaborative data annotation system of this application is used to implement the aforementioned multimodal data collaborative data annotation method. Therefore, the specific implementation of the multimodal data collaborative data annotation system can be found in the embodiment section of the multimodal data collaborative data annotation method above. The specific implementation can be referred to the description of the corresponding embodiments, which will not be repeated here.
[0169] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application is shown.
[0170] The electronic device may include a processor 510 and a memory 520 storing computer program instructions.
[0171] Specifically, the processor 510 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0172] Memory 520 may include mass storage for data or instructions. For example, and not limitingly, memory 520 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 520 may include removable or non-removable (or fixed) media. Where appropriate, memory 520 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 520 is non-volatile solid-state memory.
[0173] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of this disclosure.
[0174] The processor 510 reads and executes computer program instructions stored in the memory 520 to implement any of the multimodal data collaborative data annotation methods in the above embodiments.
[0175] In one example, the electronic device may also include a communication interface 530 and a bus 540. Wherein, such as Figure 5 As shown, the processor 510, memory 520, and communication interface 530 are connected through bus 540 and complete communication with each other.
[0176] The communication interface 530 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0177] Bus 540 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 540 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0178] The electronic device can execute the multimodal data collaborative data annotation method in the embodiments of this application, thereby realizing the multimodal data collaborative data annotation method described in conjunction with the accompanying drawings.
[0179] Furthermore, in conjunction with the multimodal data collaborative data annotation method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the multimodal data collaborative data annotation methods in the above embodiments.
[0180] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0181] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0182] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0183] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0184] The foregoing has provided a detailed description of a multimodal data collaborative data annotation method, system, electronic device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A data annotation method for multimodal data collaboration, characterized in that, include: The method acquires a color image of the target product, surface attribute text data, and a deformable grating image sequence. The surface attribute text data includes text describing the material and process of the target product. The deformable grating image sequence is obtained by projecting an coded grating pattern onto the physical surface of the target product and capturing deformable grating images modulated by the contour of the physical surface from different angles. Semantic analysis is performed on the surface attribute text data to extract keywords corresponding to the material and process, thereby obtaining a surface feature vocabulary set; By decoding the encoded information of the deformed grating image in the deformed grating image sequence, the three-dimensional spatial coordinates of multiple points on the surface of the target product are calculated, and the three-dimensional spatial coordinates are aggregated to generate a three-dimensional point cloud model representing the contour of the surface of the target product. By calculating the geometric feature parameters of local regions in the three-dimensional point cloud model, the micro-morphology of the target product surface is quantified, and a set of geometric feature descriptors is generated. By spatially registering the 3D point cloud model with the color image, a spatial position mapping relationship between the 3D point cloud model and the color image is established; Based on the spatial location mapping relationship, the geometric feature descriptor is associated and matched with the keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image.
2. The method according to claim 1, characterized in that, The method further includes: Based on the average brightness of the grating stripes at the same pixel location in multiple deformed grating images in the deformed grating image sequence, the reflection intensity value representing the surface reflectivity of each pixel location on the physical surface of the target product is calculated, and a surface reflectivity map is generated. The reflection intensity value in the surface reflectivity map is associated with the corresponding spatial coordinate point in the three-dimensional point cloud model; Based on the spatial location mapping relationship, the geometric feature descriptor is associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image, including: Based on the spatial location mapping relationship, the reflection intensity values corresponding to the spatial coordinate points in the geometric feature descriptor and the three-dimensional point cloud model are associated and matched with the keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image.
3. The method according to claim 2, characterized in that, Based on the spatial location mapping relationship, the geometric feature descriptor and the reflection intensity value corresponding to the spatial coordinate point in the 3D point cloud model are associated and matched with keywords in the surface feature vocabulary set to identify the material and process corresponding to the target image region in the color image, and annotation information including attribute labels is generated on the target image region in the color image, including: Based on the spatial location mapping relationship, a set of corresponding three-dimensional spatial coordinate points is determined for the target image region in the color image; Extract the geometric feature descriptor and the reflection intensity value corresponding to the set of three-dimensional spatial coordinate points, and combine the geometric feature descriptor and the reflection intensity value to generate a combined feature vector; Calculate the similarity between the combined feature vector and the preset feature vector corresponding to each keyword in the surface feature vocabulary set; At least one keyword with a similarity exceeding a preset similarity threshold is selected as the attribute label corresponding to the target image region, and the annotation information including the attribute label is generated in the target image region.
4. The method according to claim 1, characterized in that, The semantic analysis of the surface attribute text data is performed to extract keywords corresponding to the material and process, resulting in a surface feature vocabulary set, including: Syntactic structure analysis is performed on the surface attribute text data to identify and extract noun phrases and adjective phrases describing the physical attributes of the goods, resulting in a candidate phrase set; Each phrase in the candidate phrase set is matched with a preset material terminology library and a preset process terminology library, and the successfully matched words are classified to generate a material keyword subset and a process keyword subset. The material keyword subset and the process keyword subset are combined to generate the surface feature vocabulary set.
5. The method according to claim 1, characterized in that, The step of decoding the encoded information of the deformed grating image in the deformed grating image sequence, calculating the three-dimensional spatial coordinates of multiple points on the surface of the target product, and aggregating the three-dimensional spatial coordinates to generate a three-dimensional point cloud model representing the contour of the target product surface includes: By decoding the binary encoded pattern used to determine the stripe period in the deformed grating image sequence, a unique stripe period sequence number is determined for each pixel position. Based on the brightness change at the same pixel position presented by the sinusoidal fringe pattern used for positioning in the deformed grating image sequence, a wrapping phase value that cycles within a single fringe period is calculated. The integer multiple of the phase period to which the wrapped phase value belongs is determined by using the stripe period number, and the wrapped phase value is superimposed on the integer multiple of the phase period to calculate an absolute phase value without periodic jumps. Based on the absolute phase value and preset system parameters representing the geometric relationship between the coded grating pattern projection and the deformed grating image capture, the three-dimensional spatial coordinates are calculated by performing triangulation calculations on the absolute phase value and its corresponding pixel position. Under a unified coordinate system, the three-dimensional point cloud model is constructed based on all the three-dimensional spatial coordinates.
6. The method according to claim 1, characterized in that, The geometric feature parameters include local curvature and the rate of change of the normal vector; The method involves calculating the geometric feature parameters of local regions in the 3D point cloud model to quantify the microscopic morphology of the target product surface, generating a set of geometric feature descriptors, including: Taking each three-dimensional spatial coordinate point in the three-dimensional point cloud model as the center coordinate point, select neighboring three-dimensional spatial coordinate points to form a local region point set corresponding to each three-dimensional spatial coordinate point. For each set of points in the local region, an approximate surface of the set of points in the local region is obtained by surface fitting, and the local curvature of the center coordinate point is obtained by differential geometry calculation based on the approximate surface. Based on the approximate surface, the surface normal vectors of all three-dimensional spatial coordinate points in the local region point set are analyzed, and the rate of change of the normal vector of the center coordinate point is obtained by calculating the angular dispersion of the surface normal vector of each three-dimensional spatial coordinate point in the local region point set. The local curvature value and the rate of change of the normal vector are combined to obtain the geometric feature descriptor of the center coordinate point, and a set of geometric feature descriptors is formed based on the geometric feature descriptors of all three-dimensional space coordinate points.
7. The method according to claim 1, characterized in that, The step of establishing a spatial position mapping relationship between the 3D point cloud model and the color image by spatially registering the 3D point cloud model and the color image includes: Using a preset camera parameter model that represents the transformation relationship between the coordinate system of the color image and the coordinate system of the three-dimensional point cloud model, each three-dimensional spatial coordinate point in the three-dimensional point cloud model is projected onto the two-dimensional image coordinate system of the color image to obtain the two-dimensional projection coordinates corresponding to each three-dimensional spatial coordinate point. For each pixel in the color image, find at least one three-dimensional spatial coordinate point within the neighborhood of the two-dimensional projection coordinates corresponding to the pixel. Each pixel in the color image is associated with the at least one found three-dimensional spatial coordinate point to establish the spatial position mapping relationship.
8. A multimodal data collaborative data annotation system, characterized in that, include: The acquisition module is used to acquire color images, surface attribute text data, and deformable grating image sequences of the target product. The surface attribute text data includes text describing the material and process of the target product. The deformable grating image sequence is obtained by projecting an coded grating pattern onto the physical surface of the target product and capturing deformable grating images modulated by the contour of the physical surface from different angles. The extraction module is used to perform semantic analysis on the surface attribute text data to extract keywords corresponding to the material and process, and obtain a surface feature vocabulary set. The generation module is used to calculate the three-dimensional spatial coordinates of multiple points on the surface of the target product by decoding the encoding information of the deformed grating image in the deformed grating image sequence, and to aggregate the three-dimensional spatial coordinates to generate a three-dimensional point cloud model representing the contour of the surface of the target product. The generation module is also used to quantify the micromorphology of the target product surface by calculating the geometric feature parameters of local regions in the three-dimensional point cloud model, and generate a set of geometric feature descriptors; The construction module is used to establish a spatial position mapping relationship between the 3D point cloud model and the color image by spatially registering the 3D point cloud model with the color image; The matching module is used to associate and match the geometric feature descriptor with keywords in the surface feature vocabulary set based on the spatial location mapping relationship, so as to identify the material and process corresponding to the target image region in the color image, and generate annotation information including attribute labels on the target image region in the color image.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the data annotation method for multimodal data collaboration as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the implementation of the multimodal data collaborative data annotation method as described in any one of claims 1 to 7.