Object recognition system and method based on multi-modal data fusion

By establishing a spatiotemporal synchronization model and multimodal data processing technology, semantic and texture features of objects are extracted, and the problems of poor synchronization accuracy and spatial feature extraction in the existing technology are solved, achieving a more accurate and robust object recognition effect.

CN119992216APending Publication Date: 2025-05-13JINING ZHENGHE INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510204829.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the synchronization accuracy of different modal data in multimodal data fusion, and the spatial characteristics of transparent objects cannot be fully mined through simple edge detection and general feature extraction, resulting in poor recognition effect.

Method used

By establishing a spatiotemporal synchronization model, synchronize the acquisition equipment, obtain two-dimensional image data of objects under different angles and lighting conditions and depth data at different distances, perform denoising, enhancement and normalization processing, extract image semantic features and image texture features, and combine density clustering algorithms and voxel unit division to extract the spatial characteristics of objects.

Benefits of technology

It improves the synchronization accuracy and recognition accuracy of multimodal data, can capture the characteristics of objects more comprehensively, and improves the robustness and accuracy of object recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992216A_ABST
    Figure CN119992216A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image analysis, and discloses an object recognition system and method based on multi-modal data fusion. The method comprises the following steps: establishing a space-time synchronization model for acquisition equipment; collecting two-dimensional image data and depth data of an object, and preprocessing the two-dimensional image data and the depth data to obtain image data and depth data; performing denoising processing, enhancement processing and normalization processing on the two-dimensional image data of the object to obtain image data; performing filtering processing and smoothing processing on the depth data to obtain a depth point; analyzing the image data to obtain image semantic features; obtaining depth data of a corresponding object surface in the image data, and performing feature extraction on the depth data to obtain image texture features; taking the image semantic features and the image texture features as input of an object classification model to obtain an object recognition result; according to the method, potential category information is mined, the object category is comprehensively judged, and the method has high accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image analysis, and more specifically, to an object recognition system and method based on multimodal data fusion. Background Art

[0002] Traditional object recognition methods are mostly based on single-modal data, such as using only image data for recognition; in order to overcome the limitations of single-modal recognition, object recognition technology that fused multimodal data came into being.

[0003] A Chinese patent with publication number CN109903323A discloses a training method, device, storage medium and terminal for transparent object recognition, including: S1, establishing a first data set with multiple RGB images and a second data set with multiple depth images, the multiple RGB images correspond one-to-one to the multiple depth images respectively; S2, establishing a multimodal fusion deep convolutional neural network structure N1, the N1 is used to extract multiple RGB images for separate training and extract multiple depth images for separate training, so as to extract the first feature information of the RGB image and the second feature information of the depth image respectively to obtain a network weight model M1; S3, establishing a multimodal shared deep convolutional network structure N2, inputting the first feature information and the second feature information into the N2 for fusion training, so as to output the classification parameter information and position coordinate information of the object and obtain the network weight model M2; S4, re-inputting other multiple pairs of RGB images and depth images to adjust the parameters of the network weight model M1 and the network weight model M2 to obtain optimized network weight models M11 and M22. This invention trains data of different modalities (RGB images and depth images) separately, learns the characteristics of the modality itself through a series of neural networks, and then connects them through fusion. After a series of shared convolutional layers, the characteristics of each modality are complementary learned. The fusion of RGB information and depth information can improve the recognition of transparent objects.

[0004] Although the above method can meet most scenarios, research and practical application of the above method and existing technology have found that the above method and existing technology have at least the following defects:

[0005] It is difficult to ensure the synchronization accuracy of data of different modalities; extracting the spatial features of objects only by using simple edge detection, general feature extraction networks based on deep learning, etc., cannot fully explore the spatial features of transparent objects, resulting in the extracted features not being able to reflect the spatial characteristics of the object well.

[0006] In view of this, the present invention proposes an object recognition system and method based on multimodal data fusion to solve the above problems. Summary of the invention

[0007] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: an object recognition system and method based on multimodal data fusion, the method comprising the following steps:

[0008] Establishing a spatiotemporal synchronization model for acquisition equipment; the acquisition equipment includes image acquisition equipment and depth acquisition equipment;

[0009] The synchronized image acquisition device is used to acquire two-dimensional image data of the object at different angles and under different lighting conditions, and the synchronized depth acquisition device is used to scan and acquire depth data at different distances;

[0010] Perform denoising, enhancement and normalization on the two-dimensional image data of the object to obtain image data; perform filtering and smoothing on the depth data to obtain depth points;

[0011] Using the image data as input of the first extraction model to obtain image semantic features;

[0012] Obtaining depth data of the surface of the object corresponding to the image data, and performing feature extraction on the depth data to obtain image texture features;

[0013] The image semantic features and image texture features are used as input to the object classification model to obtain the object recognition results.

[0014] Furthermore, the method for obtaining depth data corresponding to the surface of the object in the image data includes:

[0015] A nested loop including an outer loop and an inner loop is preset to traverse all pixel points on the surface of the object corresponding to the image data, wherein the outer loop is used to traverse each row of the image data, and the inner loop is used to traverse each column of the image data, and the depth points of the coordinates of all pixel points on the surface of the object corresponding to the XYZ-O three-dimensional rectangular coordinate system are obtained in turn.

[0016] Furthermore, the method for obtaining image texture features includes:

[0017] A density clustering algorithm is used to perform a preliminary analysis of the depth points, identify different density areas, and divide them using voxel units. The connectivity of the voxel units, the Euler number of the object, and the curvature of the object are analyzed to extract the spatial characteristics of the object; the spatial characteristics include the local density peak position, voxel connectivity characteristics, Euler number and curvature characteristics; the spatial characteristics are used as image texture features.

[0018] Furthermore, the depth points are preliminarily analyzed by density clustering algorithm to identify different density areas, and the method of dividing them using voxel units includes:

[0019] A density clustering algorithm based on a preset first neighborhood radius and a preset first minimum number of points traverses the depth point set, counts the number of points within the preset first neighborhood radius of each depth point, and compares the number of points with the preset first minimum number of points. When the number of points is higher than the preset first minimum number of points, the corresponding depth point is marked as a core point, and a new cluster is created, and the points within the preset first neighborhood radius of the core point are added to the cluster in sequence until all points are marked as belonging to a certain cluster;

[0020] According to the range and resolution of the depth points, the three-dimensional space is divided into voxel units, and the depth points are mapped to the corresponding voxel units.

[0021] Furthermore, the method of dividing the three-dimensional space into voxel units according to the depth point range and resolution includes:

[0022] Traverse the depth points and obtain the minimum value X of the depth point in the X-axis, Y-axis and Z-axis directions of the XYZ-O three-dimensional rectangular coordinate system established with the acquisition device as the origin. min , Y min and Z min ; and the maximum value X of the depth point in the X-axis, Y-axis and Z-axis directions max , Y max and Z max ;

[0023] Preset voxel resolutions ΔX, ΔY, and ΔZ in the X, Y, and Z directions;

[0024] The number of voxel units in the axial direction of the X-axis, Y-axis and Z-axis is obtained by respectively calculating the ratio of the difference between the maximum value and the minimum value in the X-axis, Y-axis and Z-axis directions of the XYZ-O three-dimensional rectangular coordinate system to the resolution;

[0025] From the starting position in the depth point range, the voxel units are gradually divided according to the resolution;

[0026] For each voxel unit, it is identified by the index (i, j, k) in the three axes, where i corresponds to the index in the X-axis direction; j corresponds to the index in the Y-axis direction; j corresponds to the index in the Z-axis direction; the range in the X-axis direction is [X min +i·ΔX,X min +(i+1)·ΔX]; the range in the Y-axis direction is [Y min +i·ΔY,Y min +(i+1)·ΔY]; the range in the Z-axis direction is [Z min +i·ΔZ,Z min +(i+1)·ΔZ].

[0027] Furthermore, the method of mapping the depth points to corresponding voxel units includes:

[0028] For each depth point (x, y, z), the corresponding voxel index (i′, j′, k′) is calculated by the ratio of the coordinate value of the depth point in the three axes to the side length of the voxel;

[0029] The method for obtaining the local density peak position includes:

[0030] The local density value is calculated by calculating the distance relationship between the depth point and the points in the neighborhood corresponding to the depth point, and the maximum local density value is selected as the density peak position.

[0031] Furthermore, the method of obtaining voxel connectivity features includes:

[0032] Define the connectivity rules between voxel units, randomly select a starting voxel unit, traverse all voxel units using a depth-first search or breadth-first search algorithm, mark the visited voxel units, and record the connectivity paths between voxel units; identify the number of holes in the object based on the connectivity path record results; count the number of connectivity paths and the number of voxel units contained in each connectivity path, and splice them as voxel connectivity features;

[0033] Methods for obtaining Euler numbers include:

[0034] The Euler number is calculated by counting the number of vertices, edges, faces and internal cavities of the voxel unit.

[0035] Methods for obtaining the curvature characteristics of an object include:

[0036] Select a voxel point U, select V voxel points around the voxel point U, and select a parabola equation;

[0037] Substitute the coordinates of each selected voxel point into the above parabola equation, solve the equation group by the least square method, obtain the fitting coefficient that minimizes the sum of square errors, substitute the fitting coefficient into the parabola equation to obtain the fitted parabola equation; select J position points on the surface of the object, and calculate the curvature of the J position points according to the fitted parabola equation.

[0038] Furthermore, the method of obtaining the number of vertices V by counting the connection corner points between voxel units includes:

[0039] Traverse each vertex of each voxel unit, analyze whether the vertex between the current voxel unit and the adjacent voxel unit is occupied; and record the coordinates of the counted vertices through the data structure. Whenever a vertex is determined to be an object vertex, traverse whether it exists in the data structure. If not, add one to the count and add it to the data structure; if it already exists, skip it; count the count value output by the data structure as the number of vertices V;

[0040] The method of counting the shared edges between adjacent voxel units to obtain the edge number E′ includes:

[0041] Traverse all voxel units in the three coordinate axis directions of the XYZ-O three-dimensional rectangular coordinate system and the shared edges with adjacent voxel units; if at least one of the two voxel units where an edge is located is occupied, then this edge belongs to the edge of the object; record the coordinates of the two endpoints of the edge through the data structure, and each time an edge is counted, traverse whether it exists in the data structure, if not, add one to the count and add it to the data structure; if it already exists, skip it; count the count value output by the statistical data structure as the number of edges E′;

[0042] Methods for obtaining the number of faces F by counting the planes formed by voxel units include:

[0043] The plane is divided into planes perpendicular to the three coordinate axes of the XYZ-O three-dimensional rectangular coordinate system, and all voxel units are traversed to check whether the voxel unit perpendicular to any coordinate axis forms a plane with the adjacent voxel units on the other two coordinate axis planes. If θ adjacent voxels are continuously occupied on the plane, a plane area belonging to the object is formed, and the plane area is counted as a face; the coordinates of any three non-collinear depth points in the plane are recorded through a data structure, and the corresponding normal vector is calculated and recorded. Each time a face is counted, the plane vector composed of the depth coordinates of any two points in the plane is obtained, and it is determined whether the plane vector is perpendicular to the normal vector recorded in the data structure. If not, it is determined to be a new plane, and the face count is increased by 1. If it is perpendicular, it is determined to be a repeated plane and it is skipped; the count value output by the statistical data structure is used as the face number F.

[0044] Furthermore, the method for establishing a spatiotemporal synchronization model for the acquisition device includes:

[0045] Step 1: Obtain the coordinate system of each acquisition device;

[0046] Step 2, define a unified world coordinate system; the unified world coordinate system is to establish an XYZ-O three-dimensional rectangular coordinate system with the acquisition device as the origin;

[0047] Step 3: Use the rotation matrix R and the translation vector t to represent the device's posture. The rotation matrix R is used to describe the rotation relationship between the device coordinate system and the world coordinate system. The rotation matrix R satisfies R T R = I, T is the transpose of the vector, and I is the unit matrix; the translation vector t is used to represent the position of the origin of the device coordinate system in the world coordinate system;

[0048] Step 4: Express the coordinates of the data point Q collected by the acquisition device as homogeneous coordinates Then use the matrix transformation formula to transform the homogeneous coordinates Convert to the world coordinate system to obtain the world coordinate of data point Q.

[0049] Furthermore, the method of expressing the coordinates of the data points collected by the collection device as homogeneous coordinates includes:

[0050] For image acquisition devices: The internal parameters of the image acquisition device camera are obtained through the camera calibration method, including the x-axis focal length f x , y-axis focal length f y and the principal point coordinates (p x ,p y );

[0051] For a pixel point (u, v) in the image data, where u and v are the pixel point coordinates, the pixel point (u, v) is converted to its coordinates (x k ,y k ,z k ) is converted to homogeneous coordinates by the conversion formula Among them, x k is the x-axis coordinate of the pixel point (u, v) in the camera coordinate system; y k is the y-axis coordinate of the pixel point (u, v) in the camera coordinate system; z k is the z-axis coordinate of the pixel point (u, v) in the camera coordinate system; is the homogeneous coordinate of the pixel point (u, v) after the x-axis coordinate in the camera coordinate system is transformed; is the homogeneous coordinate of the pixel point (u, v) in the camera coordinate system after the y-axis coordinate is transformed; is the homogeneous coordinate of the pixel point (u, v) in the camera coordinate system after the z-axis coordinate is transformed;

[0052] For the depth acquisition device: the coordinates (a, b, c) of the data points of the depth acquisition device are expressed as homogeneous coordinates (a, b, c, 1); where a is the x-axis coordinate in the three-dimensional coordinate system of the depth acquisition device; b is the y-axis coordinate in the three-dimensional coordinate system of the depth acquisition device; and c is the z-axis coordinate in the three-dimensional coordinate system of the depth acquisition device.

[0053] Furthermore, the method for obtaining the two-dimensional image data of the object at different angles and different lighting conditions and the depth data at different distances includes:

[0054] Determine the angle range of the collected objects and plan according to the preset angle intervals;

[0055] Under the preset environment, simulate using artificial light sources of different types and intensities according to the preset lighting conditions;

[0056] According to the size of the object and the application scenario requirements, the distance range between the object and the collection device is divided, and the distance is planned according to the preset distance interval;

[0057] Build a closed collection space around the object, arrange lights around and on top of the object, and install shading facilities;

[0058] The two-dimensional image data of the object is collected according to the preset angle interval and the preset lighting conditions, and the depth data of the object is collected according to the preset distance interval.

[0059] An object recognition system based on multimodal data fusion is used to implement an object recognition method based on multimodal data fusion, including:

[0060] Device synchronization module: establish a spatiotemporal synchronization model for acquisition devices; acquisition devices include image acquisition devices and depth acquisition devices;

[0061] Data acquisition module: Use the synchronized image acquisition device to acquire the two-dimensional image data of the object at different angles and under different lighting conditions, and use the synchronized depth acquisition device to scan and acquire the depth data at different distances;

[0062] Data processing module: denoising, enhancing and normalizing the two-dimensional image data of the object to obtain image data; filtering and smoothing the depth data to obtain depth points;

[0063] The first extraction module: takes the image data as the input of the first extraction model to obtain the image semantic features;

[0064] The second extraction module: obtains the depth data of the object surface corresponding to the image data, and performs feature extraction on the depth data to obtain image texture features;

[0065] Object recognition module: Use image semantic features and image texture features as input to the object classification model to obtain object recognition results.

[0066] Technical effects and advantages of the object recognition system and method based on multimodal data fusion of the present invention:

[0067] The present invention establishes a spatiotemporal synchronization model specifically for acquisition devices to ensure that the data acquired by different acquisition devices are precisely corresponding in time and space, which is crucial for the subsequent accurate fusion of different modal data and feature extraction; it also collects two-dimensional image data of objects at different angles and different lighting conditions, as well as depth data at different distances. Such a rich data acquisition strategy can comprehensively cover the characteristic performance of objects in various scenarios, which helps to improve the robustness and accuracy of subsequent recognition; it also comprehensively extracts image semantic features and image texture features extracted from depth points from the image data, and provides rich discrimination basis for accurate object recognition through multi-dimensional features, thereby obtaining accurate recognition results of objects through object classification models. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a flow chart of an object recognition method based on multimodal data fusion of the present invention;

[0069] Figure 2 A schematic diagram of a process for obtaining the number of faces of a plane formed by counting voxel units according to the present invention;

[0070] Figure 3 This is a block diagram of the object recognition system based on multimodal data fusion of the present invention. DETAILED DESCRIPTION

[0071] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0072] Example 1

[0073] See also Figure 1 As shown, the object recognition method based on multimodal data fusion described in this embodiment includes:

[0074] A spatiotemporal synchronization model is established for the acquisition equipment; the acquisition equipment includes an image acquisition equipment and a depth acquisition equipment.

[0075] The method of establishing a spatiotemporal synchronization model for acquisition equipment includes:

[0076] Step 1: Get the coordinate system of each acquisition device.

[0077] Step 2: Define a unified world coordinate system; the unified world coordinate system is to establish an XYZ-O three-dimensional rectangular coordinate system with the acquisition device as the origin.

[0078] Step 3: Use the rotation matrix R and the translation vector t to represent the device's posture. The rotation matrix R is used to describe the rotation relationship between the device coordinate system and the world coordinate system. The rotation matrix R satisfies R T R=I, T is the transpose of the vector, and I is the unit matrix; the translation vector t is used to represent the position of the origin of the device coordinate system in the world coordinate system.

[0079] Step 4: Express the coordinates of the data point Q collected by the acquisition device as homogeneous coordinates Then use the matrix transformation formula to transform the homogeneous coordinates Convert to the world coordinate system to obtain the world coordinate of data point Q.

[0080] The matrix transformation formula is as follows:

[0081]

[0082] in, is the world coordinate of data point Q.

[0083] Methods for expressing the coordinates of data points collected by a collection device as homogeneous coordinates include:

[0084] For image acquisition devices: The internal parameters of the image acquisition device camera are obtained through the camera calibration method, including the x-axis focal length f x , y-axis focal length f y and the principal point coordinates (p x ,p y );

[0085] For a pixel point (u, v) in the image data, where u and v are the pixel coordinates, the coordinates (x k ,y k ,z k ) is converted to homogeneous coordinates through the camera calibration transformation formula Among them, x k is the x-axis coordinate of the pixel point (u, v) in the camera coordinate system; y k is the y-axis coordinate of the pixel point (u, v) in the camera coordinate system; z k is the z-axis coordinate of the pixel point (u, v) in the camera coordinate system; is the homogeneous coordinate of the pixel point (u, v) after the x-axis coordinate in the camera coordinate system is transformed; is the homogeneous coordinate of the pixel point (u, v) in the camera coordinate system after the y-axis coordinate is transformed; is the homogeneous coordinate of the pixel point (u, v) in the camera coordinate system after the z-axis coordinate is transformed;

[0086] The camera calibration conversion formula is as follows:

[0087]

[0088] For the depth acquisition device: the coordinates (a, b, c) of the data points of the depth acquisition device are expressed as homogeneous coordinates (a, b, c, 1); where a is the x-axis coordinate in the three-dimensional coordinate system of the depth acquisition device; b is the y-axis coordinate in the three-dimensional coordinate system of the depth acquisition device; and c is the z-axis coordinate in the three-dimensional coordinate system of the depth acquisition device.

[0089] During the data collection process, different sensors have their own independent coordinate systems. Converting the data they collect into a unified world coordinate system can accurately fuse the feature information about objects from different modalities, avoid information dislocation and mismatch caused by coordinate system differences, and better utilize the complementarity of data from different modalities to provide more comprehensive information for subsequent object recognition, thereby improving recognition accuracy.

[0090] Use the synchronized acquisition equipment to collect two-dimensional image data of objects at different angles and under different lighting conditions, as well as depth data at different distances;

[0091] The method for the acquisition device to acquire two-dimensional image data of an object at different angles and under different lighting conditions, as well as depth data at different distances includes:

[0092] Determine the angle range of the object to be collected and plan according to the preset angle interval; for example, with the object as the center, collect a set of data every 10° in the horizontal direction, and the angle range covers 0° to 360°; in the vertical direction, collect a set of data every 15°, ranging from the bottom to the top of the object. In this way, different angles of each surface of the object can be covered more comprehensively.

[0093] In a preset environment, artificial light sources of different types and intensities are used for simulation according to preset lighting conditions. For example, incandescent lamps and LED lamps with adjustable brightness are used to create preset lighting scenes ranging from dim to bright, from uniform light to side light, backlight, etc. by changing the number, position, angle and brightness of the lights.

[0094] According to the size of the object and the application scenario requirements, the distance range between the object and the collection device is divided, and the distance is planned according to the preset distance interval. For example, if the distance range is set between 1 meter and 2 meters, the collection device is moved according to the preset distance interval of 0.5 meters to change the distance.

[0095] Build a closed collection space around the object, arrange lights around and on top of the object, and install blackout facilities such as blackout curtains to better control the lighting environment and avoid interference from external natural light.

[0096] The two-dimensional image data of the object is collected according to the preset angle interval and the preset lighting conditions, and the depth data of the object is collected according to the preset distance interval.

[0097] Through the above method, we can comprehensively collect two-dimensional image data of objects at different angles and lighting conditions, as well as depth data at different distances from the camera, providing a rich and high-quality data foundation for subsequent object recognition tasks based on multimodal data fusion.

[0098] Perform denoising, enhancement and normalization on the two-dimensional image data of the object to obtain image data; perform filtering and smoothing on the depth data to obtain depth points;

[0099] Denoising the two-dimensional image data of the object can be done by using methods such as Gaussian filtering and median filtering to remove noise interference in the two-dimensional image data of the object, improve the quality of the two-dimensional image data of the object, and improve the accuracy of object recognition. Image enhancement operations such as histogram equalization can be performed on the two-dimensional image data of the object to improve the contrast of the image, make the features of the object more obvious, and improve the accuracy of object recognition. Normalizing the two-dimensional image data of the object can map the pixel values ​​of the two-dimensional image of the object to a specific interval, which is convenient for data processing and calculation in the subsequent object recognition process. Filtering the depth data can remove abnormal points caused by measurement errors or environmental interference, improve the quality of the depth data, and improve the accuracy of object recognition. Smoothing the depth data can make the three-dimensional structure of the object more continuous and smooth, which is convenient for subsequent spatial feature extraction and analysis to improve the efficiency of object recognition.

[0100] Using the image data as input of the first extraction model to obtain image semantic features;

[0101] The training method of the first extraction model includes:

[0102] M groups of image training data are collected in advance, and the image training data includes image data and image semantic features corresponding to the image data.

[0103] Each set of image training data is used as the input of the video analysis model, and the video analysis model uses the image semantic features corresponding to each set of image data as output, and uses the actual image semantic features corresponding to each set of image data as the prediction target; the training target is to minimize the sum of the prediction errors of all image semantic features; the video analysis model is trained until the sum of the prediction errors reaches convergence, and the training is stopped; the video analysis model is a deep neural network model. The loss function value of the video analysis model is the mean square error.

[0104] Prediction error formula: m is the mth group of image data; M is the number of groups of image data; is the image semantic feature corresponding to the mth group of image data; r m is the actual image semantic feature corresponding to the mth group of image data.

[0105] Image semantic features can provide high-level semantic information about the category to which an object belongs. For example, through the first extraction model, it is possible to directly identify objects in an image as "cars", "pedestrians", "buildings" and other categories. This semantic-based category determination is the core step of object recognition, which can quickly classify objects, narrow the scope of subsequent recognition, and provide a basis for more accurate recognition.

[0106] Obtain the depth data of the object surface corresponding to the image data, and extract the features of the depth data to obtain the image texture features; the image texture features are the spatial features corresponding to the object, and the spatial features include the local density peak position, voxel connectivity features, Euler number and curvature features; starting from the texture image, extracting the depth data points in combination with the three-dimensional depth data can not only realize the fusion of multimodal data, but also provide richer spatial information for object recognition, and effectively distinguish objects of different shapes and structures;

[0107] The method of obtaining the depth data of the surface of the object corresponding to the image data includes:

[0108] A nested loop including an outer loop and an inner loop is preset to traverse all pixel points on the surface of the object corresponding to the image data, wherein the outer loop is used to traverse each row of the image data, and the inner loop is used to traverse each column of the image data, and the depth points of the coordinates of all pixel points on the surface of the object corresponding to the XYZ-O three-dimensional rectangular coordinate system are obtained in turn.

[0109] Methods for obtaining image texture features include:

[0110] A density clustering algorithm is used to perform a preliminary analysis of the depth points, identify different density areas, and divide them into voxel units. The connectivity of the voxel units, the Euler number of the object, and the curvature of the object are analyzed to extract the spatial characteristics of the object. The spatial characteristics include the local density peak position, voxel connectivity characteristics, Euler number, and curvature characteristics. The spatial characteristics are used as image texture features.

[0111] Methods for preliminarily analyzing depth points through density clustering algorithms, identifying different density areas, and dividing them using voxel units include:

[0112] A density clustering algorithm based on a preset first neighborhood radius and a preset first minimum number of points traverses the depth point set, counts the number of points within the preset first neighborhood radius of each depth point, and compares the number of points with the preset first minimum number of points. When the number of points is higher than the preset first minimum number of points, the corresponding depth point is marked as a core point, and a new cluster is created. The points within the preset first neighborhood radius of the core point are added to the cluster in turn, and the above steps are repeated until all points are marked as belonging to a certain cluster;

[0113] According to the range and resolution of the depth points, the three-dimensional space is divided into voxel units, and the depth points are mapped to the corresponding voxel units;

[0114] Methods for dividing the 3D space into voxel units based on the depth point range and resolution include:

[0115] Establish an XYZ-O three-dimensional rectangular coordinate system with the acquisition device as the origin;

[0116] Traverse the depth points and obtain the minimum value X of the depth point in the X-axis, Y-axis and Z-axis directions respectively. min , Y min and Z min ; and the maximum value X of the depth point in the X-axis, Y-axis and Z-axis directions max , Y max and Z max ;

[0117] Preset voxel resolutions ΔX, ΔY, and ΔZ in the X, Y, and Z directions;

[0118] Calculate the number of voxel units along the X-axis, Y-axis, and Z-axis respectively:

[0119]

[0120] Among them, N X N is the number of voxel units in the X-axis direction; Y N is the number of voxel units in the Y-axis direction; Z is the number of voxel units in the Z-axis direction;

[0121] From the starting position in the depth point range, the voxel units are gradually divided according to the resolution;

[0122] Each voxel unit is identified by the index (i, j, k) in the three axes, where i corresponds to the index in the X-axis direction; j corresponds to the index in the Y-axis direction; k corresponds to the index in the Z-axis direction; the range in the X-axis direction is [X min +i·ΔX,X min +(i+1)·ΔX]; the range in the Y-axis direction is [Y min +i·ΔY,Ymin +(i+1)·ΔY]; the range in the Z-axis direction is [Z min +i·ΔZ,Z min +(i+1)·ΔZ].

[0123] Methods for mapping depth points to corresponding voxel units include:

[0124] For each depth point (x, y, z), the corresponding voxel index (i′, j′, k′) is calculated by the following formula:

[0125]

[0126] Among them, size is the side length of the voxel; is rounded down; i′ is the voxel index of the x-axis corresponding to the depth point (x, y, z); j′ is the voxel index of the y-axis corresponding to the depth point (x, y, z); k′ is the voxel index of the z-axis corresponding to the depth point (x, y, z); in this way, each depth point can be mapped to a unique voxel unit.

[0127] The density clustering algorithm can distinguish the high-density area where the object is located from the low-density area of ​​the background based on the distribution density of the depth points. In many practical scenes, the object and the background have obvious differences in depth information. For example, in a scene containing a table and an object placed on the table, the depth point density of the object part is usually high, while the background (such as the surrounding air, the wall in the distance) has a low depth point density. Through this separation, the area that may contain the object can be quickly located, reducing the search range of subsequent recognition and improving recognition efficiency. For complex objects, their interior may be composed of multiple parts with different densities. By identifying different density areas through the density clustering algorithm, these parts can be preliminarily divided. For example, in the recognition of mechanical parts, a complex assembly may contain parts of different materials or structures, which will differ in the density of depth points. Identifying these different density areas helps the subsequent individual recognition of each part and the overall understanding of the entire complex object, providing a basis for fine object recognition.

[0128] Dividing depth points into voxel units can quantify the continuous three-dimensional space and convert it into a structured representation that is easier to process and analyze. This structured voxel representation is similar to the concept of pixels in three-dimensional images and can describe the spatial distribution of objects in a regular way. For example, in the field of medical imaging, the three-dimensional model of human organs can be represented by voxel units, and each voxel unit can contain relevant information about organ tissues, which enables computers to more conveniently process information such as the shape and position of the organs, which helps to identify organs in medical images. After the voxel units are divided, it is convenient to extract and calculate various spatial features based on each unit or unit combination. For example, the distribution of depth information in each voxel unit can be counted, and then the local density peak position, voxel connectivity and other spatial features of the object can be calculated. These features are very important for object recognition. For example, when identifying a building, the structural integrity of the building can be judged by analyzing the connectivity of the voxel units, and the position of the main structure of the building (such as load-bearing walls, etc.) can be determined by the local density peak position, thereby achieving accurate recognition of the building type and structure.

[0129] Methods for obtaining the local density peak position include:

[0130] The density peak position is determined by calculating the local density maximum:

[0131]

[0132] Wherein, ρ(n) is the local density value of depth point n; K is the number of depth points within the neighborhood radius of depth point n; s(n,n(q)) is the distance from depth point n to the qth depth point n(q) within its neighborhood radius; the local density values ​​of each depth point in each voxel unit are compared, and the depth point coordinates corresponding to the maximum local density value in each voxel unit are selected as the density peak position.

[0133] The local density peak position can help determine the key parts of an object with higher density. By locating these key parts, the structure and composition of the object can be understood more accurately, and then matched with the key structures of the known object model to achieve object recognition. For objects with similar shapes but different internal structures, the local density peak position has a good ability to distinguish. For example, a solid sphere and a sphere with a cavity inside will have different distributions of density peak positions. The density peak of a solid sphere may be at the center of the sphere and distributed more evenly, while the density peak of a sphere with a cavity is more obvious near the outer surface. In the process of object recognition, this difference in density peak position can help distinguish different object shapes.

[0134] Methods for obtaining voxel connectivity features include:

[0135] Define the connectivity rules between voxel units (preferably 6-connected, that is, in three-dimensional space, a voxel is adjacent to voxels in six directions: up, down, left, right, front and back), randomly select a starting voxel unit, use a depth-first search or breadth-first search algorithm to traverse all voxel units, mark the visited voxel units, and record the connectivity paths between voxel units.

[0136] Use a connectivity analysis-based method to identify holes; mark unoccupied voxel units as background, and use a connected component analysis algorithm to find connected areas in background voxel units; distinguish between internal holes and external holes; check the connection relationship between holes and object boundary voxels. If the hole is completely surrounded by object voxel units, it is an internal hole; if it is connected to the external space, it is an external hole. Count the number of connected paths and the number of voxel units contained in each path and splice them as voxel connectivity features.

[0137] Voxel connectivity features can be used to determine whether an object is complete. In three-dimensional space, if the various parts of an object are connected, then they should also show a certain degree of connectivity at the voxel level. For example, when identifying a broken ceramic vase and a complete ceramic vase, the difference between them can be easily found by analyzing voxel connectivity. The voxels of the complete vase are highly connected, while the voxel connectivity of the broken vase is destroyed at the broken part. This feature provides a basis for judging the integrity of the object and helps to accurately identify the state of the object. By studying voxel connectivity, we can also gain a deeper understanding of the internal structure of the object. For objects with complex internal channels or multiple independent substructures, such as engine blocks or honeycomb structures, voxel connectivity can clearly show the connectivity of these internal structures. According to different connectivity patterns, different types of engine blocks or honeycomb structures can be distinguished, thereby assisting object recognition, especially for those objects with similar appearances but different internal structures.

[0138] Methods for obtaining Euler numbers include:

[0139] On the basis of voxel unit representation, the number of vertices V is obtained by counting the connecting corners between voxel units; the number of edges E′ is obtained by counting the shared edges between adjacent voxel units; the number of faces F is obtained by counting the planes formed by the voxel units;

[0140] The Euler number E of an object is calculated by the following formula:

[0141] E = VE′ + FC;

[0142] Where C is the number of cavities inside the object.

[0143] The Euler number is a topological invariant that can concisely describe the topological structure of an object. For example, for simple polyhedrons, the Euler number can help distinguish objects of different shapes; in object recognition, by calculating and comparing the Euler number, objects with different topological structures can be quickly distinguished, which is very effective for identifying some objects with strange shapes or complex topological structures. Since the Euler number is a topological invariant, it is not affected by the geometric deformation of the object (such as non-tearing deformation such as stretching and bending). This makes it very robust in object recognition. For example, an object made of rubber may change its appearance size and shape after being squeezed or stretched, but the Euler number remains unchanged. Using this feature, objects can still be accurately identified even when they are deformed to a certain extent, improving the reliability of the object recognition system in complex environments.

[0144] Methods for obtaining the number of vertices V by counting the connected corner points between voxel units include:

[0145] Traverse each vertex of each voxel unit (8 vertices for 6-connected units), analyze the status of the current voxel unit and the adjacent voxel units, that is, whether the vertex is occupied (that is, whether it belongs to an object); and record the coordinates of the counted vertices through the data structure. Whenever a vertex is judged to be an object vertex, first check whether it is already in the data structure. If not, add one to the count and add it to the data structure; if it already exists, skip it to ensure the accuracy of vertex statistics; finally, calculate the count value output by the data structure as the vertex number V.

[0146] The method of counting the shared edges between adjacent voxel units to obtain the edge number E′ includes:

[0147] Traverse all voxel units in the three coordinate axis directions of the XYZ-O three-dimensional rectangular coordinate system and the shared edges with adjacent voxel units (6-connectivity corresponds to a maximum of 12 edges, which can be divided into three groups of edges corresponding to the three coordinate axis directions of the XYZ-O three-dimensional rectangular coordinate system). If at least one of the two voxel units where an edge is located is occupied (belongs to the object), then this edge belongs to the object, and the edge count is increased by one; record the coordinates of the two endpoints of the edge through a data structure (such as a hash table). Each time an edge is counted, first check whether it has been recorded. If not, count and mark it. If it already exists, skip it. Finally, count the count value output by the statistical data structure as the number of edges E′.

[0148] Reference Figure 2 , the method of obtaining the number of faces F by counting the planes formed by the voxel units includes:

[0149] The plane is divided into planes perpendicular to the three coordinate axes of the XYZ-O three-dimensional rectangular coordinate system, and all voxel units are traversed to check the plane formed by the voxel unit perpendicular to a certain coordinate axis and the adjacent voxel units on the other two axis planes. If multiple adjacent voxels are continuously occupied on this plane, a plane area belonging to the object is formed, which is counted as a face; the coordinates of any three non-collinear depth points in the plane are recorded through the data structure, and the normal vector of the plane is calculated (based on obtaining the coordinates of the three non-collinear depth points to obtain the cross vector in the two planes, and the normal vector of the plane is obtained through the cross product operation), and recorded through the data structure. Each time a face is counted, the plane vector composed of the depth coordinates of any two points in the plane is obtained, and it is judged whether the plane vector is perpendicular to the normal vector recorded in the data structure to determine whether it is a repeated plane. If not perpendicular, it is judged as a new plane, and the face count is increased by 1. If it is perpendicular, it is judged as a repeated plane and skipped. Finally, the count value output by the statistical data structure is counted as the face number F.

[0150] Methods for obtaining the curvature characteristics of an object include:

[0151] Select voxel point U, select W voxel points around it, select the parabola equation Z = A × X 2 +B×Y 2 +C×X+F×Y+G, where A, B, C, F and G are fitting coefficients; X, Y and Z are the coordinates of the voxel point in the XYZ-O three-dimensional rectangular coordinate system;

[0152] Substitute the coordinates of each selected voxel point into the above parabola equation, solve the equation group by the least square method, obtain the fitting coefficient that minimizes the sum of square errors, substitute the fitting coefficient into the parabola equation to obtain the fitted parabola equation; select J position points on the surface of the object, and calculate the curvature of the J position points according to the fitted parabola equation.

[0153] Curvature features play a key role in identifying the boundaries and contours of objects. The curvature at the boundaries of objects usually changes. By analyzing the size and change trend of the curvature, the boundaries of the objects can be accurately located. For example, when identifying an irregularly shaped leaf, the change in curvature at the edge of the leaf can help determine the outline shape of the leaf and distinguish it from other objects of different shapes. Moreover, for objects with complex contours, such as carved artworks, curvature features can help identify the fine contours of the carved parts, thereby achieving more accurate recognition of the object. Curvature can also reflect the detailed information of the shape of the object. For objects with similar shapes but different details, such as different styles of car fronts, their curvature distribution will be different. Some car fronts may have more rounded curves, while others may have sharper broken lines. By analyzing the curvature features, these differences in shape details can be captured, which helps to distinguish objects of different styles or models during object recognition and achieve more refined object recognition.

[0154] The image semantic features and image texture features are used as input to the object classification model to obtain the object recognition results.

[0155] The training methods for object classification models include:

[0156] H groups of recognition training data are collected in advance, and the recognition training data include image semantic features, image texture features and object recognition results.

[0157] Each group of recognition training data is used as the input of the object classification model, the object classification model uses the object recognition results corresponding to each group of image semantic features and image texture features as output, and uses the actual object recognition results corresponding to each group of image semantic features and image texture features as prediction targets; minimizing the prediction error of the object recognition results output by the object classification model each time is used as the training target; the object classification model is trained until the sum of the prediction errors reaches convergence, and the training is stopped; the object classification model is a deep neural network model.

[0158] Example 2

[0159] See also Figure 3 As shown, the object recognition system based on multimodal data fusion described in this embodiment includes:

[0160] Device synchronization module: establish a spatiotemporal synchronization model for acquisition devices; acquisition devices include image acquisition devices and depth acquisition devices;

[0161] Data acquisition module: Use the synchronized image acquisition device to acquire the two-dimensional image data of the object at different angles and under different lighting conditions, and use the synchronized depth acquisition device to scan and acquire the depth data at different distances;

[0162] Data processing module: denoising, enhancing and normalizing the two-dimensional image data of the object to obtain image data; filtering and smoothing the depth data to obtain depth points;

[0163] The first extraction module: takes the image data as the input of the first extraction model to obtain the image semantic features;

[0164] The second extraction module: obtains the depth data of the object surface corresponding to the image data, and performs feature extraction on the depth data to obtain image texture features;

[0165] Object recognition module: Use image semantic features and image texture features as input to the object classification model to obtain object recognition results.

[0166] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

[0167] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An object recognition method based on multimodal data fusion, characterized in that: The steps include: Establishing a spatiotemporal synchronization model for acquisition equipment; the acquisition equipment includes image acquisition equipment and depth acquisition equipment; The synchronized image acquisition device is used to acquire two-dimensional image data of the object at different angles and under different lighting conditions, and the synchronized depth acquisition device is used to scan and acquire depth data at different distances; Performing denoising, enhancement and normalization processing on the two-dimensional image data of the object to obtain image data; Filtering and smoothing the depth data to obtain depth points; Using the image data as input of the first extraction model to obtain image semantic features; Obtaining depth data of the surface of the object corresponding to the image data, and performing feature extraction on the depth data to obtain image texture features; The image semantic features and image texture features are used as input to the object classification model to obtain the object recognition results.

2. The object recognition method based on multimodal data fusion according to claim 1, characterized in that: The method for obtaining depth data corresponding to the surface of an object in the image data comprises: A nested loop including an outer loop and an inner loop is preset to traverse all pixel points corresponding to the surface of the object in the image data, wherein the outer loop is used to traverse each row of the image data, and the inner loop is used to traverse each column of the image data, and the depth points of the coordinates of all pixel points corresponding to the surface of the object in the XYZ-O three-dimensional rectangular coordinate system are obtained in sequence; The method for obtaining image texture features comprises: A density clustering algorithm is used to perform a preliminary analysis of the depth points, identify different density areas, and divide them using voxel units. The connectivity of the voxel units, the Euler number of the object, and the curvature of the object are analyzed to extract the spatial characteristics of the object; the spatial characteristics include the local density peak position, voxel connectivity characteristics, Euler number and curvature characteristics; the spatial characteristics are used as image texture features.

3. The object recognition method based on multimodal data fusion according to claim 2, characterized in that: The method of performing preliminary analysis on the depth points by using a density clustering algorithm, identifying different density areas, and dividing them using voxel units includes: A density clustering algorithm based on a preset first neighborhood radius and a preset first minimum number of points traverses the depth point set, counts the number of points within the preset first neighborhood radius of each depth point, and compares the number of points with the preset first minimum number of points. When the number of points is higher than the preset first minimum number of points, the corresponding depth point is marked as a core point, and a new cluster is created. The points within the preset first neighborhood radius of the core point are added to the cluster in turn, and the above steps are repeated until all points are marked as belonging to a certain cluster; According to the range and resolution of the depth points, the three-dimensional space is divided into voxel units, and the depth points are mapped to the corresponding voxel units.

4. The object recognition method based on multimodal data fusion according to claim 3, characterized in that: The method of dividing the three-dimensional space into voxel units according to the depth point range and resolution includes: Traverse the depth points and obtain the minimum value X of the depth point in the X-axis, Y-axis and Z-axis directions of the XYZ-O three-dimensional rectangular coordinate system established with the acquisition device as the origin. min , Y min and Z min ; and the maximum value X of the depth point in the X-axis, Y-axis and Z-axis directions max , Y max and Z max ; Preset voxel resolutions ΔX, ΔY, and ΔZ in the X, Y, and Z directions; The number of voxel units in the axial direction of the X-axis, Y-axis and Z-axis is obtained by respectively calculating the ratio of the difference between the maximum value and the minimum value in the X-axis, Y-axis and Z-axis directions of the XYZ-O three-dimensional rectangular coordinate system to the resolution; From the starting position in the depth point range, the voxel units are gradually divided according to the resolution; Each voxel unit is identified by the index (i, j, k) in the three axes, where i corresponds to the index in the X-axis direction; j corresponds to the index in the Y-axis direction; k corresponds to the index in the Z-axis direction; the range in the X-axis direction is [X min +i·ΔX,X min +(i+1)·ΔX]; the range in the Y-axis direction is [Y min +i·ΔY,Y min +(i+1)·ΔY]; the range in the Z-axis direction is [Z min +i·ΔZ,Z min +(i+1)·ΔZ].

5. The object recognition method based on multimodal data fusion according to claim 4, characterized in that: The method of mapping depth points to corresponding voxel units comprises: For each depth point (x, y, z), the corresponding voxel index (i′, j′, k′) is calculated by the ratio of the coordinate value of the depth point in the three axes to the side length of the voxel; The method for obtaining the local density peak position includes: The local density value is calculated by calculating the distance relationship between the depth point and the points in the neighborhood corresponding to the depth point, and the maximum local density value is selected as the density peak position; Methods for obtaining voxel connectivity features include: Define the connectivity rules between voxel units, randomly select a starting voxel unit, traverse all voxel units using a depth-first search or breadth-first search algorithm, mark the visited voxel units, and record the connectivity paths between voxel units; identify the number of holes in the object based on the connectivity path record results; count the number of connectivity paths and the number of voxel units contained in each connectivity path, and splice them as voxel connectivity features; Methods for obtaining Euler numbers include: The Euler number is calculated by counting the number of vertices, edges, faces and internal cavities of the voxel unit. Methods for obtaining the curvature characteristics of an object include: Select a voxel point U, select V voxel points around the voxel point U, and select a parabola equation; Substitute the coordinates of each selected voxel point into the above parabola equation, solve the equation group by the least square method, obtain the fitting coefficient that minimizes the sum of square errors, substitute the fitting coefficient into the parabola equation to obtain the fitted parabola equation; select J position points on the surface of the object, and calculate the curvature of the J position points according to the fitted parabola equation.

6. The object recognition method based on multimodal data fusion according to claim 5, characterized in that: Methods for obtaining the number of vertices V by counting the connected corner points between voxel units include: Traverse each vertex of each voxel unit and analyze whether the vertex between the current voxel unit and the adjacent voxel unit is occupied; and record the coordinates of the counted vertices through the data structure. Whenever a vertex is determined to be an object vertex, traverse the data structure to see whether it exists. If not, add one to the count and add it to the data structure; if it already exists, skip it; The count value output by the statistical data structure is used as the vertex number V; The method of counting the shared edges between adjacent voxel units to obtain the edge number E′ includes: Traverse all voxel units in the three coordinate axis directions of the XYZ-O three-dimensional rectangular coordinate system and the shared edges with adjacent voxel units; if at least one of the two voxel units where an edge is located is occupied, then this edge belongs to the edge of the object; record the coordinates of the two endpoints of the edge through a data structure, and each time an edge is counted, traverse whether it exists in the data structure, if not, add one to the count and add it to the data structure; if it already exists, skip it; count the count value output by the statistical data structure as the number of edges E′; Methods for obtaining the number of faces F by counting the planes formed by voxel units include: The plane is divided into planes perpendicular to the three coordinate axes of the XYZ-O three-dimensional rectangular coordinate system, and all voxel units are traversed to check whether the voxel unit perpendicular to any coordinate axis forms a plane with the adjacent voxel units on the other two coordinate axis planes. If θ adjacent voxels are continuously occupied on the plane, a plane area belonging to the object is formed, and the plane area is counted as a face; the coordinates of any three non-collinear depth points in the plane are recorded through a data structure, and the corresponding normal vector is calculated and recorded. Each time a face is counted, the plane vector composed of the depth coordinates of any two points in the plane is obtained, and it is determined whether the plane vector is perpendicular to the normal vector recorded in the data structure. If not, it is determined to be a new plane, and the face count is increased by 1. If it is perpendicular, it is determined to be a repeated plane and it is skipped; the count value output by the statistical data structure is used as the face number F.

7. The object recognition method based on multimodal data fusion according to claim 1, characterized in that: The method of establishing a spatiotemporal synchronization model for acquisition equipment includes: Step 1: Obtain the coordinate system of each acquisition device; Step 2, define a unified world coordinate system; the unified world coordinate system is to establish an XYZ-O three-dimensional rectangular coordinate system with the acquisition device as the origin; Step 3: Use the rotation matrix R and the translation vector t to represent the device's posture. The rotation matrix R is used to describe the rotation relationship between the device coordinate system and the world coordinate system. The rotation matrix R satisfies R T R = I, T is the transpose of the vector, and I is the unit matrix; the translation vector t is used to represent the position of the origin of the device coordinate system in the world coordinate system; Step 4: Express the coordinates of the data point Q collected by the acquisition device as homogeneous coordinates Then use the matrix transformation formula to transform the homogeneous coordinates Convert to the world coordinate system to obtain the world coordinate of data point Q.

8. The object recognition method based on multimodal data fusion according to claim 7, characterized in that: The method of expressing the coordinates of the data points collected by the collection device as homogeneous coordinates comprises: For image acquisition devices: The internal parameters of the image acquisition device camera are obtained through the camera calibration method, including the x-axis focal length f x , y-axis focal length f y and the principal point coordinates (p x ,p y ); For a pixel point (u, v) in the image data, where u and v are the pixel coordinates, the coordinates (x k ,y k ,z k ) is converted to homogeneous coordinates Among them, x k is the x-axis coordinate of the pixel point (u, v) in the camera coordinate system; y k is the y-axis coordinate of the pixel point (u, v) in the camera coordinate system; z k is the z-axis coordinate of the pixel point (u, v) in the camera coordinate system; is the homogeneous coordinate of the pixel point (u, v) after the x-axis coordinate in the camera coordinate system is transformed; is the homogeneous coordinate of the pixel point (u, v) in the camera coordinate system after the y-axis coordinate is transformed; is the homogeneous coordinate of the pixel point (u, v) in the camera coordinate system after the z-axis coordinate is transformed; For the depth acquisition device: the coordinates (a, b, c) of the data points of the depth acquisition device are expressed as homogeneous coordinates (a, b, c, 1); where a is the x-axis coordinate in the three-dimensional coordinate system of the depth acquisition device; b is the y-axis coordinate in the three-dimensional coordinate system of the depth acquisition device; and c is the z-axis coordinate in the three-dimensional coordinate system of the depth acquisition device.

9. The object recognition method based on multimodal data fusion according to claim 6, characterized in that: The method for obtaining the two-dimensional image data of the object at different angles and different lighting conditions and the depth data at different distances includes: Determine the angle range of the collected objects and plan according to the preset angle intervals; Under the preset environment, simulate using artificial light sources of different types and intensities according to the preset lighting conditions; According to the size of the object and the application scenario requirements, the distance range between the object and the collection device is divided, and the distance is planned according to the preset distance interval; Build a closed collection space around the object, arrange lights around and on top of the object, and install shading facilities; The two-dimensional image data of the object is collected according to the preset angle interval and the preset lighting conditions, and the depth data of the object is collected according to the preset distance interval.

10. An object recognition system based on multimodal data fusion, used to implement the object recognition method based on multimodal data fusion according to any one of claims 1 to 9, characterized in that: include: Device synchronization module: used to establish a spatiotemporal synchronization model for acquisition devices; acquisition devices include image acquisition devices and depth acquisition devices; Data acquisition module: Use the synchronized image acquisition device to acquire the two-dimensional image data of the object at different angles and under different lighting conditions, and use the synchronized depth acquisition device to scan and acquire the depth data at different distances; Data processing module: used to perform denoising, enhancement and normalization processing on the two-dimensional image data of the object to obtain image data; Filtering and smoothing the depth data to obtain depth points; The first extraction module is used to use the image data as the input of the first extraction model to obtain the image semantic features; The second extraction module: obtains the depth data of the object surface corresponding to the image data, and performs feature extraction on the depth data to obtain image texture features; Object recognition module: Use image semantic features and image texture features as input to the object classification model to obtain object recognition results.

Citation Information

Patent Citations

  • Training method and device for transparent object recognition, storage medium and terminal

    CN109903323A

Cited By

  • Target detection method and system based on machine vision

    CN121121081A