Real-time reasoning method for physical attributes of object

By combining multimodal non-contact sensors and deep neural networks, real-time and accurate reasoning of the physical properties of objects is achieved, solving the problems of insufficient real-time performance and security in existing technologies and expanding the application scenarios of intelligent robots.

CN121390282APending Publication Date: 2026-01-23HANGZHOU ZHIYUE QIANXING TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511442139.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing technologies lack real-time performance, security, and automation when acquiring the physical properties of objects, and cannot meet the needs of embodied intelligence for autonomous interaction in complex and dynamic environments.

Method used

The system uses multimodal non-contact sensors to collect raw perception data of objects, preprocesses and initially fuses the data through a cloud processor, extracts multi-dimensional physical feature vectors, and uses deep neural networks to perform multi-layer nonlinear transformations and feature mappings to build a physical attribute database, which is then output to the decision control system of the embodied intelligent agent.

Benefits of technology

It enables real-time and accurate reasoning of the physical properties of objects, securely acquires multi-dimensional information, expands the application areas and market potential of intelligent robots, and improves the level of intelligence and the smoothness of human-machine collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121390282A_ABST
    Figure CN121390282A_ABST
Patent Text Reader

Abstract

The invention provides a real-time reasoning method for physical attributes of an object, and relates to the technical field of data processing, and the method comprises the steps: collecting original sensing data of a target object through a multi-mode non-contact sensor, carrying out the preprocessing of the original sensing data of the target object, obtaining multi-source heterogeneous sensing data, and carrying out the real-time reasoning of the physical attributes of the object; performing time synchronization, space alignment and data cleaning on the multi-source heterogeneous sensing data to obtain pre-processed data, and extracting visual features, geometric features and thermodynamic features associated with physical attributes of an object from the pre-processed data to obtain a multi-dimensional physical feature vector; through collaborative design of multi-mode non-contact perception, data standardization processing, multi-dimensional feature extraction, deep neural network reasoning and database output, real-time inference of physical attributes of an object is realized, and reliable data is provided for intelligent decision control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a real-time reasoning method for the physical properties of objects. Background Technology

[0002] Currently, the technology for obtaining the physical properties of objects mainly relies on contact measurement or offline analysis. The specific implementation schemes and their disadvantages are as follows: Contact sensor measurement: This method uses torque sensors, tactile sensors, etc., to directly contact the object and measure its stiffness, friction, etc., through physical feedback. The disadvantages of this method are that it requires physical contact with the object, making it unsuitable for high-temperature, sterile, fragile, or distant objects; additionally, the sensor itself may interfere with the object's state, and it typically only measures a single-dimensional attribute. Manual judgment and input: Operators pre-set or manually measure the object's physical parameters based on experience and then input them into the system. This method is inefficient, heavily reliant on subjective experience, cannot achieve automation or real-time response, and is completely ineffective for unknown or entirely new objects. Offline simulation and analysis: After obtaining a precise 3D model of the object using equipment such as a 3D scanner, physical simulation analysis is performed on a computer to deduce its physical properties. While this method can obtain relatively accurate data, it is complex, requires expensive equipment, and is extremely time-consuming, completely failing to meet the needs of embodied intelligent agents for real-time decision-making and interaction with the environment.

[0003] Existing technologies have significant shortcomings in terms of real-time performance, security, applicability, and automation, and cannot meet the needs of modern embodied intelligence for autonomous and intelligent interaction in complex and dynamic environments. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a real-time reasoning method for the physical properties of objects. Through the collaborative design of multimodal non-contact sensing, data standardization processing, multi-dimensional feature extraction, deep neural network reasoning and database output, the method realizes real-time inference of the physical properties of objects, providing reliable data for embodied intelligence decision control.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a real-time reasoning method for the physical properties of an object, the method comprising: Step 1: Collect raw sensing data of the target object using a multimodal non-contact sensor, and preprocess the raw sensing data of the target object to obtain multi-source heterogeneous sensing data. Step 2: Perform time synchronization, spatial alignment, and data cleaning on the multi-source heterogeneous sensing data to obtain preprocessed data; Step 3: Extract visual features, geometric features, and thermodynamic features associated with the physical properties of the object from the preprocessed data to obtain a multi-dimensional physical feature vector; Step 4: Perform multi-layer nonlinear transformation and feature mapping on the multi-dimensional physical feature vector through a deep neural network. Based on the correlation between appearance features and physical attributes, infer a physical attribute of the target object and its corresponding confidence level. Step 5: Based on a physical attribute of the target object and its confidence level, construct a physical attribute database, obtain the corresponding confidence score, and output the physical attribute database and confidence score to the decision control system of the embodied intelligent agent to realize real-time reasoning on the physical attributes of the object.

[0006] Furthermore, raw sensing data of the target object is acquired through multimodal non-contact sensors, and this raw sensing data is preprocessed to obtain multi-source heterogeneous sensing data, including: Step 1.1: Acquire raw sensing data of the target object using the multimodal non-contact sensor; Step 1.2: Based on the collected raw sensing data, preprocessing is performed by cloud processor to remove noise and enhance the data, resulting in improved sensing data. Step 1.3: The improved multi-source sensing data is initially fused by the cloud processor to obtain multi-source heterogeneous sensing data.

[0007] Furthermore, the multi-source heterogeneous sensing data undergoes time synchronization, spatial alignment, and data cleaning to obtain preprocessed data, including: Step 2.1: Based on the obtained multi-source heterogeneous sensing data, time-synchronized sensing data is obtained by synchronizing the data streams from different sensors. Step 2.2: The obtained time-synchronized sensing data is processed by spatial alignment to unify the spatial coordinate system of each sensor data, thus obtaining spatially aligned sensing data. Step 2.3: The obtained spatially aligned perception data is cleaned to remove outliers and noise, resulting in preprocessed data.

[0008] Furthermore, visual, geometric, and thermodynamic features associated with the physical properties of objects are extracted from the preprocessed data to obtain a multi-dimensional physical feature vector, including: Step 3.1: Based on the obtained standardized perception data, extract visual features associated with the physical properties of the object; Step 3.2: Further extract geometric features associated with the physical properties of objects from the obtained visual features and standardized perceptual data; Step 3.3: Further extract thermodynamic features associated with the physical properties of the object from the obtained geometric features and the standardized sensing data; Step 3.4: The extracted visual features, geometric features, and thermodynamic features are fused and vectorized to obtain a multi-dimensional physical feature vector.

[0009] Furthermore, by performing multi-layer nonlinear transformations and feature mappings on multi-dimensional physical feature vectors using deep neural networks, and based on the correlation between appearance features and physical attributes, a physical attribute of the target object and its corresponding confidence level are inferred, including: Step 4.1: Based on the obtained multi-dimensional physical feature vectors, perform multi-layer nonlinear transformations through the deployed deep neural network to obtain high-dimensional abstract features; Step 4.2: For the obtained high-dimensional abstract features, feature mapping is performed through a deep neural network. Based on the pre-learned correlation between appearance features and physical properties, preliminary inference results of physical properties are obtained. Step 4.3: Calculate the confidence level of the preliminary inference results of the obtained physical properties using a deep neural network, and output a physical property of the target object and its corresponding confidence level.

[0010] Furthermore, based on a physical attribute of the target object and its confidence level, a physical attribute database is constructed, and the corresponding confidence score is obtained, including: Step 5.1: Based on a physical attribute of the target object output by the deep neural network and its corresponding confidence level, an initial physical attribute database is obtained through structured storage; Step 5.2: For the initial physical attribute database, the physical attributes are bound to their confidence scores through confidence score calculation and association to obtain a standardized physical attribute database; Step 5.3: Through data indexing and query optimization, the standardized physical attribute database is used to obtain the final physical attribute database and corresponding confidence scores that can be used for real-time interaction.

[0011] Furthermore, a real-time reasoning method for the physical properties of objects is characterized by outputting a physical property database and confidence scores to the decision control system of an embodied intelligent agent to achieve real-time reasoning for the physical properties of objects, including: Step 5.4: Based on the final physical attribute database that can be used for real-time interaction and the corresponding confidence scores, the data is encapsulated and formatted for transmission through the cloud processor to obtain a standardized data stream; Step 5.5: The standardized data stream is transmitted through the real-time communication link established between the embodied intelligent agent decision control system. Step 5.6: For the data successfully transmitted to the decision control system, through parsing and application, drive the embodied intelligent agent to perform real-time interactive tasks that match the physical properties of the target object, and complete the real-time inference closed loop of physical properties.

[0012] Secondly, a real-time reasoning system for the physical properties of objects includes: The multimodal non-contact sensing module is used to collect raw sensing data of the target object through multimodal non-contact sensors, and to preprocess the raw sensing data of the target object to obtain multi-source heterogeneous sensing data. The data preprocessing and feature fusion module is used to perform time synchronization, spatial alignment and data cleaning on multi-source heterogeneous sensing data to obtain preprocessed data. Visual features, geometric features and thermodynamic features associated with the physical properties of objects are extracted from the preprocessed data to obtain multi-dimensional physical feature vectors. The physical attribute real-time inference engine module is used to perform multi-layer nonlinear transformation and feature mapping on multi-dimensional physical feature vectors through deep neural networks, and infer a physical attribute of the target object and its corresponding confidence level based on the correlation between appearance features and physical attributes. The physical attribute database and output module is used to construct a physical attribute database based on a physical attribute of the target object and its confidence level, obtain the corresponding confidence score, and output the physical attribute database and confidence score to the decision control system of the embodied intelligent agent to realize real-time reasoning on the physical attributes of the object.

[0013] Thirdly, a computing device includes: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0014] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0015] The above-described solution of the present invention has at least the following beneficial effects: This solution utilizes deep learning models for real-time reasoning, reducing the offline analysis process that previously took minutes or even hours to milliseconds. It achieves instantaneous acquisition of physical properties, enabling embodied intelligent agents to abandon clumsy "trial and error" interactions and gain a full understanding of the object's physical nature before operation. This allows them to plan the most efficient and safest action strategies, resulting in a qualitative leap in intelligence and operational efficiency. Since no physical contact is required, this technology can be safely applied to various scenarios where traditional contact sensors are inadequate. These applications include operating sterile instruments in the medical field, handling high-temperature or chemically corrosive materials in the industrial field, digitizing fragile cultural relics in the cultural heritage field, and interacting with diverse unknown objects in home services. This greatly expands the application areas and market potential of intelligent robots.

[0016] This solution makes machine behavior more consistent with physical laws and human intuition. When a robot can, like a human, determine from observation that a cup is made of glass and filled with water, and thus pick it up with a smooth and gentle motion, its behavior is predictable, reliable, and natural to human observers or collaborators, improving the smoothness of human-machine collaboration and the user's interactive experience. This invention is a deep cross-application of multiple cutting-edge fields such as computer vision, deep learning, sensor technology, and robotics. The need for precise reasoning about physical properties will, in turn, drive the development of high-precision multimodal sensors, the design of more efficient and lightweight neural network models, and the innovation of more robust robot control theories, thereby driving the collaborative progress of the entire technological ecosystem. This solution mainly relies on versatile and relatively low-cost vision and depth sensors, and achieves the effects that previously required a combination of multiple expensive and dedicated physical sensors through advanced software algorithms. The core capability of the system lies in its iterative and updatable AI model, which means that adapting to new objects or improving performance only requires retraining the model without replacing hardware. The system's flexibility, scalability, and economy are all superior to traditional solutions. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a real-time reasoning method for physical properties of an object, provided by an embodiment of the present invention.

[0018] Figure 2 This invention presents a real-time reasoning flowchart for the physical properties of an object.

[0019] Figure 3 A schematic diagram of the traditional separate physical property acquisition method. Detailed Implementation

[0020] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0021] like Figure 1 As shown, an embodiment of the present invention proposes a real-time reasoning method for the physical properties of an object, the method comprising the following steps: Step 1: Collect raw sensing data of the target object using a multimodal non-contact sensor, and preprocess the raw sensing data of the target object to obtain multi-source heterogeneous sensing data. Step 2: Perform time synchronization, spatial alignment, and data cleaning on the multi-source heterogeneous sensing data to obtain preprocessed data; Step 3: Extract visual features, geometric features, and thermodynamic features associated with the physical properties of the object from the preprocessed data to obtain a multi-dimensional physical feature vector; Step 4: Perform multi-layer nonlinear transformation and feature mapping on the multi-dimensional physical feature vector through a deep neural network. Based on the correlation between appearance features and physical attributes, infer a physical attribute of the target object and its corresponding confidence level. Step 5: Based on a physical attribute of the target object and its confidence level, construct a physical attribute database, obtain the corresponding confidence score, and output the physical attribute database and confidence score to the decision control system of the embodied intelligent agent to realize real-time reasoning on the physical attributes of the object.

[0022] In this embodiment of the invention, the invention employs a multimodal non-contact sensor to collect raw perception data of the target object, which is then preprocessed to obtain multi-source heterogeneous perception data. This multi-source heterogeneous perception data is then time-synchronized, spatially aligned, and cleaned to obtain standardized perception data. Subsequently, visual, geometric, and thermodynamic features are extracted from the preprocessed data to form a multi-dimensional physical feature vector. This feature vector is then subjected to multi-layer nonlinear transformations and feature mapping using a deep neural network to infer the physical properties of the target object and their corresponding confidence levels in real time. Finally, structured physical attribute data is generated based on the attributes and confidence levels and output to the embodied intelligent agent for decision-making and control. This technological approach overcomes the risks of object damage and sensor interference associated with existing contact measurement methods, the inability of manual judgment or offline analysis to meet real-time requirements, the limitations of single feature extraction in simultaneously acquiring multi-dimensional physical attributes of objects, and the disconnect between perceived data and decision-making control by embodied intelligent agents. It achieves the technical effects of safely collecting data without physical contact, accurately inferring multi-dimensional physical attributes of objects in real time, and providing standardized structured data support for embodied intelligent agents to achieve efficient real-time action planning and control. This effectively adapts to the interactive needs of embodied intelligence in dynamic and complex scenarios, improving operational safety and intelligence levels.

[0023] In a preferred embodiment of the present invention, step 1 above may include: Step 1.1 involves acquiring raw perception data of the target object using the multimodal non-contact sensors. Specifically, this includes configuring a multimodal non-contact sensor combination based on the application scenario of the embodied intelligent agent. This combination includes a depth camera, a high-resolution RGB camera, and a thermal imaging camera. The sensors are deployed within a safe range of 0.5 to 5 meters from the target object to avoid physical contact. After activating all sensors, the depth camera acquires the three-dimensional geometric shape data and depth information of the target object in real time. The high-resolution RGB camera simultaneously captures the surface color data and texture detail data of the target object, while the thermal imaging camera acquires the surface temperature distribution data of the target object. The raw data acquired by the three types of sensors is transmitted to the data receiving end in real time, completing the acquisition of the raw perception data of the target object.

[0024] Step 1.2: Based on the collected raw sensing data, the cloud processor performs data denoising and data enhancement preprocessing to obtain improved sensing data. Specifically, this includes: uploading the collected raw sensing data to the cloud processor, which first performs targeted denoising processing for different types of raw data: for flying point noise that may exist in depth data, a Gaussian filtering algorithm is used for smoothing and removing invalid data points; for salt-and-pepper noise that may exist in RGB image data, a median filtering algorithm is used to remove it and restore the clear texture of the image; for temperature fluctuation noise that may exist in thermal imaging data, a moving average filtering algorithm is used for stabilization to ensure the accuracy of temperature data. After denoising, the cloud processor then performs data enhancement processing: for RGB images, a histogram equalization algorithm is used to improve image contrast and highlight the texture details and color differences on the glass shell surface; for depth data, a data completion algorithm is used to fill in local data gaps caused by occlusion to ensure the integrity of the three-dimensional geometric shape data; for thermal imaging data, a temperature range stretching algorithm is used to enhance the identification of temperature differences in different areas, ultimately obtaining improved sensing data.

[0025] Step 1.3 involves the cloud processor performing preliminary fusion of the improved multi-source sensing data to obtain multi-source heterogeneous sensing data. Specifically, after acquiring the improved multi-source sensing data, the cloud processor first correlates and matches sensing data collected by different sensors at the same time based on the sensor's acquisition timestamp. This ensures that each set of data corresponds to the state of the target object at the same point in time, avoiding data misalignment due to acquisition time differences. Then, the cloud processor uses a data association mapping algorithm to correlate the 3D coordinate information in the depth data with the pixel coordinate information in the RGB image data, establishing a spatial correlation between geometric shape and visual appearance. Simultaneously, the temperature distribution information in the thermal imaging data is bound to the 3D position information in the depth data, clarifying the temperature attributes corresponding to different spatial positions of the target object. Through this correlation processing, the originally independent depth data, RGB image data, and thermal imaging data are integrated into a set containing multi-dimensional information about the target object's geometry, vision, and thermodynamics, completing the preliminary fusion of the multi-source sensing data and obtaining multi-source heterogeneous sensing data.

[0026] In this embodiment of the invention, the present invention overcomes the technical problems of existing contact measurement methods, such as easy damage to the target object, noise or incomplete information in the original sensing data affecting the accuracy of subsequent analysis, and isolated and unrelated multi-source sensor data that cannot form a comprehensive understanding of the object. It adopts a technical means of acquiring original sensing data of the target object with multi-modal non-contact sensors, preprocessing the original sensing data with noise and data enhancement based on cloud processor, and finally performing preliminary fusion of the improved multi-source sensing data by cloud processor to generate multi-source heterogeneous sensing data.

[0027] In a preferred embodiment of the present invention, step 2 above may include: Step 2.1: Based on the obtained multi-source heterogeneous sensing data, time-synchronized sensing data is obtained by synchronizing the data streams from different sensors. Specifically, this includes: first, acquiring the obtained multi-source heterogeneous sensing data, which comes from depth cameras, high-resolution RGB cameras, and thermal imaging cameras in the scenario of sorting micro-sensor components with precision glass shells; first, extracting the acquisition timestamp of each data frame in each sensor data stream, where the depth camera generates timestamps at a frequency of 30 frames per second, the RGB camera generates timestamps at a frequency of 25 frames per second, and the thermal imaging camera generates timestamps at a frequency of 20 frames per second; then, calculating the acquisition latency between different sensors, for example, comparing the data frame timestamps of the depth camera and the RGB camera near the same moment, and finding that the average latency of the two is about 10 milliseconds, and the average latency of the depth camera and the thermal imaging camera is about 15 milliseconds. Subsequently, using the timestamp of the depth camera as a reference, time compensation is performed on the data streams of the RGB camera and the thermal imaging camera. The data frames of the RGB camera are shifted forward or backward by the corresponding time according to the time delay, and the data frames of the thermal imaging camera are done in the same way. This adjusts the data frames with similar timestamps in the three sensors to data at the same moment under a unified time reference. Finally, frame matching is performed on the adjusted data stream to ensure that each set of data frames corresponds to the state of the micro-sensor element in the precision glass shell at the same moment. This avoids the problem of mismatch between the geometry of the glass shell and the surface texture and temperature distribution in subsequent analysis due to time misalignment, and finally obtains the time-synchronized sensing data.

[0028] Step 2.2: The obtained time-synchronized perception data is processed through spatial alignment to unify the spatial coordinate system of each sensor data, resulting in spatially aligned perception data. Specifically, based on the obtained time-synchronized perception data, a unified spatial coordinate system is first determined. The three-dimensional spatial coordinate system of the depth camera is selected as the reference coordinate system. This coordinate system has the center of the depth camera lens as the origin, the horizontal rightward axis as the X-axis, the vertical upward axis as the Y-axis, and the vertical outward axis as the Z-axis. It has been pre-calibrated according to the location of the sorting station in the workshop to ensure that it can accurately reflect the actual spatial position of the precision glass shell at the sorting station. Next, the relative position parameters of the depth camera, RGB camera, and thermal imaging camera during installation are obtained, including the distance and angular offset between each camera. These parameters are obtained through the initial equipment installation and calibration in the workshop. For example, the RGB camera is installed 10 cm to the right of the depth camera, at a 5-degree angle to the depth camera lens plane, and the thermal imaging camera is installed 15 cm below the depth camera, at a 3-degree angle to the depth camera lens plane. Then, based on these relative position parameters, a coordinate transformation algorithm is used to map the two-dimensional pixel coordinate data collected by the RGB camera to a reference three-dimensional coordinate system, and the three-dimensional spatial coordinates corresponding to each pixel are calculated. At the same time, the pixel positions corresponding to the temperature data collected by the thermal imaging camera are also transformed to the reference three-dimensional coordinate system according to the relative position parameters, so that each temperature value in the thermal imaging data can correspond to the specific position of the precision glass shell in three-dimensional space. Finally, all the transformed data are spatially calibrated and verified. Three feature points on the glass shell are selected to check whether the spatial coordinates of these feature points in the depth, RGB, and thermal imaging data are consistent. If there is a deviation, the transformation parameters are fine-tuned until the spatial coordinate error of the three is less than 0.5 mm, and finally, the spatially aligned perception data is obtained.

[0029] Step 2.3 involves cleaning the spatially aligned sensing data to remove outliers and noise, resulting in preprocessed data. Specifically, this includes outlier and noise removal for depth data, RGB image data, and thermal imaging data. For depth data, a normal range is defined based on the known dimensions of the precision glass housing. For example, in a sorting station scenario, the depth of the glass housing from the sensor should be between 0.8 meters and 1.2 meters. Depth values ​​outside this range are considered outliers and removed. Simultaneously, for isolated flying point noise caused by dust reflection in the depth data, a neighborhood mean method is used to calculate the average depth of the eight neighboring points around each depth point. If the difference between the depth point and the average value exceeds 0.3 millimeters, it is identified as a noise point and replaced with a neighboring point. For RGB image data, to address the light spot noise generated by workshop lights reflecting off the glass shell surface, the difference between pixel grayscale values ​​and surrounding pixels is analyzed. If a pixel's grayscale value exceeds 20% of the average grayscale value of its surrounding 3×3 neighboring pixels, it is identified as light spot noise, and interpolation is used to replace the noisy pixel with the grayscale value of the surrounding pixels. Simultaneously, isolated noise points caused by camera sensor errors are removed. These noise points differ significantly from the normal transparency or printed color of the glass shell; they are filtered by color space range, and pixels exceeding this range are identified as noise points and removed. For thermal imaging data, based on the normal temperature range of the glass shell in a workshop environment, temperature values ​​exceeding 30℃ or falling below 15℃ are identified as outliers and removed. A sliding window filtering method is used to smooth the thermal imaging data with a 5×5 window, eliminating temperature noise caused by ambient airflow fluctuations and ensuring that the thermal imaging data accurately reflects the temperature distribution on the glass shell surface. After removing all outliers and noise, the processed data undergoes an integrity check to ensure no critical areas are missing, resulting in the final preprocessed data.

[0030] In this embodiment of the invention, the present invention employs a technical approach of time synchronization of data streams from different sensors based on multi-source heterogeneous sensing data, followed by spatial alignment of the time-synchronized data to unify the spatial coordinate systems of each sensor, and finally data cleaning of the spatially aligned data to remove outliers and noise. This overcomes the technical problems of time misalignment caused by acquisition delays in multi-source heterogeneous sensing data streams, spatial mismatch caused by differences in the spatial coordinate systems of each sensor due to parameter variations, and residual outliers and noise affecting subsequent processing accuracy after spatial alignment. This achieves the technical effect of obtaining standardized sensing data with consistent time dimension, unified spatial dimension, and pure data quality. This provides a basis for accurately extracting visual, geometric, and thermodynamic features related to the physical properties of objects from the standardized sensing data, and for reliably inferring the physical properties of target objects through deep neural networks.

[0031] In a preferred embodiment of the present invention, step 3 above may include: Step 3.1: Based on the obtained standardized perception data, visual features associated with the physical properties of the object are extracted. Specifically, this includes extracting visual features associated with the physical properties of the micro-sensor components in the precision glass shell, using the RGB image data in the obtained standardized perception data as the core. First, the RGB color channel distribution of the image is analyzed, and the pixel value range of the red, green, and blue channels is statistically analyzed to determine the basic color attributes of the glass shell. If the pixel values ​​are concentrated in the high brightness range and the values ​​of the three channels are similar, it is determined to be the transparency characteristic of the glass shell. If there are local clusters of specific color pixels, they are identified as possible printed marks or markings on the shell surface. This color information can help determine the purity of the shell material and thus correlate the influence of material uniformity on fragility. Next, the texture features of the shell surface are extracted by scanning the grayscale changes of the image pixels row by row and column by column. If the grayscale values ​​are generally stable and the fluctuation range is small, it indicates that the shell surface is smooth. If there are continuous areas with sudden changes in local grayscale values, they are determined to be surface scratches or defects. The smoothness of the texture is directly related to the subsequent inference of the friction coefficient, while scratches and other defects are related to the fragility threshold assessment. Finally, the gloss characteristics of the outer shell surface are calculated, and the reflective areas caused by workshop lights in the image are captured. The area ratio of the reflective areas and the peak brightness of the pixels are counted. If the reflective areas are concentrated and the peak brightness is high, it indicates that the gloss of the outer shell surface is high. Conversely, the gloss is low. The gloss characteristics can help verify whether the outer shell material is glass. At the same time, it reflects the influence of surface flatness on the force distribution during contact. Finally, the color attributes, texture features and gloss characteristics are integrated to form a preliminary set of visual features.

[0032] Step 3.2 involves further extracting geometric features associated with the physical properties of the object from the obtained visual features and standardized perceptual data. Specifically, this includes combining the obtained visual features with the depth image data in the standardized perceptual data to extract geometric features associated with the physical properties of the micro-sensor element in the precision glass shell. First, a 3D point cloud model of the object is constructed based on the depth data. By traversing all data points in the point cloud, the maximum length, maximum width, and maximum height of the shell in 3D space are calculated. Then, the approximate volume of the shell is obtained by multiplying these three values. This volume data is a crucial foundation for subsequent inferences about the object's mass. Next, the shape features of the 3D point cloud are analyzed, and the matching degree between the point cloud outline and standard geometry is compared. If the point cloud outline is close to a cuboid and the edges and corners are clear, it is judged as a regular shape. If there are curved surfaces or irregular protrusions, it is recorded as a special shape structure. The regularity of the shape is directly related to the stiffness characteristics of the object. For example, the stiffness of a regular thin-shell structure is usually lower than that of a thick-walled regular structure. Then, the flatness of the shell surface is calculated. Multiple evenly distributed sampling areas are selected, and the fluctuation range of depth data in each area is statistically analyzed. If the fluctuation range is less than a set threshold, it indicates that the surface is flat. If the fluctuation range is large, it is judged that the surface is uneven or deformed. The surface flatness will affect the uniformity of force during grasping, and thus be related to the fragility assessment. Finally, the characteristic structure of the shell is identified, such as whether there are interfaces for installation, protruding pins or recessed slots. The three-dimensional dimensions and positions of these structures are recorded. The presence and size of the characteristic structure will affect the center of gravity distribution of the object, and thus help to infer the mass distribution and grasping stability. Finally, the volume, shape features, surface flatness, and characteristic structure information are integrated to form a complete set of geometric features.

[0033] Step 3.3 involves further extracting thermodynamic features associated with the physical properties of the object from the obtained geometric features and the standardized sensing data. Specifically, this includes: based on the obtained geometric features and thermal imaging data from the standardized sensing data, extracting thermodynamic features associated with the physical properties of the micro-sensor element within the precision glass shell. First, the overall surface temperature of the shell is collected from the thermal imaging data. The average value of all temperature measurement points is calculated to obtain the average surface temperature of the shell. The difference between this temperature and the ambient temperature in the workshop is compared. If the difference is less than 1°C, it indicates that the shell has weak thermal conductivity, consistent with the characteristics of glass. If the difference is large, it is necessary to combine the geometric features to determine whether there are internal metal components. Next, the temperature uniformity of the shell surface is analyzed, and the standard deviation of the temperature at all temperature measurement points is calculated. If the standard deviation is small, it indicates that the surface temperature of the shell is uniform and the material is consistent. If the standard deviation is large, there are local temperature anomalies. It is necessary to combine visual and geometric features to check whether the temperature anomalies are caused by surface stains, internal structural differences, or shell damage. Temperature uniformity can help judge the purity of the material and the integrity of the internal structure, and then correlate it with stiffness and fragility. Then observe the rate of temperature change of the shell over a certain time interval, select thermal imaging data at multiple time points, and calculate the temperature change amplitude at the same location. If the change amplitude is small, it indicates that the shell has good thermal stability. If the change amplitude is large, there may be material defects. Thermal stability is related to the heat resistance and structural stability of the object. Finally, integrate the average surface temperature, temperature uniformity, and temperature change rate to form a set of thermodynamic characteristics.

[0034] Step 3.4 involves fusing and vectorizing the extracted visual, geometric, and thermodynamic features to obtain a multi-dimensional physical feature vector. Specifically, this includes: quantizing the extracted visual, geometric, and thermodynamic features separately; converting color attributes in the visual features into average pixel values ​​of the RGB three channels, texture features into the standard deviation of grayscale fluctuations, and gloss features into the proportion of reflective areas; converting volume in the geometric features into specific cubic millimeter values, shape features into shape matching coefficients, surface flatness into the maximum difference in depth fluctuations, and feature structure into the number of features and the dimensions of each feature; and converting the average surface temperature in the thermodynamic features into degrees Celsius values, temperature uniformity into temperature standard deviations, and temperature change rate into temperature change per unit time. The quantified values ​​are then determined, and the order of all quantified indicators is established. Following a fixed order of visual feature indicators, geometric feature indicators, and thermodynamic feature indicators, each quantified indicator is treated as a dimension and sequentially filled into the feature vector. For example, the average pixel value of the RGB three channels, the standard deviation of texture grayscale fluctuation, and the gloss and reflectivity ratio are filled in first, followed by the volume value, shape matching coefficient, surface flatness depth difference, feature quantity and size, and finally the average temperature, temperature standard deviation, and temperature change rate. If any indicator under a certain feature is missing, it is supplemented with a preset default value to ensure that the dimension of the feature vector is fixed and complete. Finally, a multi-dimensional physical feature vector containing all quantified feature indicators is formed. This vector can comprehensively cover the key information related to the physical properties of the micro sensor element with a precision glass shell, providing standardized input data for subsequent deep neural network inference calculations.

[0035] In this embodiment of the invention, the present invention overcomes the technical problems of traditional technologies, such as the inability of single feature extraction to fully cover the associated information of physical attributes of objects, the isolated existence of different types of features leading to a one-sided understanding of the physical attributes of objects, and the impact on the accuracy of attribute inference, by employing a technical means of first extracting visual features associated with the physical attributes of objects based on standardized perceptual data, then further extracting geometric features by combining visual features with standardized perceptual data, and then further extracting thermodynamic features by combining geometric features with standardized perceptual data, and finally fusing and vectorizing the visual features, geometric features, and thermodynamic features. At the same time, it solves the defect that the lack of effective integration of multi-dimensional features makes it difficult to associate appearance with physical attributes in subsequent reasoning, thereby achieving the technical effect of constructing a multi-dimensional physical feature vector that contains multi-dimensional associated information of objects in terms of vision, geometry, and thermodynamics.

[0036] In a preferred embodiment of the present invention, step 4 above may include: Step 4.1: Based on the obtained multi-dimensional physical feature vectors, a multi-layer nonlinear transformation is performed through a deployed deep neural network to obtain high-dimensional abstract features. Specifically, this includes: first, inputting the obtained multi-dimensional physical feature vectors into a deep neural network pre-deployed on a cloud processor. This network has been optimized for the sorting scenario of precision glass shell micro-sensor components in electronic component manufacturing workshops. The network structure includes three convolutional layers and two fully connected layers. The first convolutional layer performs local correlation extraction on the visual and geometric features in the feature vectors, such as capturing the synergistic relationship between smooth texture and high surface flatness. These relationships are directly related to the material uniformity of the glass shell. The second convolutional layer further integrates thermodynamic features with the local features output from the previous layer to establish a low-temperature... The potential correlation between standard deviation and high gloss strengthens the feature combination related to physical properties. The third convolutional layer deeply fuses the features output by the first two layers, filtering out feature combinations that are strongly correlated with the physical properties of precision glass components, such as small volume, low temperature change rate, and smooth texture. Then, two fully connected layers perform nonlinear transformation on the features output by the convolutional layer. Each layer uses the ReLU activation function to break the linear constraints between features, transforming low-dimensional feature combinations such as volume of 9.8 cubic millimeters, temperature standard deviation of 0.25℃, and reflectivity of 28% into high-dimensional abstract features that reflect intrinsic physical properties such as stiffness and fragility. Finally, a high-dimensional abstract feature set containing complex correlation information between features is obtained, providing deep data support for subsequent physical property inference.

[0037] Step 4.2: For the obtained high-dimensional abstract features, feature mapping is performed using a deep neural network. Based on the pre-learned correlation between appearance features and physical properties, preliminary inferences of physical properties are obtained. Specifically, the obtained high-dimensional abstract features are input into the feature mapping layer of the deep neural network. This mapping layer has pre-learned the correlation between appearance features and physical properties using massive samples. The training samples include 6000 sets of data on micro-sensor elements with different specifications of precision glass shells. Each set of samples is labeled with visual, geometric, and thermodynamic appearance features and corresponding real physical properties. The feature mapping layer first matches the most similar samples in the pre-learned correlation database based on the core combinations in the current high-dimensional abstract features. The clusters of physical properties are concentrated in stiffness of 208-220 MPa and fragility threshold of 6.5-7.0 N. The values ​​are then refined through the network's property prediction branches: the stiffness prediction branch, combined with a shape matching coefficient of 0.93, determines that the glass shell is a thin-shell, regular structure, with a stiffness value biased towards the lower limit of the cluster, initially inferred to be 210 MPa; the fragility prediction branch, combined with a texture grayscale fluctuation standard deviation of 4 and good temperature uniformity, determines that the glass material has high purity and no obvious defects, with a fragility threshold biased towards the upper limit of the cluster, initially inferred to be 6.9 N. Finally, integrating the results of the two branches, the preliminary inference results of the physical properties are obtained: stiffness approximately 210 MPa and fragility threshold approximately 6.9 N.

[0038] Step 4.3: For the preliminary inference results of the obtained physical properties, calculate their confidence level using a deep neural network and output a physical property of the target object and its corresponding confidence level. Specifically, this includes: evaluating the reliability of the preliminary inference results using the confidence calculation module of the deep neural network. First, calculate the error between the preliminary inference results and similar sample clusters in the pre-learning association database: the error between the current inferred stiffness of 210 MPa and the cluster average stiffness of 214 MPa is 1.8%, and the error between the inferred fragility threshold of 6.9 N and the cluster average threshold of 6.7 N is 2.9%. Both errors are less than the preset 8% threshold, laying the foundation for confidence. Next, check the input multi-dimensional physical features... The integrity of the data was assessed, confirming that the quantitative indicators of visual, geometric, and thermodynamic features were all complete and met the requirements, further enhancing the confidence level. The current sorting environment was then compared with the training sample collection environment: the actual workshop temperature was 24.9℃ and the illumination was 485 lux, showing minimal difference from the training environment, with no interference factors such as dust or light obstruction. This environmental consistency improved the confidence level. Finally, the confidence calculation module comprehensively considered the error, integrity, and environmental consistency to output the confidence level of the current physical property inference results: stiffness 210 MPa and fragility threshold 6.9 N. These two sets of physical property-confidence data were used as the physical property inference results for the target object, providing data support for subsequent decision-making and control.

[0039] In this embodiment of the invention, the present invention overcomes the technical problems of existing contact measurement methods, which can only obtain a single physical attribute at a time, offline simulation analysis is time-consuming and cannot meet the real-time requirements of dynamic scenes, and lacks reliability assessment of physical attribute inference results. This invention adopts a technique that uses multi-dimensional physical feature vectors obtained through multi-layer nonlinear transformation of deployed deep neural networks to obtain high-dimensional abstract features, and then uses deep neural networks to perform feature mapping on the high-dimensional abstract features and obtain preliminary inference results based on the correlation between pre-learned appearance features and physical attributes. Finally, the invention uses deep neural networks to calculate the confidence level of the preliminary inference results and outputs a physical attribute of the target object and its corresponding confidence level. This invention achieves the technical effect of being able to infer at least one physical attribute of the target object in real time and accurately in high-precision, high-dynamic scenes, and can intuitively reflect the reliability of the inference results through confidence level, providing accurate and reliable data support for the embodied intelligent agent to formulate adaptive action strategies.

[0040] In a preferred embodiment of the present invention, step 5 above may include: Step 5.1: Based on a physical attribute of the target object output by the deep neural network and its corresponding confidence level, an initial physical attribute database is obtained through structured storage. Specifically, this includes: the first output physical attribute of the target object and its corresponding confidence level. For the precision glass shell micro-sensor element in the electronic component manufacturing workshop, the output physical attributes at this stage include stiffness inference results and fragility threshold inference results. The stiffness inference value of 212 MPa corresponds to a confidence level of 90%, and the fragility threshold inference value of 6.8 N corresponds to a confidence level of 87%. Then, these data are classified and organized according to the type of physical attribute, separating stiffness-related attributes... The inferred values ​​and confidence scores are grouped together, while the inferred values ​​and confidence scores related to the fragility threshold are grouped together in another group. This ensures that the attribute type and the corresponding confidence score correspond one-to-one within each data group, avoiding data confusion. Then, each group of data is packaged using a data packaging method. During the packaging process, a temporary identifier is added to each data group. The temporary identifier contains the time information of data generation and the sensor number information. The time information is accurate to milliseconds to distinguish the component data collected and inferred at different times. The sensor number corresponds to the multimodal sensor group number that collected the data of that component. Finally, a packaged data packet containing attribute type, inferred value, confidence score, and temporary identifier is formed.

[0041] Step 5.2: Based on the initial physical attribute database, the physical attributes are bound to their confidence scores through confidence score calculation and association to obtain a standardized physical attribute database. Specifically, this includes: First, based on the obtained packaged data packet, assigning a unique object identifier to each precision glass-shell micro-sensor element. The object identifier is pre-set by the workshop production system and includes the element's production batch, specifications, and serial number information, such as SG20240901001, where SG represents the precision glass shell element, 202409 represents the September 2024 production batch, and 01001 represents the first element in that batch. Next, the contents of the packaged data packet are split and supplemented to clarify the specific content of each field: the first field is the object... The first field is the identifier, filled with a unique assigned number; the second field is the physical attribute type, labeled as stiffness or fragility threshold; the third field is the inferred physical attribute value, labeled with the specific numerical value and corresponding unit, such as 212MPa or 6.8N; the fourth field is the confidence score, presented as a percentage, such as 90% or 87%. The fields are then arranged in a fixed order: object identifier, physical attribute type, inferred physical attribute value, and confidence score, forming structured data in a unified format. Simultaneously, it is ensured that all data have standardized unit labeling and consistent number of decimal places, for example, inferred values ​​retain one decimal place, and confidence scores retain an integer. This results in standardized structured physical attribute data, ensuring that the embodied intelligent agent's decision control system can quickly identify and analyze it.

[0042] Step 5.3: Based on the standardized physical attribute database, data indexing and query optimization are used to obtain the final physical attribute database and corresponding confidence scores that can be used for real-time interaction. Specifically, the obtained structured physical attribute data is transmitted to the decision control terminal of the embodied intelligent agent via the industrial Ethernet in the workshop. Here, the embodied intelligent agent is the intelligent robot decision control system responsible for sorting tasks. During the transmission, a real-time data transmission protocol is used to ensure that the data arrives at the decision control system within 100 milliseconds to match the sorting rhythm of 15 pieces per minute on the production line. After receiving the structured data, the decision control system first parses the data, identifies the identity of the precision glass shell micro-sensor element to be sorted by recognizing the object identifier, and then extracts the physical attribute type, inferred value, and confidence score corresponding to the element. For example, when the fragility threshold of object identifier SG20240901001 is parsed to be 6.8N with a confidence level of 87%, the decision control system will match this data with the preset action strategy library to determine that the gripping force for this component should be controlled below 6N to avoid exceeding the fragility threshold and causing scratches on the shell. At the same time, the system will record this structured data for traceability of subsequent sorting actions. If a sorting abnormality occurs, the corresponding physical attribute information can be inferred by querying the object identifier. Finally, the structured physical attribute data is output to the decision control system, providing accurate data support for the robot to adjust its sorting actions in real time.

[0043] In this embodiment of the invention, the present invention overcomes the technical problems of existing technologies where physical attribute data is scattered, disordered, and inconsistently formatted, making it difficult for embodied intelligent agents to accurately match target objects with corresponding attributes, and lacking standardized data input to support efficient action planning. This is achieved by using a technical means to encapsulate data based on the physical attributes of the output target object and their corresponding confidence scores, and then standardizing and formatting the encapsulated data to generate structured physical attribute data containing object identifiers, physical attribute types, inferred values, and confidence scores. Finally, this structured physical attribute data is output to the embodied intelligent agent decision control.

[0044] In a preferred embodiment of the present invention, step 5 above may include: Step 5.4: Based on the final physical attribute database and corresponding confidence scores that can be used for real-time interaction, the data is encapsulated and formatted for transmission by the cloud processor to obtain a standardized data stream. Specifically, the decision control system of the embodied intelligent agent first receives the output structured physical attribute data, which includes the object identifier SG20240901001 of the precision glass shell micro-sensor element, the physical attribute type and inferred value stiffness 210MPa with a confidence level of 91%, and fragility threshold 6.9N with a confidence level of 88%. The decision control system first analyzes the data to confirm that the core physical properties of the component to be sorted are low and the fragility threshold is high, so it is necessary to avoid excessive gripping force that could scratch the casing. The stiffness is at a medium level, and the inertial impact during the operation needs to be controlled to prevent structural damage. Then, the decision control system calls the preset control strategy library, which is built based on historical data of electronic component sorting scenarios. It includes action principles corresponding to different fragility thresholds and stiffness ranges: for example, a fragility threshold of 5-7N corresponds to a low-force gripping strategy, and a stiffness of 200-220MPa corresponds to a smooth movement strategy. Combining the current component properties, the decision control system matches a preliminary control strategy of low-force gripping, smooth movement, and buffer placement. It specifies that the gripping force must be lower than the fragility threshold of 6.9N, the movement speed must be lower than 60% of the normal sorting speed, and a buffer distance must be reserved during placement. Finally, a preliminary control strategy adapted to the physical properties of the component is formed.

[0045] Step 5.5: For the standardized data stream, data is transmitted through the real-time communication link established between the embodied intelligent agent decision control system. Specifically, the decision control system performs real-time action planning based on the obtained preliminary control strategy and the geometric feature data of the precision glass shell component. First, determine the gripping parameters of the robotic gripper: based on the length and width dimensions of the component, set the opening degree of the robotic gripper to 30% to ensure that the gripper fingers can stably wrap the component without squeezing the outer shell; referencing the fragility threshold of 6.9N, precisely set the gripping force to 5.5N. Next, plan the movement path and speed: combined with the sorting rhythm of 15 pieces per minute on the production line, set the movement speed to 0.2 meters per second, and plan the path to be a straight line that avoids other components on the production line to reduce inertial impact during turning; after reaching above the sorting station, set the descent buffer distance to 2 centimeters to prevent the robotic gripper from directly hitting the sorting table and causing vibration damage to the component. Finally, plan the placement action parameters: after descending to the buffer distance, the robotic gripper slowly releases at an opening degree increase of 10% per second until the component is completely released, preventing friction scratches on the outer shell during the release moment. Through the above planning, the preliminary control strategy is transformed into specific control instructions including the robotic gripper opening degree, gripping force, movement speed, descent distance, and release rate.

[0046] Step 5.6: For the data successfully transmitted to the decision control system, through parsing and application, drive the embodied intelligent agent to perform real-time interactive tasks that match the physical properties of the target object, completing the real-time inference closed loop of physical properties. Specifically, the decision control system transmits the generated specific control commands to the action execution mechanism of the embodied intelligent agent in real time via the industrial Ethernet in the workshop. This execution mechanism includes a mechanical gripper drive module and a motion drive module. The transmission delay is controlled within 50 milliseconds to match the sorting rhythm of the assembly line. After receiving the command, the mechanical gripper drive module drives the internal servo motor to adjust the gripper position so that the opening degree reaches 30%. At the same time, the built-in pressure feedback unit monitors the gripping force in real time. When the force reaches 5.5N, it automatically maintains the pressure stability to avoid force fluctuations exceeding the safe range. After receiving the command, the motion drive module controls the robot arm to move along the planned straight path at a speed of 0.2 meters per second. During the process, the encoder corrects the position deviation in real time to ensure that the movement accuracy is within 0.5 millimeters. When the robotic arm reaches the sorting station, the motion drive module controls the arm to descend 2 centimeters, triggering the mechanical gripper drive module to perform a release action. The gripper fingers slowly open at an increasing rate of 10% per second until the component is stably placed in the designated area of ​​the sorting table. Throughout the entire process, the motion actuator provides real-time feedback on its operating status to the decision control system. If a slight deviation in gripping force occurs, the decision control system immediately fine-tunes the drive commands to ensure that the component remains in a safe interactive state at all times.

[0047] In this embodiment of the invention, the present invention overcomes the technical problems of existing technologies, such as relying on contact measurement to adjust actions, which easily leads to damage to the target object; offline simulation to adjust parameters, which is time-consuming and cannot match the high dynamic rhythm of 15 pieces per minute on the production line; and the disconnect between the decision control of the embodied intelligent agent and the physical properties of the object, which leads to a lack of adaptability of interactive actions. This is because the present invention adopts structured physical attribute data based on output, generates a preliminary control strategy adapted to the physical attributes of the target object through decision control of the embodied intelligent agent, performs real-time action planning on the preliminary control strategy to obtain specific control instructions, and finally outputs the specific control instructions to the action execution mechanism of the embodied intelligent agent.

[0048] Embodiments of the present invention also provide a real-time reasoning system for the physical properties of objects, comprising: The multimodal non-contact sensing module is used to collect raw sensing data of the target object through multimodal non-contact sensors, and to preprocess the raw sensing data of the target object to obtain multi-source heterogeneous sensing data. The data preprocessing and feature fusion module is used to perform time synchronization, spatial alignment and data cleaning on multi-source heterogeneous sensing data to obtain preprocessed data. Visual features, geometric features and thermodynamic features associated with the physical properties of objects are extracted from the preprocessed data to obtain multi-dimensional physical feature vectors. The physical attribute real-time inference engine module is used to perform multi-layer nonlinear transformation and feature mapping on multi-dimensional physical feature vectors through deep neural networks, and infer a physical attribute of the target object and its corresponding confidence level based on the correlation between appearance features and physical attributes. The physical attribute database and output module is used to construct a physical attribute database based on a physical attribute of the target object and its confidence level, obtain the corresponding confidence score, and output the physical attribute database and confidence score to the decision control system of the embodied intelligent agent to realize real-time reasoning on the physical attributes of the object.

[0049] In embodiments of the present invention, the technical problem to be solved by the present invention includes: Limitations and security risks of information acquisition methods: Existing technologies heavily rely on physical contact, which means that the measurement process itself carries extremely high risks when dealing with fragile, high-temperature, electrically charged, or chemically hazardous objects, potentially damaging the object or the sensor itself. The primary problem this invention aims to solve is how to overcome these limitations and achieve non-contact, undisturbed physical property sensing while ensuring absolute safety.

[0050] The lack of real-time interaction and decision-making capabilities: Current physical property analysis methods, especially those based on 3D modeling and simulation, have long processing times and huge computational demands, falling into the category of offline analysis and failing to meet the needs of embodied intelligent agents for real-time task decision-making. This latency severely limits the interactive experience, preventing robots from making immediate and appropriate physical operations in dynamically changing environments, such as quickly sorting packages of different materials and weights on an assembly line.

[0051] Poor system robustness and environmental adaptability: Systems based on fixed physical models or manually set parameters exhibit extremely poor robustness. When lighting conditions change, objects are partially occluded, or the system encounters entirely new objects it has never seen before, its accuracy drops drastically or it may even fail completely. Constructing a perception system that can adapt to complex and changing environments and possesses good generalization ability for unknown objects is a key bottleneck in achieving truly intelligent interaction.

[0052] The inherent limitations and incompleteness of our understanding of the physical world: Traditional sensors are often single-function; for example, force sensors can only measure force, and vision sensors can only perceive color and shape. This "blind men and the elephant" approach to data acquisition leads to a fragmented understanding of objects. Intelligent interaction, however, is a comprehensive decision-making process that requires considering multiple attributes of an object simultaneously, such as its mass, fragility, and surface friction. The core challenge in resolving the "clumsy" and "unnatural" nature of interactive behavior lies in how to integrate multi-source, heterogeneous sensory information into a complete and three-dimensional understanding of physical properties.

[0053] The Gap Between Intelligent Behavior and Physical Cognition: A significant gap exists in current embodied intelligence technology: the disconnect between "perception" and "cognition." A robot can accurately reconstruct a three-dimensional model of an object (perception), but it doesn't know whether the object is "heavy" or "light," "hard" or "soft" (cognition). This lack of physical cognition leads to behaviors that do not conform to physical laws, inefficient interactions, and an inability to complete delicate and complex physical manipulation tasks. This fundamentally limits the depth and breadth of embodied intelligence's application in the real world.

[0054] In embodiments of the present invention, such as Figure 2 As shown, the complete product structure of the present invention includes: Modal non-contact sensing module: integrates a depth camera, a high-resolution RGB camera, and (optional) a thermal imaging camera. The depth camera is used to acquire the object's three-dimensional geometry, size, and depth information; the RGB camera is used to capture the object's color, texture, material appearance, and other two-dimensional information; the thermal imaging camera is used to acquire the object's surface temperature distribution to help determine the material (such as metal or plastic).

[0055] The data preprocessing and feature fusion module synchronizes and aligns the data stream from the multimodal perception module in real time. It denoises and enhances the data, extracting key feature vectors, such as volume, shape contour, and surface smoothness from 3D data; and texture complexity, gloss, and color histograms from 2D images.

[0056] Real-time Physical Property Inference Engine: The core of this invention. This engine is based on a pre-trained deep neural network (such as a Convolutional Neural Network (CNN) or a Graph Neural Network (GNN)) using a fused feature vector as input. This network is obtained through supervised learning on a dataset of objects containing a large number of known physical properties (mass, material, fragility, coefficient of friction, etc.). After receiving real-time data, the engine can instantly infer multiple physical properties of the current target object and their confidence levels.

[0057] Physical Attribute Database and Output Module: This module stores the structured data output by the inference engine, including the object ID, the inferred values ​​of various physical attributes (such as mass: 0.5kg, fragility: 85%), and the corresponding confidence scores, and provides them to the embodied intelligence decision control system.

[0058] In embodiments of the present invention, regarding the background art as follows: Figure 3 The differences between the present invention and the prior art shown are as follows: The fundamental difference in sensing methods: This invention completely abandons physical contact and uses non-contact methods of vision and depth perception as the source of information, fundamentally solving the limitations of contact measurement.

[0059] Reasoning instead of direct measurement: Existing technologies rely on direct measurement of physical feedback, while this invention uses an AI model to learn the relationship between "appearance" and "intrinsic properties" to achieve "reasoning"-style property acquisition, which has the ability to generalize to unknown objects.

[0060] Real-time and comprehensive: Existing technologies are either real-time but single-function (contact sensors) or comprehensive but not real-time (offline simulation). This invention, through a highly efficient deep learning model, elevates the analysis of comprehensive physical properties to a real-time level for the first time, and can simultaneously output multi-dimensional information such as mass, material, and fragility.

[0061] Artificial intelligence driven: This invention applies artificial intelligence technology, especially deep learning models, for data analysis and reasoning, enabling the system to learn autonomously from data rather than relying on fixed physical formulas or artificial rules, which greatly improves the robustness and adaptability of the technology.

[0062] In embodiments of the present invention, the technical means and effects of the present invention include: Multimodal non-contact sensing technology: Technical means: It comprehensively utilizes depth cameras, high-resolution RGB cameras, and thermal imaging cameras to simultaneously acquire 3D geometric, 2D visual, and thermodynamic distribution information of target objects from a safe distance. This technology serves as the information entry point for the entire system. It overcomes the limitation of traditional methods that require physical contact, enabling rapid and comprehensive data acquisition of any object without risk or interference, providing a rich and multi-dimensional data foundation for subsequent intelligent reasoning.

[0063] Multi-source heterogeneous data fusion and feature extraction technology: Develop specialized algorithms to perform time synchronization, spatial alignment, and data cleaning on heterogeneous data streams from different sensors, and extract deep features that indirectly reflect physical properties, such as extracting volume from 3D contours, associating friction with texture complexity, and using thermodynamic properties to aid in material identification. This technology serves as a bridge between "raw perception" and "advanced cognition." It transforms messy, multi-source raw data into standardized, high-information-density feature vectors that AI models can understand and process, a prerequisite for accurate reasoning.

[0064] The invention employs a deep learning-based physical property reasoning technique: It constructs and trains a deep neural network model that learns the complex mapping relationships between a vast amount of "object appearance features" and "real physical properties." This is the "brain" and core innovation of the invention. It enables leapfrog reasoning from an object's external appearance to its intrinsic physical properties. Because it is based on an AI model, its reasoning speed is extremely fast, and it also possesses excellent generalization ability for novel, unseen objects.

[0065] Embodied Intelligence Decision-Making Empowerment Technology: This technology provides structured physical attribute data with confidence levels, output by the inference engine, to the motion planning and control system of embodied intelligence (such as robots) in real time through a standard interface. This technology is the ultimate embodiment of the value of this invention. It transforms abstract inference results into decision-making basis that can be executed by the machine, enabling the robot to "know not only what, but also why," thereby autonomously generating more refined, safer, and smarter interactive actions, such as adjusting the grasping force based on fragility and estimating the amount of force applied based on mass.

[0066] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A real-time reasoning method for the physical properties of an object, characterized in that, The method includes: Step 1: Collect raw sensing data of the target object using a multimodal non-contact sensor, and preprocess the raw sensing data of the target object to obtain multi-source heterogeneous sensing data. Step 2: Perform time synchronization, spatial alignment, and data cleaning on the multi-source heterogeneous sensing data to obtain preprocessed data; Step 3: Extract visual features, geometric features, and thermodynamic features associated with the physical properties of the object from the preprocessed data to obtain a multi-dimensional physical feature vector; Step 4: Perform multi-layer nonlinear transformation and feature mapping on the multi-dimensional physical feature vector through a deep neural network. Based on the correlation between appearance features and physical attributes, infer a physical attribute of the target object and its corresponding confidence level. Step 5: Based on a physical attribute of the target object and its confidence level, construct a physical attribute database, obtain the corresponding confidence score, and output the physical attribute database and confidence score to the decision control system of the embodied intelligent agent to realize real-time reasoning on the physical attributes of the object.

2. The real-time reasoning method for the physical properties of an object according to claim 1, characterized in that, Raw sensing data of the target object is acquired through multimodal non-contact sensors, and this raw sensing data is preprocessed to obtain multi-source heterogeneous sensing data, including: Step 1.1: Acquire raw sensing data of the target object using the multimodal non-contact sensor; Step 1.2: Based on the collected raw sensing data, preprocessing is performed by cloud processor to remove noise and enhance the data, resulting in improved sensing data. Step 1.3: The improved multi-source sensing data is initially fused by the cloud processor to obtain multi-source heterogeneous sensing data.

3. The real-time reasoning method for the physical properties of an object according to claim 2, characterized in that, Multi-source heterogeneous sensing data undergoes time synchronization, spatial alignment, and data cleaning to obtain preprocessed data, including: Step 2.1: Based on the obtained multi-source heterogeneous sensing data, time-synchronized sensing data is obtained by synchronizing the data streams from different sensors. Step 2.2: The obtained time-synchronized sensing data is processed by spatial alignment to unify the spatial coordinate system of each sensor data, thus obtaining spatially aligned sensing data. Step 2.3: The obtained spatially aligned perception data is cleaned to remove outliers and noise, resulting in preprocessed data.

4. The real-time reasoning method for the physical properties of an object according to claim 3, characterized in that, Visual features, geometric features, and thermodynamic features associated with the physical properties of objects are extracted from the preprocessed data to obtain a multi-dimensional physical feature vector, including: Step 3.1: Based on the obtained standardized perception data, extract visual features associated with the physical properties of the object; Step 3.2: Further extract geometric features associated with the physical properties of objects from the obtained visual features and standardized perceptual data; Step 3.3: Further extract thermodynamic features associated with the physical properties of the object from the obtained geometric features and the standardized sensing data; Step 3.4: The extracted visual features, geometric features, and thermodynamic features are fused and vectorized to obtain a multi-dimensional physical feature vector.

5. The real-time reasoning method for the physical properties of an object according to claim 4, characterized in that, By employing deep neural networks to perform multi-layer nonlinear transformations and feature mapping on multi-dimensional physical feature vectors, and based on the correlation between appearance features and physical attributes, a physical attribute of the target object and its corresponding confidence level are inferred, including: Step 4.1: Based on the obtained multi-dimensional physical feature vectors, perform multi-layer nonlinear transformations through the deployed deep neural network to obtain high-dimensional abstract features; Step 4.2: For the obtained high-dimensional abstract features, feature mapping is performed through a deep neural network. Based on the pre-learned correlation between appearance features and physical properties, preliminary inference results of physical properties are obtained. Step 4.3: Based on the preliminary inference results of the obtained physical properties, calculate their confidence level using a deep neural network, and output a physical property of the target object and its corresponding confidence level.

6. The real-time reasoning method for the physical properties of an object according to claim 5, characterized in that, Based on a physical attribute of the target object and its confidence level, a physical attribute database is constructed, and the corresponding confidence score is obtained, including: Step 5.1: Based on a physical attribute of the target object output by the deep neural network and its corresponding confidence level, an initial physical attribute database is obtained through structured storage; Step 5.2: For the initial physical attribute database, the physical attributes are bound to their confidence scores through confidence score calculation and association to obtain a standardized physical attribute database; Step 5.3: Through data indexing and query optimization, the standardized physical attribute database is used to obtain the final physical attribute database and corresponding confidence scores that can be used for real-time interaction.

7. The real-time reasoning method for the physical properties of an object according to claim 6, characterized in that, The physical attribute database and confidence scores are then output to the embodied agent's decision control system to enable real-time reasoning about the physical attributes of objects, including: Step 5.4: Based on the final physical attribute database that can be used for real-time interaction and the corresponding confidence scores, the data is encapsulated and formatted for transmission through the cloud processor to obtain a standardized data stream; Step 5.5: For the standardized data stream, data is transmitted through the real-time communication link established between the embodied intelligent agent decision control system; Step 5.6: The data successfully transmitted to the decision control system is analyzed and applied to drive the embodied intelligent agent to perform real-time interactive tasks that match the physical properties of the target object, thus completing the real-time inference loop of physical properties.

8. A real-time reasoning system for the physical properties of an object, the system implementing the method as described in any one of claims 1 to 7, characterized in that, include: The multimodal non-contact sensing module is used to collect raw sensing data of the target object through multimodal non-contact sensors, and to preprocess the raw sensing data of the target object to obtain multi-source heterogeneous sensing data. The data preprocessing and feature fusion module is used to perform time synchronization, spatial alignment and data cleaning on multi-source heterogeneous sensing data to obtain preprocessed data. Visual features, geometric features and thermodynamic features associated with the physical properties of objects are extracted from the preprocessed data to obtain multi-dimensional physical feature vectors. The physical attribute real-time inference engine module is used to perform multi-layer nonlinear transformation and feature mapping on multi-dimensional physical feature vectors through deep neural networks, and infer a physical attribute of the target object and its corresponding confidence level based on the correlation between appearance features and physical attributes. The physical attribute database and output module is used to construct a physical attribute database based on a physical attribute of the target object and its confidence level, obtain the corresponding confidence score, and output the physical attribute database and confidence score to the decision control system of the embodied intelligent agent to realize real-time reasoning on the physical attributes of the object.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Space intelligent visual physical process inference method based on implicit physical large model

    CN121706990A