Information body construction method of environment understanding type service robot
By constructing object-level, relation-level, and scenario-level information bodies, the problem of service robots lacking unified semantic expression in complex environments is solved, enabling dynamic understanding of the environment and adaptive decision-making capabilities.
Patent Information
- Application Number
- CN202610030413.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2046-01-12
AI Technical Summary
In existing technologies, service robots lack a unified expression of environmental semantic information in complex environments, making it impossible to comprehensively describe objects, relationships, and scenarios. This results in incomplete environmental understanding and limited task execution and decision-making capabilities.
Construct object-level, relation-level, and scenario-level information bodies, and generate an environmental semantic knowledge base through multi-source environmental data fusion and cross-modal association, supporting service robots in calling upon these knowledge bodies during task execution and environmental decision-making.
It achieves a unified, dynamic, and evolvable semantic representation of the environment, enhancing the robot's autonomous task execution and adaptive decision-making capabilities in complex environments.
Smart Images

Figure CN121480560B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent service robots, and particularly relates to an information body construction method of an environment understanding type service robot. BACKGROUND
[0002] With the continuous expansion of the application of service robots in complex environments such as hotels, hospitals and families, robots need to have the ability to autonomously perform tasks in unknown or dynamically changing environments. In the prior art, service robots mainly rely on geometric map construction, object recognition based on vision or sensors, and knowledge graph or rule base assisted task planning to understand the environment, wherein the understanding of the environment by the geometric map construction method is through SLAM technology, and the robot can generate a two-dimensional or three-dimensional geometric map of the environment for obstacle avoidance and path planning, but such a method can only provide geometric structure information of the environment and cannot identify the categories, functions and mutual relationships of objects in the environment; the object recognition method based on vision or sensors can detect and locate some environmental objects, but is mostly based on single modal data and is difficult to fuse multi-source perception information, and has poor adaptability to missing perception or environmental changes; some research attempts to assist task planning through a knowledge graph or rule base to achieve semantic-level environmental understanding, but such a method usually relies on static predefined information and cannot dynamically adjust semantic content according to environmental changes. The prior art generally has problems of scattered environmental semantic information, missing hierarchical association and fixed expression structure, and it is difficult to construct an information body that can uniformly describe objects, scenes and their mutual relationships, thereby limiting the comprehensive understanding and adaptive decision-making ability of service robots to the environment. SUMMARY
[0003] In view of the above problems, the present application provides an information body construction method of an environment understanding type service robot, which introduces an object level information body, a relationship level information body and a scene level information body construction mechanism and an environmental semantic knowledge base integration method, solves the problem that traditional service robots lack uniformity and cannot comprehensively describe objects, relationships and scenes in a dynamic environment, resulting in incomplete environmental understanding and limited task execution and decision-making ability.
[0004] To achieve the above purpose, the present application provides an information body construction method of an environment understanding type service robot, comprising the following steps,
[0005] acquiring multi-source environmental data collected during the operation of the service robot, dividing the multi-source environmental data into different modalities according to types, and constructing a single modality feature data set of each object under each modality based on the division result;
[0006] The single-modal feature data set of each perceivable object in the environment in each modality is cross-modally associated, and each modal feature data is fused based on the cross-modal association result to construct an object-level information body of each perceivable object for representing the category, spatial position and attribute of each perceivable object.
[0007] A life cycle identifier is assigned to the object-level information body, and the state of the object-level information body is managed based on a time stamp, so that the object-level information body can be dynamically updated with the change of the environment.
[0008] Based on the dynamically updated object-level information body, the spatial relationship, functional association and interaction relationship between each perceivable object in the environment are analyzed, and a relationship-level information body is generated based on the analysis result for representing the spatial relationship, functional relationship and interaction relationship features of each object in the environment.
[0009] Based on the object-level information body and the relationship-level information body, a scene-level information body is constructed by analyzing the spatial layout and relationship connection of the object for reflecting the overall spatial structure, object distribution and semantic features of the environment.
[0010] The object-level information body, the relationship-level information body and the scene-level information body are uniformly stored as an environmental semantic knowledge base, and the service robot is supported to call in the task execution and environmental decision process.
[0011] The technical scheme provided in the present application has the following technical effects or advantages: facing the demand of service robots in complex, unknown or dynamic environment for autonomous task execution, by constructing the object-level information body, the relationship-level information body and the scene-level information body, and uniformly storing them as an environmental semantic knowledge base, the unified, dynamic and evolvable semantic representation of the environment is realized. Specifically, first, multi-source environmental data acquisition and cross-modal association method is adopted to fuse multi-modal features of each perceivable object in the environment, to construct the object-level information body and realize comprehensive understanding of a single object; then, based on the object-level information body, the spatial relationship, functional association and interaction relationship between objects are analyzed to generate the relationship-level information body, and an object relationship network is constructed to systematically represent the spatial structure and semantic connection between objects; further, based on the spatial features and object attribute features of the object relationship network, the functional area distribution and spatial hierarchical structure in the environment are identified to form the scene-level information body to comprehensively reflect the overall semantic structure of the environment; finally, the various information bodies are integrated and stored in the environmental semantic knowledge base to support the service robot to call in the task execution and environmental decision process, to realize real-time dynamic understanding and adaptive decision of the environment. Compared with the prior art, the present application can realize real-time, accurate and evolvable semantic perception and task support in the environment with multi-modal, multi-object and multi-functional area, and significantly improves the ability of robot environment understanding and adaptive decision. BRIEF DESCRIPTION OF DRAWINGS
[0012] Brief Description of the Drawings
[0013] Figure 1 An information body construction method of an environment understanding type service robot according to an embodiment of the present application. DETAILED DESCRIPTION
[0014] The present application realizes unified representation of objects, relationships and scenes by introducing an information body construction mechanism, improves the accuracy and adaptability of environment understanding through multi-source perception data fusion and dynamic state management, realizes structured expression of environment semantics based on functional area division and spatial level analysis, and realizes autonomous and dynamic decision-making ability of the service robot in a complex environment through integration and calling of the environment semantic knowledge base.
[0015] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. The embodiments are only some of the embodiments of the present application, rather than all the embodiments, and the terms "comprise" and "have" and any variants thereof are intended to cover non-exclusive inclusion.
[0016] Embodiment one, as shown in the figure, an information body construction method of an environment understanding type service robot, comprising the following steps, Figure 1
[0017] S1. acquiring multi-source environment data collected in the running process of the service robot, dividing the multi-source environment data into different modalities according to the type, and constructing a feature data set of each perceivable object in the environment under each modality based on the division result;
[0018] Specifically, during the execution of the task by the service robot, multi-source environmental data in the environment is collected in real time, including visual image data, depth point cloud data, spatial positioning data, acoustic data, contact pressure data, and environmental parameter data. The visual image data is collected by an RGB camera installed on the robot, with a set collection frame rate of 30 frames per second and an image resolution of 1920x1080 pixels. The depth point cloud data is collected by a depth camera, with a point cloud output frequency set to 10 times per second, and a single frame of point cloud containing three-dimensional coordinate information of objects in the environment. The spatial positioning data includes inertial measurement unit output and odometer data, with an inertial measurement unit sampling frequency set to 100 Hz and an odometer update frequency set to 10 Hz. The acoustic data is collected by a microphone array with a sampling rate of 48 kHz. The contact pressure data is collected by force / pressure sensors installed on the robot arm or end effector, with a sampling frequency set to 100 Hz. The environmental parameter data includes temperature, humidity, light intensity, and air quality parameters, wherein the air quality parameters specifically include concentration, PM2.5 concentration, VOC concentration; temperature and humidity are sampled by digital temperature and humidity sensors with a sampling interval set to 1 time per second; light intensity is sampled by a light sensor with a sampling interval set to 1 time per second; air quality parameters are sampled by corresponding gas sensors with a sampling interval set to 1 time per second. The collected multi-source environmental data is preprocessed and used in subsequent environmental understanding process. The preprocessing includes denoising processing of various types of data to eliminate random interference generated during sampling, normalization processing of data amplitude to eliminate dimensional differences between different data types, and time alignment operation according to each data sampling timestamp. Through the time synchronization mechanism, different types of data are unified into the same time window, thereby ensuring that multi-source data reflects consistent environmental state at the same time,
[0019] After the preprocessing and time synchronization of multi-source environmental data are completed, the processed multi-source environmental data are divided by type into visual modalities, depth modalities, spatial modalities, acoustic modalities, tactile modalities, and environmental modalities, each of which is used to represent different attribute characteristics of a perceivable object in the environment. The visual modalities represent the appearance attributes of the object; the depth modalities represent the geometric attributes of the object; the spatial modalities represent the position information and motion state of the object; the acoustic modalities represent the acoustic characteristics of the object; the tactile modalities represent the contact attributes of the object; and the environmental modalities represent the background parameters of the environment in which the object is located. After the modalities are divided, corresponding preprocessing operations are performed on each modality data: the visual modalities are subjected to distortion correction and image denoising to improve image clarity and feature extraction accuracy; the depth modalities are subjected to point cloud filtering and ground segmentation to remove invalid points and extract effective spatial structures; the spatial modalities are subjected to coordinate system transformation and odometer error correction to ensure the global consistency of the position information; the acoustic modalities are subjected to beamforming and energy spectrum analysis to enhance the identification ability of the target sound source; the tactile modalities are subjected to force signal filtering and contact feature decomposition to eliminate high-frequency noise and identify stable contact features; and the environmental modalities are subjected to parameter smoothing and outlier removal to ensure the continuity and stability of the environmental parameter changes. The preprocessed modality data are uniformly stored and synchronized to the same time reference,
[0020] After the multi-modal division and feature preprocessing are completed, object recognition and feature extraction are performed on the divided modal data within the same time window in which the service robot performs the task. In the visual modality, the position area of each perceivable object in the environment is recognized through image segmentation or edge detection, and two-dimensional appearance features including color distribution, texture structure, and edge contour information are extracted to form a visual feature vector of the object. In the depth modality, three-dimensional object boundary and spatial shape features of the object are extracted using point cloud clustering or contour detection to generate a three-dimensional geometric feature vector of the object. In the spatial modality, the position, pose, and motion state of the perceivable object in the global coordinate system are recognized through the attitude information provided by the inertial measurement unit and the displacement data measured by the odometer, and spatial position features, pose features, and motion features are extracted through time series analysis, velocity calculation, and acceleration calculation to form a spatial feature vector of the object. In the acoustic modality, the spatial position of the sound source is recognized through a sound source positioning algorithm, and an acoustic feature vector is extracted, including spectral distribution, sound pressure level, and time domain envelope features. In the tactile modality, the object in contact with the robot is detected through force / pressure signal change detection, and the contact force, contact distribution, and contact duration are calculated to form a tactile feature vector of the object through force / pressure signal mapping. In the environmental modality, environmental parameter data are collected through temperature, humidity, light, and gas sensors, the instantaneous state of each parameter is recognized, and based on the time series of sliding average filtering and outlier rejection processing, numerical features, statistical features, and change features of the environmental parameters are calculated to form an environmental feature vector that can represent the environmental state and the potential influence on the perceivable object. In the process of extracting modal features, a uniform identifier is assigned to each perceivable object, and the feature vector of the object in each modality is recorded to form a single-modal feature data set of each object in each modality.
[0021] S2. The single-modal feature data set of each perceivable object in each modality is cross-modally associated, and based on the cross-modal association result, the modal feature data is fused to construct an object-level information body of each perceivable object, which is used to represent the class, spatial position, and attributes of each perceivable object.
[0022] Specifically, first, a corresponding feature vector is extracted from each single-modal data set, including a visual feature vector, a three-dimensional geometric feature vector, a spatial feature vector, an acoustic feature vector, and an environmental feature vector, to ensure that high-dimensional information of the features accurately represents the attributes of each perceivable object. The single-modal feature vector data is synchronized with the timestamps of each modality to ensure that the features of the same object in different modalities can be accurately aligned. Meanwhile, spatial position correction is performed on the modality feature data that has a spatial deviation, so that all modality feature data can be matched in the same coordinate system. The processed single-modal feature vector data is matched by using a cross-modal matching algorithm to correspondingly match the feature data from different modalities, to establish a corresponding relationship between the object features of each modality, and the corresponding relationship between the object features of each modality is determined for consistency. When the determination result meets a preset consistency condition, a modality feature corresponding combination is determined, and the modality feature corresponding combinations are collected as a cross-modal correlation result. Then, each modality feature vector after cross-modal correlation is extracted from the cross-modal feature corresponding combinations between each modality of the cross-modal correlation result, and is marked as a modality correlation feature vector, including a visual correlation feature vector, a three-dimensional geometric correlation feature vector, a spatial correlation feature vector, an acoustic correlation feature vector, a tactile correlation feature vector, and an environmental correlation feature vector, and each modality correlation feature vector is normalized. The consistency scores of each modality correlation feature vector in all cross-modal feature combinations are counted, the average consistency score of the cross-modal feature corresponding combination corresponding to each modality correlation feature vector is taken as the consistency value of the modality correlation feature vector, the consistency values of the modality correlation feature vectors are normalized to determine the weights of the modality correlation feature vectors. Based on the modality correlation feature vectors and their corresponding weights, the linear combination of the correlation feature vectors of each object in each modality is performed according to the weights to obtain a fusion feature vector of the object, and the fusion feature vectors of the objects are collected to form an object fusion feature set. Finally, for the category features, the matching between each object fusion feature vector and the fusion feature vector of each preset object category template is performed by calculating the Euclidean distance between them. The smaller the vector distance, the higher the similarity. When the vector distance is less than the category matching threshold, it is determined that the object matches the category template, and the category feature of the object is determined. The sources of each preset object category template are obtained based on the statistical results of the historical multi-modal data, and the source of the category matching threshold is that the training samples corresponding to the category of the preset object category template are selected, the similarity between the fusion feature vector of the training sample and the fusion feature vector of the corresponding preset object category template is calculated, the similarity set of all training samples in the category is obtained, the set is sorted according to the similarity, and the 95th percentile of the similarity distribution is taken as the category matching threshold of the object.For the spatial position feature and the attribute feature, components reflecting the spatial position and the object attribute are extracted from the fusion feature vector of each object, and are taken as the spatial position feature and the attribute feature respectively; the category feature, the spatial position feature and the attribute feature of each perceptible object are combined as corresponding fields according to an information organization format, to form an object-level information body of the object, so as to uniformly represent the category, the spatial position and the attribute information of the object, and the information organization format includes a category field for identifying the object, a coordinate field for recording the spatial position of the object and an attribute field for representing the object characteristics.
[0023] Further, the cross-modal association process includes the following steps,
[0024] S200. The single-modal feature data set of each perceptible object in each modality is analyzed, and different modal features belonging to the same perceptible object are correspondingly matched, so as to establish the object feature correspondence relationship between modalities, specifically,
[0025] In the visual modality, the three-dimensional geometric features of the object in the depth modality are mapped to the visual image coordinate system by using the spatial re-projection method and the object region mapping method, and the object feature correspondence relationship between the visual modality and the depth modality is established; in the tactile modality, the surface topology mapping method and the coordinate alignment method are used to correspondingly match the contact pressure distribution, the surface deformation feature perceived by the tactile perception and the three-dimensional point cloud geometric structure in the spatial modality, and the object feature correspondence relationship between the tactile modality and the spatial modality is established; in the acoustic modality, the sound source positioning algorithm, the time-space alignment method and the spatial positioning method are used to respectively establish the object feature correspondence relationship between the acoustic modality and the visual modality, the depth modality and the spatial modality.
[0026] S201. The object feature correspondence relationship between modalities is respectively determined for consistency, and the consistency score of the corresponding combination of modal features is calculated, when the determination result meets the preset consistency condition, the corresponding modal feature combination is generated, and the set of cross-modal feature combinations between modalities forms the cross-modal association result, and the cross-modal association result is determined to include,
[0027] a. For the determination process of the corresponding combination of features of the visual and depth modalities, firstly, the three-dimensional geometric feature vector of the object in the depth modality is mapped to the visual image coordinate system by the spatial re-projection method, so as to determine the corresponding object region in the visual image, realize the object region mapping mode, and obtain the spatial position matching degree, the contour shape matching degree and the texture structure direction matching degree by comparing the spatial position correspondence, the boundary contour overlap degree and the surface structure direction distribution of the mapping region and the object region in the visual image. Since the spatial position error directly affects the accuracy of object positioning, the contour shape error is for object boundary recognition, and affects the object appearance judgment, and the texture structure error has less influence on the overall object appearance judgment, therefore, the spatial position matching degree, the contour shape matching degree and the texture structure direction matching degree are respectively assigned weights of 0.5, 0.3 and 0.2;
[0028] Then, the three indicators are normalized so that the values of each indicator are in the range of 0 to 1, the normalized indicator values are multiplied by the corresponding weight values respectively, and the weighted sum of the product results is obtained, so as to obtain the corresponding combination consistency score. The closer the score value is to 1, the higher the consistency of the two modal features in the spatial position, the contour shape and the texture structure direction is.
[0029] Finally, the corresponding combination consistency score is compared with its preset threshold value. When the corresponding combination consistency score is higher than its preset threshold value, the feature of the cross-modal association is regarded as the corresponding combination of features of the visual and depth modalities. The preset threshold value of the corresponding combination consistency score is obtained by statistical analysis of a large number of labeled sample data. For 500 samples, the feature data thereof is obtained under the visual and depth modalities, and the corresponding combination consistency score of each object is calculated. According to the probability density distribution analysis of the corresponding combination consistency score, the distribution approximately obeys the normal distribution, the mean value is 0.853, and the standard deviation is 0.079. The mean value of the high consistency matching sample score is concentrated in the interval of 0.8 to 1.0, and the mean value of the low consistency matching sample is concentrated in the interval of 0.3 to 0.6. There is an obvious boundary between them. According to the normal distribution characteristics, about 95% of the real matching object consistency scores are above the position of the mean value minus 1.64 times the standard deviation. The preset threshold value of the corresponding combination consistency score is set to 0.75,
[0030] b. For the determination process of the corresponding combination of the features of the haptic modality and the space modality, first, based on the contact distribution information in the haptic feature vector, the contact point perceived by the haptic modality is determined, and the contact point is mapped in the space geometric coordinate system through the coordinate alignment method to obtain the corresponding position of the haptic contact point in the space geometric coordinate system, and the surface topological mapping method is adopted to compare the differences of the contact position, contact pressure direction and local surface curvature of the corresponding points of the haptic modality and the space modality in the same object region, and the contact position matching degree, contact direction matching degree and surface curvature matching degree are obtained respectively. Among them, since the contact position error directly affects the registration accuracy of the haptic and the space geometry, and the contact direction error affects the judgment of the contact surface attitude, the influence degree on the overall object positioning is moderate, and the surface curvature error reflects the consistency of the local surface form, and the influence on the overall registration is small, therefore, the contact position matching degree, the contact direction matching degree and the surface curvature matching degree are respectively assigned weights of 0.5, 0.3 and 0.2;
[0031] Then, the three indicators are normalized, each normalized indicator is multiplied by the corresponding weight value, and the weighted sum of the product results is obtained to obtain the corresponding combination consistency score;
[0032] Finally, the corresponding combination consistency score is compared with its preset threshold value, when the corresponding combination consistency score is higher than its preset threshold value, the feature of the cross-modal association is taken as the corresponding combination of the features of the haptic modality and the space modality, wherein the preset threshold value of the corresponding combination consistency score is obtained by statistical analysis on a large number of labeled sample data, 500 samples are selected, the feature data thereof is obtained under the haptic modality and the space modality, and the corresponding combination consistency score of each object is calculated. According to the probability density distribution analysis of the consistency score, the result is approximately normally distributed, the discrimination between the mean values of the high consistency matching samples and the low consistency matching samples is relatively small, mainly concentrated in the interval of 0.6 to 0.8, the mean value is 0.762, and the standard deviation is 0.107. Since the haptic perception is affected by the contact attitude, surface friction and local deformation, the discrimination of the corresponding combination consistency score is relatively reduced. Based on the statistical characteristics, in order to ensure the effective matching recall rate while avoiding excessive misjudgment, the range of the corresponding combination consistency score lower than the mean value minus 0.7 times the standard deviation is taken as the acceptable threshold limit, and the preset threshold value of the corresponding combination consistency score is set to 0.68,
[0033] c.For the determination process of the corresponding combination of the acoustic modal and the visual modal, first, the acoustic features of the object in the acoustic modal are mapped to the visual image coordinate system by the sound source positioning algorithm to determine the corresponding object region in the visual image. By comparing the mapping region and the object region in the visual image in terms of spatial position, contour boundary and dynamic feature distribution, the spatial position matching degree, the contour shape matching degree and the dynamic feature matching degree are obtained respectively. Since the spatial position error directly affects the accuracy of object positioning, the contour shape error affects the object boundary recognition, and the dynamic feature error has relatively small influence on the overall object judgment, the spatial position matching degree, the contour shape matching degree and the dynamic feature matching degree are respectively assigned weights of 0.5, 0.3 and 0.2;
[0034] Then, the three indicators are normalized, each normalized indicator is multiplied by the corresponding weight value, and the product is weighted and summed to obtain the corresponding combination consistency score;
[0035] Finally, the corresponding combination consistency score is compared with its preset threshold value. When the corresponding combination consistency score is higher than its preset threshold value, the cross-modal related features are taken as the corresponding combination of the acoustic modal and the visual modal. The preset threshold value of the corresponding combination consistency score is obtained by statistical analysis of a large number of labeled sample data. For 500 samples, their feature data are obtained under the acoustic modal and the visual modal, and the corresponding combination consistency score of each object is calculated. According to the probability density distribution analysis of the consistency score, the distribution approximately obeys the normal distribution, the mean is 0.782, and the standard deviation is 0.089. The high consistency matching sample score is mainly concentrated in the interval of 0.7 to 0.9, and the low consistency matching sample score is mainly concentrated in the interval of 0.5 to 0.7. There is an obvious boundary between them. In order to control the mis-matching ratio within 5% in practical application, the preset threshold value of the corresponding combination consistency score is set to 0.72;
[0036] d. Regarding the process of determining the feature correspondence combination between acoustic and depth modes, firstly, the acoustic features are mapped to the three-dimensional geometric feature vectors of objects in the depth mode using a time-space alignment method. This determines the correspondence between the corresponding object region in the depth mode space and the time window of the acoustic event. Based on the mapped correspondence, for the features of the acoustic and depth modes within the same object region, the differences between the acoustic energy distribution, spectral feature pattern, and sound source direction features and the spatial volume distribution, surface contour, and geometric center position of the object in the depth mode are calculated. This yields the acoustic-depth spatial center matching degree, acoustic-depth contour matching degree, and acoustic-depth spectral pattern matching degree. Since spatial position matching directly affects the spatial correlation of acoustic-depth features, contour matching has a moderate impact on object recognition, while spectral pattern matching has a small impact on overall consistency, the acoustic-depth spatial center matching degree, acoustic-depth contour matching degree, and acoustic-depth spectral pattern matching degree are assigned weights of 0.5, 0.3, and 0.2, respectively.
[0037] Then, the three indicators are normalized, each normalized indicator is multiplied by its corresponding weight, and the product results are weighted and summed to obtain the consistency score of the corresponding combination.
[0038] Finally, the consistency score of the corresponding combination is compared with its preset threshold. When the consistency score of the corresponding combination is higher than the preset threshold, the cross-modal associated feature is taken as the feature corresponding combination of acoustic mode and depth mode. The preset threshold of the consistency score of the corresponding combination is obtained by statistical analysis of a large number of labeled sample data. 500 samples are selected, and their feature data are obtained in acoustic mode and depth mode respectively. The consistency score of the corresponding combination for each object is calculated. According to the probability density distribution analysis of the consistency score, the result approximately follows a normal distribution. The discrimination between the mean scores of high consistency matching samples and low consistency matching samples is relatively small, mainly concentrated in the range of 0.62 to 0.80, with a mean of about 0.74 and a standard deviation of about 0.10. In order to avoid too many false positives while ensuring effective matching recall, the range of consistency score below the mean minus 0.6 times the standard deviation is taken as the acceptable threshold. The preset threshold of the consistency score of the corresponding combination is set to 0.68.
[0039] e. For the determination process of the corresponding combination of the acoustic modal and the spatial modal, first, the acoustic features of the object in the acoustic modal are mapped to the spatial geometric coordinate system by a spatial positioning method to determine the corresponding position area of the sound source in the space, and according to the mapping result, the differences in the spatial position, sound pressure direction and spatial distribution characteristics of the corresponding points of the acoustic modal and the spatial modal in the same object area are compared to obtain the spatial position matching degree, the sound pressure direction matching degree and the spatial distribution matching degree. Since the spatial position matching degree most directly affects the object positioning accuracy, the sound pressure direction matching degree affects the accurate judgment of the object sound source direction, and the overall registration is generally affected, and the spatial distribution matching degree has less influence on the overall object contour and positioning, therefore, the spatial position matching degree, the sound pressure direction matching degree and the spatial distribution matching degree are respectively assigned weights of 0.5, 0.3 and 0.2;
[0040] Then, the three matching degree indexes are normalized, each normalized index is multiplied by the corresponding weight, and the weighted sum of the product results is obtained to obtain the corresponding combination consistency score,
[0041] Finally, the corresponding combination consistency score is compared with its preset threshold value, when the corresponding combination consistency score is higher than the preset threshold value of the corresponding combination consistency score, the cross-modal related features are taken as the corresponding combination of the acoustic modal and the spatial modal, wherein the preset threshold value of the corresponding combination consistency score is obtained by statistical analysis on a large number of labeled sample data, 500 samples are selected, their feature data are obtained under the acoustic modal and the spatial modal, and the corresponding combination consistency score of each object is calculated. According to the probability density distribution analysis, the results approximately obey the normal distribution, the score distinction between high consistency matching samples and low consistency matching samples is relatively small, mainly concentrated in the interval of 0.58 to 0.78, the mean is about 0.71, and the standard deviation is about 0.13. Based on the statistical characteristics, in order to ensure the effective matching recall rate and control the mismatching rate, the range of the consistency score lower than the mean minus 0.25 times the standard deviation is taken as the acceptable threshold, and the preset threshold value of the corresponding combination consistency score is set to 0.68;
[0042] f. For the determination process of the corresponding combination of the environmental modal and the visual modal, first, the environmental feature vector of the object in the environmental modal is mapped to the visual image coordinate system by a spatial alignment method to determine the corresponding object area in the visual image, and the environmental modal features and the visual modal features are compared in the corresponding area to obtain the environmental parameter matching degree, the illumination distribution matching degree and the surface color distribution matching degree. Considering that the environmental parameter has the greatest influence on the overall identification of the object, the illumination distribution has the second greatest influence on the visual performance of the object, and the surface color distribution has less influence on the overall matching, therefore, the environmental parameter matching degree, the illumination distribution matching degree and the surface color distribution matching degree are respectively assigned weights of 0.5, 0.3 and 0.2.
[0043] Then, the three matching degree indexes are normalized respectively, each normalized index is multiplied by a corresponding weight, and the product result is weighted and summed to obtain the corresponding combination consistency score;
[0044] Finally, the corresponding combination consistency score is compared with its preset threshold value, when the corresponding combination consistency score is higher than its preset threshold value, the cross-modal associated feature is taken as the corresponding combination of the environmental modal and the visual modal feature, wherein the preset threshold value of the corresponding combination consistency score is obtained by statistical analysis on a large number of labeled sample data, 500 samples are selected, the feature data thereof is obtained under the environmental modal and the visual modal, and the corresponding combination consistency score of each object is calculated, according to the consistency score probability density distribution analysis, the result approximately obeys the normal distribution, the mean is 0.812, and the standard deviation is 0.091, the scores of high consistency matching samples are concentrated in the interval of 0.75 to 0.95, and the scores of low consistency matching samples are concentrated in the interval of 0.5 to 0.7, based on the statistical characteristics, in order to control the mis-matching proportion within 5% while ensuring the effective matching recall rate, the range of the consistency score lower than the mean minus about 1 times the standard deviation is taken as the acceptable threshold limit, and the preset threshold value of the corresponding combination consistency score is set to 0.72;
[0045] g.For the confirmation process of the corresponding combination of the environmental modal and the depth modal, first, the object three-dimensional geometric feature vector obtained by point cloud clustering or contour detection in the depth modal is directly mapped to the environmental feature vector of the environmental modal through a spatial registration method, the object three-dimensional region within the environmental perception range is determined, and for the mapping region, comparison and analysis are respectively performed from the energy distribution correspondence, the reflection response correspondence and the spatial form correspondence to obtain the energy response matching degree, the surface reflection matching degree and the spatial form matching degree, the correspondence degree between the environmental energy distribution and the object geometric structure has the greatest influence on the overall matching result, the consistency of the optical feature response plays an important role in the surface feature alignment, the spatial form reflects the coincidence degree of the spatial geometric contour between the two modalities, and has a relatively small influence on the overall matching result, and the weight is 0.20, therefore, the energy response matching degree, the surface reflection matching degree and the spatial form matching degree are respectively assigned weights of 0.45, 0.35 and 0.20;
[0046] Then, the three matching degree indexes are normalized respectively, each normalized index is multiplied by a corresponding weight, and the product result is weighted and summed to obtain the corresponding combination consistency score;
[0047] Finally, the corresponding combination consistency score is compared with its preset threshold value, when the corresponding combination consistency score is higher than its preset threshold value, the feature of the cross-modal association is taken as the corresponding combination of the features of the environmental modality and the depth modality, wherein the preset threshold value of the corresponding combination consistency score is obtained by statistical analysis on a large number of labeled sample data, 500 samples are selected, the feature data thereof is obtained under the environmental modality and the depth modality respectively, and the corresponding combination consistency score of each object is calculated, according to the analysis result of the probability density distribution of the consistency score, it is known that the distribution approximately obeys the normal distribution, the mean value is 0.787, and the standard deviation is 0.094, the high consistency matching sample score is mainly concentrated in the interval of 0.75 to 0.9, and the low consistency matching sample score is concentrated in the interval of 0.5 to 0.7, based on the statistical characteristics, in order to ensure the true matching sample recognition rate and avoid too high false matching rate, the range of the consistency score lower than the mean value minus 0.7 times the standard deviation is taken as the acceptable threshold, and the preset threshold value of the corresponding combination consistency score is set to 0.72;
[0048] h. For the confirmation process of the feature corresponding combination of the environmental modality and the spatial modality, first, the environmental feature vector of the object in the environmental modality is mapped to the spatial geometric coordinate system, so that each parameter in the environmental feature vector is mapped to the corresponding position in the spatial geometric structure, and a spatial mapping-spatial alignment combination method is used to determine the object region corresponding to the environmental feature vector in the spatial geometric coordinate system. By comparing the corresponding relationship between the environmental parameter distribution in the region and the surface morphology feature in the spatial geometric structure, the environmental action direction matching degree, the environmental distribution gradient matching degree and the environmental distribution gradient matching degree are obtained. Since the environmental action direction matching degree can directly reflect the sensitive response of the environmental change to the object surface stress and illumination direction, it has the greatest influence on the multi-modal fusion result, the environmental distribution gradient matching degree reflects the responsiveness of the environmental change to the local difference of the object surface, and it has a medium influence on the multi-modal fusion result, and the spatial position matching degree mainly reflects the spatial alignment degree of the object region, and it has a relatively small influence on the multi-modal fusion result, therefore, weights of 0.4, 0.35 and 0.25 are respectively assigned;
[0049] Then, the three matching degree indexes are normalized respectively, each normalized index is multiplied by the corresponding weight, and the weighted sum of the product results is obtained to obtain the comprehensive matching score of the corresponding combination;
[0050] Finally, the corresponding combination consistency score is compared with its preset threshold value, and when the corresponding combination consistency score is higher than the preset threshold value, the feature corresponding combination of the cross-modal correlation is taken as the feature corresponding combination of the environmental modal and the spatial modal, wherein the preset threshold value of the corresponding combination consistency score is obtained by statistical analysis of a large number of labeled sample data. Specifically, 500 samples are selected, and feature data thereof is obtained under the environmental modal and the spatial modal, and the corresponding combination consistency score of each object is calculated. According to the probability density distribution analysis result of the consistency score, it is known that the distribution approximately obeys the normal distribution, the high consistency matching sample score is concentrated in the interval of 0.75 to 0.95, the low consistency matching sample score is concentrated in the interval of 0.4 to 0.65, and there is an obvious boundary between the two. Based on the statistical characteristics, in order to control the false matching rate within 5% while ensuring the effective matching recall rate, the range of the consistency score lower than the mean minus about 1.64 times the standard deviation is taken as the acceptable threshold, and the preset threshold value of the corresponding combination consistency score is set to 0.70.
[0051] Further, the process of fusing the feature data of each modal based on the cross-modal correlation result comprises,
[0052] S210. From the cross-modal feature corresponding combination between each modal of the cross-modal correlation result, the cross-modal correlation feature vector of each modal is extracted, and is marked as the modal correlation feature vector, which is used for modal feature fusion. Specifically,
[0053] From the cross-modal feature corresponding combination between each modal of the cross-modal correlation result, the cross-modal correlation feature vector of each modal is extracted, which is used for modal feature fusion, and each modal feature vector is normalized to keep the feature dimension consistent, forming a fusion feature data set of each object under each modal. Specifically, the cross-modal correlation visual feature vector is extracted from the feature corresponding combination of the visual modal and the depth modal, the acoustic modal and the visual modal, and the environmental modal and the visual modal, and is marked as the visual correlation feature vector V i The cross-modal correlation three-dimensional geometric feature vector is extracted from the feature corresponding combination of the visual modal and the depth modal, the acoustic modal and the depth modal, and the environmental modal and the depth modal, and is marked as the three-dimensional geometric correlation feature vector G i The cross-modal correlation spatial feature vector is extracted from the feature corresponding combination of the tactile modal and the spatial modal, the acoustic modal and the spatial modal, and the environmental modal and the spatial modal, and is marked as the spatial correlation feature vector S i The cross-modal correlation acoustic feature vector is extracted from the feature corresponding combination of the acoustic modal and the visual modal, the acoustic modal and the depth modal, and the acoustic modal and the spatial modal, and is marked as the acoustic correlation feature vectorA i ; extract the cross-modal associated tactile feature vector from the characteristic corresponding combination of the tactile modality and the spatial modality, marked as the tactile associated feature vector T i ; extract the cross-modal associated environment feature vector from the characteristic corresponding combination of the environment modality and the visual modality, the environment modality and the depth modality, and the environment modality and the spatial modality, marked as the environment associated feature vector E i , wherein i indicates the serial number of the object, in order to ensure the comparability of different modal characteristics, the extracted modal feature vectors are respectively normalized by the z-score standardization method, so as to eliminate the influence of the dimension and numerical range difference of each modal characteristic, so that the characteristic values of each modality are located in the same range, and the characteristic dimension is consistent;
[0054] S211. Based on the consistency score of the cross-modal characteristic corresponding combination between each modality corresponding to each modal associated feature vector, the weight of each modal associated feature vector is determined. Specifically, for each modal associated feature vector, the consistency score in all involved cross-modal characteristic combinations is counted. The higher the consistency score, the higher the contribution of the modality in the object representation. The average value of the consistency score of the cross-modal characteristic corresponding combination corresponding to each modal associated feature vector is calculated to obtain the consistency value of each modal associated feature vector. For the consistency value of the visual associated feature vector , the average value of the consistency score of the characteristic corresponding combination of the visual modality and the depth modality, the acoustic modality and the visual modality, and the environment modality and the visual modality is calculated as the consistency value; for the consistency value of the three-dimensional geometric associated feature vector , the average value of the consistency score of the characteristic corresponding combination of the visual modality and the depth modality, the acoustic modality and the depth modality, and the environment modality and the depth modality is calculated as the consistency value; for the consistency value of the spatial associated feature vector , the average value of the consistency score of the characteristic corresponding combination of the tactile modality and the spatial modality, the acoustic modality and the spatial modality, and the environment modality and the spatial modality is calculated as the consistency value; for the consistency value of the acoustic associated feature vector , the average value of the consistency score of the characteristic corresponding combination of the acoustic modality and the visual modality, the acoustic modality and the depth modality, and the acoustic modality and the spatial modality is calculated as the consistency value; for the consistency value of the tactile associated feature vector , the consistency score value of the characteristic corresponding combination of the tactile modality and the spatial modality is taken as the consistency value; for the consistency value of the environment associated feature vector The consistency score is calculated as the average of the consistency scores for feature correspondence combinations between environmental modality and visual modality, environmental modality and depth modality, and environmental modality and spatial modality. The consistency scores of the feature vectors associated with each modality are then proportionally normalized to ensure that the sum of the weights of the feature vectors associated with each modality is 1, thus obtaining the weights of the feature vectors associated with each modality.
[0055] The process of normalizing the consistency values of the associated feature vectors of each modality is as follows:
[0056] ,in, The weights of the modal association feature vectors are represented by , where i represents the index of the modal association feature vector. The weights of the visual association feature vectors are calculated. 3D geometric correlation feature vector weights Spatial correlation feature vector weights Acoustic correlation feature vector weights tactile association feature vector weights Environmental correlation feature vector weights Finally, a set of modal association feature vector weights is formed, which serves as the weight basis for fusing the feature data of each modality, so that each modal association feature vector participates in the calculation according to the weight ratio during fusion.
[0057] S212. Based on the modal correlation feature vectors and their corresponding weights, the modal correlation feature vectors are weighted and calculated, and inter-modal fusion processing is performed to obtain the fused feature vector. Specifically,
[0058] Based on the modal feature vectors and their corresponding weights, the feature vectors of each modality are calculated, and inter-modal fusion processing is performed. The feature vectors of each object in each modality are linearly combined according to their weights to obtain the fused feature vector of the object. ,in, The fusion feature vectors of each object are organized according to the object identifier to form an object fusion feature set.
[0059] S3. Assign a lifecycle identifier to the object-level information body, and manage the state of the object-level information body based on a timestamp, so that the object-level information body can be dynamically updated as the environment changes.
[0060] Specifically, for each existing object-level information body Assign a unique lifecycle identifier The allocation method can employ either an auto-incrementing number or a unique hash value generation method to ensure that each object-level information body corresponds to a unique identifier. Create a status field and a time field, and associate them with their lifecycle identifiers. Binding, where the status field records the current survival state of the object-level information body, initially set to active state, and the status types include active state, tracking state, pending confirmation state, and invalid state; the time field records the timestamp of the most recent update of the object-level information body, initially set to the object-level information body's timestamp. The generation time will be used as a lifecycle identifier. Its status field, time field, and object-level information body Binding, forming a binding information between the object-level information body and the lifecycle identifier.
[0061] When the characteristic information of the current object in the environment is obtained, the timestamp of the object's characteristic is extracted and matched with its corresponding object-level information body. The time field is compared; if the time difference is less than the preset state update time threshold, the object-level information body is determined to be within the valid update cycle, and the state of the object-level information body is managed. Based on the acquisition of the current object's feature information within the preset window time, the object-level information body is... The status field is updated, and the timestamp is also updated to the object-level information body. The updated status field and time field will be compared with the object-level information body. Lifecycle identifier To maintain the binding relationship, the object-level information can be dynamically updated as the environment changes. The preset state update time threshold is determined based on the sampling frequency of the multi-source environmental data. This threshold must be greater than the longest sampling period among the multi-source environmental data to ensure that the state update of the object-level information covers the complete sampling period of each modality. In this embodiment, considering the multi-source environmental data collection, the sampling interval for temperature, humidity, light intensity, and air quality parameters is 1 second. Taking into account hardware data transmission and response latency factors, the preset state update time threshold is set to 1.1 seconds.
[0062] Furthermore, the specific process of managing the state of object-level information includes,
[0063] When the feature information of the current object is acquired in multiple consecutive frames within a preset window time, the object-level information body will be... The status field is updated to active status, and its time field is updated to the current timestamp; when only some frames out of multiple consecutive frames within a preset window time are obtained, the object-level information body is updated. The status field is updated to tracking status, and its time field is updated to the timestamp of the frame corresponding to the most recent successful acquisition of the current object's feature information within the preset window time; if the feature information of the current object is not acquired for multiple consecutive frames within the preset window time, the object-level information body is updated. the state field of the object-level information body is updated to a to-be-confirmed state, and the time field is kept as the time stamp when the object successfully acquires the feature information last time within the preset window time; when the feature information of the current object is not acquired within the preset window time, the object-level information body the state field of the object-level information body is updated to an invalid state, and the time field is updated to the time stamp when the invalidity is determined,
[0064] In this embodiment, considering the sampling frequency difference of multi-source perception data in the environment and the possible transient frame loss, the preset window time is set to 3 times of the longest sampling period in the multi-source perception data, that is, 3s, and the continuous frame number is set according to the frame number proportion of the visual modal in the multi-source environmental data within the preset window time, because the visual modal can provide the most stable target continuous observation. In this embodiment, the frame rate of the visual image data acquisition is 30 frames / s, when the frame number proportion of the current object feature information acquired continuously within the preset window time is greater than or equal to 90% of the total frame number, the state field of the object-level information body is updated to an active state; when the proportion is greater than or equal to 30% and less than 90%, the state field of the object-level information body is updated to a tracking state; when the proportion is greater than 0 and less than 30%, the state field of the object-level information body is updated to a to-be-confirmed state; when no object feature information frame is acquired within the preset window time, the state field of the object-level information body is updated to an invalid state.
[0065] S4. Based on the dynamically updated object-level information body, the spatial relationship, functional association and interaction relationship between each perceivable object in the environment are analyzed, and a relationship-level information body is generated based on the analysis result, which is used to represent the relative position, functional contact and interaction characteristics between objects in the environment,
[0066] Specifically, after the dynamic update of the object-level information body, the spatial relationship, functional association and interaction relationship between each perceivable object in the environment are analyzed, and the object-level information body of each perceivable object in the environment is combined according to its life cycle identifier to form a plurality of object-level information body pairs. The spatial relationship of the object-level information body pair is calculated according to the spatial position characteristics of the object-level information body, the functional relationship of the object-level information body pair is judged according to the category characteristics, and the interaction relationship of the object-level information body pair is judged through the interaction condition between the spatial relationship and the functional association. The spatial relationship, functional association and interaction relationship of the object-level information body pair are combined to generate a relationship-level information body. The relationship-level information body clearly records the spatial relationship, functional contact and interaction characteristics of the associated object-level information body, and the generated relationship information body is clearly bound to the two object-level information bodies of the object-level information body pair to form an association relationship record.
[0067] Further, the spatial relationship, functional association and interaction relationship between each perceivable object in the environment are analyzed, specifically including,
[0068] S400. For each object-level information body of the perceivable objects in the environment, according to the category feature and the spatial position feature, the object-level information body is combined with other object-level information bodies according to the life cycle identification to form an object-level information body pair, i.e., object-level information body — object-level information body , j The same represents the serial number of the object, i ≠ j ;
[0069] S401. According to the spatial position feature of the object-level information body pair, the coordinates of the spatial positions of the object-level information body and the object-level information body are obtained, the relative position parameters between the two, including the three-dimensional Euclidean distance, the azimuth angle and the height difference, are calculated according to the coordinates of the spatial positions of the two, when the calculated relative position parameters meet the preset spatial proximity determination condition, it is determined that the two exist a spatial correlation relationship, and the relative position parameters are recorded as the spatial relationship field of the object-level information body pair; wherein the setting of the spatial proximity determination condition is based on the effective action range of the service robot in the environment, and specifically, the spatial distance threshold is determined according to the ranging accuracy of the depth camera and the average distance of objects in a typical task scenario, taking the ranging accuracy of the depth camera as ±5cm and the typical operation radius of the robot in the indoor environment as 1.5m, the spatial distance threshold is set to 1.0m; the azimuth angle difference threshold is used to distinguish whether the objects are in the same azimuth area, according to the characteristics of the horizontal field of view angle 60° of the depth camera, the azimuth angle difference threshold is set to ±15°; the height difference threshold is used to distinguish the relationship between objects on the same plane or different planes, combined with the vertical movement range 0.5m of the robot operation arm, the height difference threshold is set to 0.3m, when the three-dimensional Euclidean distance of the object-level information body pair is less than 1.0m, and the azimuth angle difference is less than 15° and the height difference is less than 0.3m, it is determined that the object-level information body pair meets the spatial proximity determination condition, and the two exist a spatial correlation relationship;
[0070] S402. According to the category feature and the attribute feature of the object-level information body pair, the functional relationship between the two object-level information bodies is retrieved from a pre-established functional association rule database, and when the retrieval result shows that there is a functional association between the two, the retrieved functional relationship result is recorded as the functional association field corresponding to the object-level information body pair; wherein the establishment of the functional association rule database can be realized based on the task scene rules defined by the existing object semantic knowledge base, specifically, when constructing the functional association rule database, the category feature, the operable attribute and the co-occurrence relationship in the task scene of the object are comprehensively considered, first, according to the category feature of the object-level information body, the category set of the perceivable objects in the environment is determined, the category set is divided according to the functional attribute or the use attribute of the object in the environment, and is used to represent the possible interaction relationship between objects of different functional categories; then, the cooperative use relationship between categories in a typical task scene is statistically analyzed to form a functional association pair, for example, there is a placing relationship between the tableware category and the table category, and the functional association relationship is obtained by a rule extraction method based on the semantic knowledge base, specifically, the semantic association information between object categories in the semantic knowledge base is used to construct a mapping table between the category pair and the functional relationship type, to form the basic data of the functional association rule database;
[0071] S403. Based on the spatial relationship field and the functional association field of the object-level information body pair, according to the category feature, the attribute feature and the spatial position feature of the two object-level information bodies, the possible interaction type between the two object-level information bodies in the object-level information body pair is retrieved from a pre-established interaction rule database, when the retrieval result shows that there is an interaction type between the two, the interaction type and the related constraint information are recorded as the interaction relationship field of the object-level information body pair, and a relationship-level information body is generated, wherein the interaction rule database is a pre-established structured rule set, which is used to describe the possible interaction type and the constraint relationship between objects of different categories, and the establishment of the interaction rule database is based on the following, one is to refer to the experience rules and operation specifications about object interaction in daily life or industrial environment in the prior art; two is to obtain object interaction instances including object contact, operation sequence, placement position and the like based on the historical collected multi-source environment data statistics; three is to combine the object category feature, the spatial position feature and the attribute feature to induce and arrange common interaction modes to form rule entries.
[0072] S5. Based on the object-level information body and the relationship-level information body, the spatial layout and the relationship connection of the object are analyzed to construct a scene-level information body, which is used to reflect the spatial structure, object distribution and semantic features of the whole environment,
[0073] Specifically, first, the environment is divided into a regular grid in a horizontal plane, each grid corresponds to a spatial unit, and objects are mapped to the corresponding unit to form a correspondence between the grid and the object. The connected domain extraction and density-based clustering method are used to identify the object set that is continuous in space and similar in semantics, forming the preliminary environment region. The clustering clusters are merged and segmented in combination with the object category features and category semantic rules, so as to obtain the scene object set reflecting the spatial structure and object distribution characteristics. Then, based on the spatial relationship, functional relationship and interaction relationship of each object-level information body pair in the relationship-level information body, an object relationship network is established, and according to the relationship characteristics of the corresponding object-level information body pair, a relationship label is assigned to each connection edge in the object relationship network, forming an object relationship network composed of network nodes and connection edges with relationship labels. Next, the spatial features of the object relationship network are used to aggregate the object set that meets the spatial proximity, containment condition and functional or interactive association to form a candidate functional region, and the functional type and functional region are determined according to the pre-established functional association rule data and the function and interaction relationship of the object-level information body and the object-level information body pair in the candidate functional region, and the spatial hierarchical relationship between the functional regions is established. Finally, the feature information of the object-level information body in the functional region and the semantic relationship field of the corresponding relationship-level information body are integrated to generate a regional semantic set, and the regional semantic sets and the semantic association between regions are fused under the spatial hierarchical framework to form a complete scene semantic structure, and recorded as a scene-level information body, which is used to represent the overall spatial structure, object distribution and semantic characteristics of the environment.
[0074] Further, the construction of the scene-level information body specifically includes,
[0075] S500. According to the spatial distribution and category characteristics of the object-level information body, the environment is regionally divided, and the related objects in each region are aggregated to form a scene object set, and the specific process is as follows,
[0076] Firstly, the environment space is divided into grid regions of equal size, each grid region corresponds to a space unit, each object is mapped to the corresponding grid unit to form the correspondence relationship between the grid and the object. Specifically, taking the global coordinate system of the robot as the reference, the environment space is regularly divided on the horizontal plane to divide the environment into grid regions of equal size, each grid region corresponds to a space unit, the edge length of each unit is set according to the environment perception accuracy and the object distribution density, and is preferably 0.5 m. According to the spatial coordinates recorded in the object-level information body, each object is mapped to the corresponding grid unit to form the correspondence relationship between the grid and the object. When the object has a certain spatial size, the projection covers multiple grid units, and the object is recorded in the corresponding multiple units. Secondly, according to the grid units occupied by the object-level information body, a set of spatially interconnected grid units is extracted to obtain a plurality of preliminary function candidate regions. The spatial coordinates of the object-level information body contained in each preliminary function candidate region are subjected to density-based clustering processing. Specifically, according to the grid units occupied by the object, a connected domain extraction operation is performed to extract a set of spatially interconnected grid units. Each set of grid units represents a preliminary spatial distribution region, thereby obtaining a plurality of preliminary function candidate regions. The spatial coordinates of the object-level information body contained in each preliminary function candidate region are subjected to density-based clustering processing to further distinguish the object aggregation distribution within the region. Preferably, the DBSCAN clustering method is used, wherein the value of the clustering radius eps is determined according to the typical spatial distance between objects. When the perception accuracy is within about 1 m, the adjacent objects can be stably distinguished, and the clustering radius eps is set to 1.0m; the minimum sample number minPts is set according to the average density of objects in the scene, and when a region contains at least two objects, the minimum sample number minPts is set to 2; next, based on the category features recorded in the object-level information body and the pre-established category semantic rules, the clustering clusters are further merged and segmented to form a plurality of environmental regions, specifically, based on the category features recorded in the object-level information body and the pre-established category semantic rules, the clustering clusters are further merged and segmented, wherein the establishment of the category semantic rules can refer to the object category relationship rules defined in the existing semantic knowledge base, when the main categories of different clustering clusters have semantic association and the spatial distance is lower than the clustering cluster merging threshold, the corresponding clustering clusters are merged into the same region; when a single clustering cluster contains a plurality of semantically unrelated category combinations, the clustering cluster is segmented according to the category semantic rules, wherein the clustering cluster merging threshold is set according to the typical spatial distance between objects and the positioning / perspection accuracy of the robot, for example, when the typical object distance is about 1m, the merging threshold can be set to 1m, forming a plurality of environmental regions; finally, by analyzing the spatial distribution and category statistical characteristics of the member objects in the environmental region, the member object list, the region centroid, the boundary, the main category, the object density and the state statistical field are extracted, and finally the scene object set is formed, specifically, the object list records the life cycle identifier of all object-level information bodies belonging to the environmental region, the region centroid is determined according to the spatial coordinates of the member objects, the boundary information is determined according to the spatial distribution range of the member objects, the main category is determined according to the category distribution of the objects in the region, the object density is determined according to the number of members in the region and the spatial range, and the state statistical field is statistically determined according to the state field of the member objects in the region, including the proportion of objects in the active state, the tracking state and the to-be-confirmed state, finally forming the scene object set, which is used for subsequent scene graph construction and semantic labeling, in order to ensure the timeliness and accuracy of the scene object set, when the spatial position or state field of the object-level information body changes, the local incremental update is performed on the corresponding grid cell and its clustering cluster; if the proportion of the number of updated objects in a certain environmental region to the total number of objects in the environmental region exceeds the global update triggering threshold, the global re-determination of the scene object set is triggered, and the global update triggering threshold is set according to the scene change statistical results in the historical environmental perception data.
[0077] S501. Based on the relationship-level information body between the objects, an object relationship network is constructed to comprehensively represent the spatial structure and semantic connection between the objects, and the specific process is as follows,
[0078] Firstly, a network node is created for each object-level information body, in which the object's lifecycle identifier, as well as the category feature, spatial location feature and attribute feature are recorded; then, the spatial relationship, functional relationship and interactive relationship of each object-level information body pair in the relationship-level information body are obtained, and the spatial proximity relationship, containing relationship or semantic association relationship between the two object-level information bodies are judged, and for the object-level information body pair meeting the conditions, an object relationship network is established, specifically, for each object-level information body pair in the generated relationship-level information body, the spatial relationship, functional relationship and interactive relationship recorded thereby are obtained, the spatial proximity relationship and containing relationship between the two object-level information bodies are determined according to the spatial relationship field, and the semantic association relationship therebetween is comprehensively judged according to the functional relationship field and the interactive relationship field, and for the object-level information body pair meeting the spatial proximity relationship, containing relationship or semantic association relationship condition, a connection edge is established between the corresponding network nodes, and the corresponding relationship type and related constraint parameters are recorded, and when all object-level information body pairs are connected, an object relationship network containing all network nodes and edges thereof is formed, and the judgment rule is that if the three-dimensional Euclidean distance between the two is less than a preset spatial proximity threshold, then a spatial proximity relationship exists, wherein the preset spatial proximity threshold is set according to the aforementioned typical spatial distance between objects and the positioning / sensing accuracy of the robot; if the spatial boundary of one object is completely or partially contained in the spatial boundary of the other object, then a containing relationship exists; if the functional association field or the interactive relationship field indicates that there is a semantic connection between the two, then a semantic association relationship exists; finally, according to the relationship features of the corresponding object-level information body pair, each connection edge in the object relationship network is assigned a relationship label, forming an object relationship network composed of network nodes and connection edges with relationship labels, specifically, for each connection edge in the established object relationship network, according to the relationship features of the corresponding object-level information body pair, a relationship label is assigned to the connection edge, including marking the spatial proximity relationship as adjacent or not adjacent, marking the containing relationship as containing or not containing, marking the semantic association relationship as a specific relationship type according to the function or interaction type, forming an object relationship network composed of network nodes and connection edges with relationship labels.
[0079] S502. Based on the spatial features of the object relationship network, the scene is analyzed for layout, and the functional area distribution and spatial hierarchical structure in the environment are identified, and the specific process is as follows,
[0080] Firstly, adjacent nodes are determined according to the spatial coordinates of each network node in the object relationship network, and a node set that meets the containment relationship and the spatial proximity condition and has a functional or interactive correlation is aggregated to form a candidate functional area, specifically, the spatial coordinates of each network node in the object relationship network are extracted, the three-dimensional Euclidean distance between the network node and other network nodes is calculated, and the nodes with a three-dimensional Euclidean distance less than a preset spatial proximity threshold are determined as adjacent nodes, for each pair of adjacent nodes, the spatial boundary of the corresponding object-level information body is obtained, and it is judged whether there is a containment relationship, if the boundary of an object-level information body is completely or partially located inside the boundary of another object-level information body, the former is marked as a contained node and is associated with the containing node, and a node set that meets the spatial proximity condition and has a functional or interactive correlation is aggregated to form a candidate functional area; then, according to a pre-established functional correlation rule database, the category features of the object-level information bodies in each candidate functional area and the functional relationship fields and the interactive relationship fields of the object-level information body pairs are mapped to corresponding candidate functional types, and the functional type and the corresponding functional area are determined according to the proportion of the object-level information body pairs of each candidate functional type in the object-level information body pairs in the corresponding candidate functional area, specifically, the category features of the object-level information bodies corresponding to all network nodes in each candidate functional area and the functional correlation and the interactive relationship recorded between the nodes are counted, according to the mapping information of the object category pairs and the functional relationship types recorded in the pre-established functional correlation rule database, the category features of the object-level information bodies in each candidate functional area and the functional relationship fields and the interactive relationship fields of the object-level information body pairs are mapped to corresponding candidate functional types, and the proportion of the number of object-level information body pairs of each candidate functional type in the number of all object-level information body pairs in the corresponding candidate functional area is counted, when the proportion of the object-level information body pairs corresponding to a certain candidate functional type meets the functional determination proportion condition, the candidate functional type is determined as the functional type of the candidate functional area, and the candidate functional area is confirmed as a functional area, wherein the functional determination proportion condition can be determined by counting the ratio of the co-occurrence times and the total times of the object-level information body pairs corresponding to different functional types in a typical task scenario; finally, the spatial boundary data corresponding to each functional area is obtained, the spatial hierarchical relationship between the functional areas is established, and a spatial hierarchical structure is formed, specifically, the spatial boundary data corresponding to each functional area is obtained, the three-dimensional Euclidean distance between the candidate functional areas is calculated, if the spatial boundary of a candidate functional area is completely or partially located inside the spatial boundary of another candidate functional area, the former is determined as a sub-area of the latter, and a hierarchical membership relationship is established; if the spatial boundaries of two candidate functional areas are adjacent and there is an interactive relationship recorded in the functional or interactive correlation field, a parallel hierarchical relationship is established, through the establishment of the spatial hierarchy, a spatial hierarchical structure containing parent-child relationship and adjacent relationship is formed, and the structure is recorded in the scene-level information body to represent the hierarchical membership relationship and the spatial distribution of each functional area in the environment.
[0081] S503. Integrating spatial layout results, object attribute characteristics, and semantic relationships, generate scene-level information volumes to represent the overall semantic structure of the environment. This includes the spatial location, hierarchical relationships, and functional types of each functional area within the environment, as well as the spatial distribution and semantic connections of object-level information volumes within those areas. The specific process is as follows:
[0082] First, based on the object-level information bodies within a functional area and their associated relation-level information bodies, a clear correspondence is established between the functional area, the object-level information bodies within the area, and the related relation-level information bodies. Specifically, for each identified functional area, its spatial boundaries and spatial hierarchy are recorded as the spatial layout result of that functional area. All object-level information bodies contained within the functional area are obtained, and the category features, spatial location features, and attribute features recorded in each object-level information body are extracted. The relation-level information bodies associated with the object-level information bodies within the functional area are obtained, and the spatial relationship, functional relationship, and interaction relationship fields recorded in each relation-level information body are extracted. A clear correspondence is established between the functional area, the object-level information bodies within the area, and the related relation-level information bodies, enabling each functional area to fully associate its internal objects and the semantic relationships between objects. Then, the object-level information bodies... Feature information is integrated with the semantic relationship fields of the corresponding relational information body to generate a regional semantic set for the corresponding functional area. This regional semantic set fully represents the semantic associations between objects within the functional area and the distribution of object attributes. Next, based on the intermediate hierarchical structure of the spatial layout results between functional areas, the semantic sets of each region are organized according to the spatial hierarchy to construct the semantic association structure between regions. Finally, within the framework of the spatial hierarchy structure, the semantic sets of each region and the semantic association structure between regions are integrated to ensure that object-level semantics, regional semantics, and spatial hierarchy semantics are consistent, thereby generating a complete scene semantic structure. The complete scene semantic structure is recorded as a scene-level information body, which includes the spatial location, hierarchical relationship, functional type, and spatial distribution and semantic connection of object-level information bodies within each functional area in the environment. This is used to represent the semantic organization, object distribution, and spatial composition relationship of the environment as a whole.
[0083] S6. Unify the storage of object-level information, relation-level information, and scene-level information into an environmental semantic knowledge base, and support service robots in calling it during task execution and environmental decision-making.
[0084] Specifically, after the construction of object-level information body, relationship-level information body and scene-level information body, the three types of information bodies are uniformly arranged and semantically processed. Firstly, the object identification, timestamp, spatial position and semantic label in each type of information body are stored in a structured manner to ensure the consistency of data format among information bodies. Then, the fields related to spatial relationship and category description are uniformly processed according to the preset semantic rules to ensure the consistency and associability of different levels of information bodies. Subsequently, the corresponding relationship fields in the object-level information body and the relationship-level information body are associated according to the hierarchical correspondence among the object-level information body, the relationship-level information body and the scene-level information body, and each functional area in the scene-level information body is corresponded to the object-level information body contained therein, so that the semantic association relationship among the multi-level information bodies is uniformly represented. Finally, the object attributes, the relationship among objects and the scene composition in the environment are stored in an integrated manner in the unified data structure to form an environmental semantic knowledge base. In the process of task execution and environmental decision, the knowledge base content can be searched according to the semantic label or spatial position to extract the corresponding object information, spatial relationship and scene structure, thereby providing semantic support for the behavior planning and environmental understanding of the service robot.
Claims
1. A method for constructing the information body of an environmentally conscious service robot, characterized in that, Includes the following steps, Acquire multi-source environmental data collected during the operation of the service robot, classify the multi-source environmental data into different modalities according to type, and construct a single-modal feature data set for each object in each modality based on the classification results; Cross-modal correlation is performed on the single-modal feature data set of each perceptible object in the environment under each modality, and the feature data of each modality are fused based on the cross-modal correlation results to construct an object-level information body for each perceptible object, which is used to characterize the category, spatial location and attributes of each perceptible object; Assign a lifecycle identifier to the object-level information body and manage the state of the object-level information body based on the timestamp, so that the object-level information body can be dynamically updated as the environment changes. Based on dynamically updated object-level information volumes, we analyze the spatial relationships, functional associations, and interaction relationships among perceptible objects in the environment, and generate relationship-level information volumes based on the analysis results to characterize the spatial relationships, functional relationships, and interaction relationship features of each object in the environment. Based on object-level information bodies and relation-level information bodies, a scene-level information body is constructed by analyzing the spatial layout and relational relationships of objects to reflect the overall spatial structure, object distribution and semantic features of the environment. The object-level information body, relation-level information body, and scenario-level information body are uniformly stored as an environmental semantic knowledge base, which supports service robots in calling it during task execution and environmental decision-making.
2. The information body construction method according to claim 1, characterized in that, The cross-modal association process includes the following steps. Analyze the single-modal feature data set of each perceptible object in each modality, match the different modal features belonging to the same perceptible object, and thus establish the object feature correspondence relationship between each modality. Consistency determination is performed on the correspondence of object features between each modality, and consistency score is calculated for the corresponding combination of modal features. When the determination result meets the preset consistency condition, the corresponding modal feature combination is generated, the corresponding cross-modal feature correspondence combination is determined, and the cross-modal association result is formed by the set of cross-modal feature combinations between each modality.
3. The information body construction method according to claim 2, characterized in that, The process of fusing feature data from various modalities based on cross-modal correlation results includes: Extract the modal feature vectors after cross-modal association from the cross-modal feature correspondence combinations between each modality in the cross-modal association results, and label them as modal association feature vectors for modal feature fusion; The weights of each modality-related feature vector are determined based on the consistency score of the cross-modal feature correspondence combinations between each modality corresponding to each modality-related feature vector. Based on the modal association feature vectors and their corresponding weights, the modal association feature vectors are weighted and calculated, and inter-modal fusion processing is performed to obtain the fused feature vector.
4. The information body construction method according to claim 3, characterized in that, The management of the state of the object-level information body includes updating the state field and time field of the object-level information body based on the acquired feature information of the current object in the environment, specifically including... When the feature information of the current object is obtained in multiple consecutive frames within a preset window time, the status field of the object-level information body is updated to the active state, and its time field is updated to the current timestamp. When the feature information of the current object is obtained in only some frames in a series of consecutive frames within a preset window time, the status field of the object-level information body is updated to the tracking status, and its time field is updated to the timestamp corresponding to the frame in the preset window time where the feature information of the current object was most recently successfully obtained. If the feature information of the current object is not obtained for multiple consecutive frames within the preset window time, the status field of the object-level information body is updated to the pending confirmation state, and the time field is kept as the timestamp of the object when the feature information was most recently successfully obtained within the preset window time. If the characteristic information of the current object is not obtained within the preset window time, the status field of the object-level information body is updated to the failure status, and its time field is updated to the timestamp when the failure was determined.
5. The information body construction method according to claim 4, characterized in that, The spatial relationships, functional associations, and interaction relationships among the perceptible objects in the analysis environment specifically include: For each perceptible object in the environment, the object-level information body is combined with other object-level information bodies in pairs according to its category characteristics and spatial location characteristics and life cycle identifiers to form object-level information body pairs. Based on the spatial location characteristics of the object-level information body pair, obtain the coordinates of the spatial locations of the two object-level information bodies, calculate the relative position parameters between them based on their spatial coordinates, and determine that there is a spatial relationship between them when the relative position parameters meet the preset spatial proximity judgment conditions, and record the relative position parameters as the spatial relationship field of the object-level information body pair. Based on the category and attribute characteristics of the object-level information body pair, the functional relationship between the two object-level information bodies is retrieved from the pre-established functional association rule database. When the retrieval result shows that there is a functional association between the two, the retrieved functional relationship result is recorded as the functional association field corresponding to the object-level information body pair. Based on the spatial relationship field and functional association field of the object-level information body pair, according to the category characteristics, attribute characteristics and spatial location characteristics of the two object-level information bodies, the interaction type between the two object-level information bodies in the pre-established interaction rule database is retrieved. When the retrieval result shows that there is an interaction type between the two, the interaction type and related constraint information are recorded as the interaction relationship field of the object-level information body pair, and a relation-level information body is generated.
6. The information body construction method according to claim 5, characterized in that, The constructed scenario-level information body includes Based on the spatial distribution and category characteristics of object-level information, the environment is divided into regions, and related objects are aggregated in each region to form a scene object set; Based on the information body of relationships between objects, an object relationship network is constructed to comprehensively represent the spatial structure and semantic connections between objects. Based on the spatial characteristics of object relationship networks, we perform layout analysis on scenes to identify the distribution of functional areas and spatial hierarchy in the environment. By integrating spatial layout results, object attribute characteristics, and semantic relationships, a scene-level information body is generated to represent the overall semantic structure of the environment, including the spatial location, hierarchical relationship, functional type of each functional area in the environment, and the spatial distribution and semantic connection of object-level information bodies within the area.
7. The information body construction method according to claim 6, characterized in that, The process of forming the scene object set specifically includes, The environment space is divided into grid regions of equal size, each grid region corresponds to a spatial unit, and each object is mapped to the corresponding grid unit to form a correspondence between grid and object; Based on the grid cells occupied by the object-level information body, extract the set of grid cells that are interconnected in space to obtain several preliminary functional candidate regions, and perform density-based clustering on the spatial coordinates of the object-level information bodies contained in each preliminary functional candidate region. Based on the category features recorded in the object-level information body and the pre-established category semantic rules, the clusters are further merged and segmented to form several environmental regions; By analyzing the spatial distribution and category statistics of member objects within the environment area, the member object list, region centroid, boundary, main category, object density, and status statistics fields are extracted to ultimately form a scene object set.
8. The information body construction method according to claim 7, characterized in that, The construction of the object relationship network specifically includes, Create a network node for each object-level information body. The network node records the object's lifecycle identifier, as well as its category characteristics, spatial location characteristics, and attribute characteristics. Obtain the spatial, functional, and interactive relationships of each object-level information body pair in the relation-level information body, and determine the spatial proximity, inclusion, or semantic association relationships between two object-level information bodies. For object-level information body pairs that meet the conditions, establish an object relationship network. Based on the relational characteristics of the corresponding object-level information pairs, a relation label is assigned to each connecting edge in the object relation network, forming an object relation network composed of network nodes and connecting edges with relation labels.
9. The information body construction method according to claim 8, characterized in that, The functional area distribution and spatial hierarchy structure in the identification environment specifically include: Based on the spatial coordinates of each network node in the object relationship network, neighboring nodes are determined. The set of nodes that simultaneously satisfy the inclusion relationship and spatial proximity conditions and have functional or interactive associations are aggregated to form candidate functional regions. Based on the pre-established functional association rule database, the category characteristics of object-level information bodies in each candidate functional area and the functional relationship fields and interaction relationship fields of object-level information body pairs are mapped to the corresponding candidate functional types. Based on the proportion of object-level information body pairs of each candidate functional type in the corresponding candidate functional area, the functional type and the corresponding functional area are determined. Acquire spatial boundary data corresponding to each functional area, establish spatial hierarchical relationships between functional areas, and form a spatial hierarchical structure.
10. The information body construction method according to claim 9, characterized in that, The construction of the scene-level information body includes, Based on the object-level information body and the related relation-level information body within the functional area, establish a clear correspondence between the functional area, the object-level information body within the functional area, and the related relation-level information bodies. The feature information of the object-level information body is integrated with the semantic relationship fields of the corresponding relation-level information body to generate the regional semantic set of the corresponding functional area; Based on the hierarchical structure of the spatial layout between functional areas, the semantic sets of each area are organized according to the spatial hierarchy to construct the semantic association structure between areas; Within the framework of spatial hierarchy, the semantic sets of each region and the semantic association structure between regions are integrated to generate a complete scene semantic structure, and the complete scene semantic structure is recorded as a scene-level information body.
Citation Information
Patent Citations
Multi-source heterogeneous corpus fusion method and system based on government affair service data
CN120493159A
Decision-making method and device based on multi-modal semantic alignment, equipment and medium
CN120954438A