A Multi-source Sensor Autonomous Intelligent Sensing Method and System in a Complex Dynamic Environment
By constructing the perception matrix and cluster analysis, the sensor update is judged, which solves the problem of low perception accuracy of multi-source sensors in complex environments, and improves the efficiency and effect of robot autonomous adaptation and emotional interaction.
Patent Information
- Application Number
- CN202510561529.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In complex dynamic environments, the multi-source sensor fusion method has low perceptual accuracy, resulting in poor robotic self-adaptation ability to environmental changes, affecting emotional interaction efficiency and user experience.
By building a perception matrix, analyzing the sensor's associative array and cluster cluster, determining whether the sensor needs to be updated, and performing program updates to improve perception accuracy and autonomous adaptability.
It improves the robot's perception accuracy and autonomous adaptability in complex dynamic environments, and improves emotional interaction efficiency and user experience.
Smart Images

Figure CN120086618B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous sensing, and particularly relates to a multi-source sensor autonomous intelligent sensing method and system in a complex dynamic environment. Background Art
[0002] With the continuous development of robot technology, robots are increasingly widely used in complex dynamic environments. In these environments, robots need to perceive the surrounding environmental information in real time and accurately in order to make reasonable decisions and actions. However, complex dynamic environments have characteristics such as large amounts of information, rapid changes, and strong uncertainties, and a single sensor is difficult to meet the comprehensive and accurate requirements of robots for environmental perception.
[0003] Currently, multi-source sensor fusion technology is widely used to improve the environmental perception ability of robots. By fusing various types of sensor data, richer and more accurate environmental information can be obtained. However, existing multi-source sensor fusion methods still have some problems in complex dynamic environments. For example, during the operation of sensors, inaccurate perception results may occur due to low working accuracy, which in turn leads to poor autonomous adaptation ability of robots to environmental changes, affects the efficiency of emotional interaction, and reduces the user experience.
[0004] Therefore, the present invention proposes a multi-source sensor autonomous intelligent sensing method and system in a complex dynamic environment. Summary of the Invention
[0005] The present invention provides a multi-source sensor autonomous intelligent sensing method and system in a complex dynamic environment, which is used to construct a perception matrix based on historical perception results to assign an association array to perception elements, and then subsequently determine whether the sensor needs to be updated by performing clustering analysis on all association arrays under the sensor, so as to avoid inaccurate perception results caused by low perception accuracy of the sensor, improve the robot's autonomous adaptation ability to environmental changes, and improve the interaction efficiency and experience effect.
[0006] The present invention provides a multi-source sensor autonomous intelligent sensing method in a complex dynamic environment, including:
[0007] Step 1: Construct a perception matrix based on historical perception results;
[0008] Step 2: According to the overlap relationship between the feature description of any one perception element and the feature descriptions of each of the remaining perception elements in each column perception vector in the perception matrix, and in combination with the acquisition association relationship of the multi-source sensor, assign an association array to each perception element, where the overlap relationship is related to the scene intersection situation of the perception scene;
[0009] Step 3: Perform clustering analysis on all associated arrays under each sensor to obtain associated clusters, and determine whether the corresponding sensor needs program update according to the shortest distance of all associated clusters under the same sensor;
[0010] If no update is required, keep the set program of the corresponding sensor unchanged;
[0011] If an update is required, determine the update type of the corresponding sensor and perform program update;
[0012] Step 4: Sense the target to be sensed based on the updated multi-source sensors, and control the robot to perform emotional interaction with the target to be sensed.
[0013] Preferably, each row in the sensing matrix is the sensing result of the same sensor for different historical sensing targets, and each column in the sensing matrix is the sensing result of different sensors for the same historical sensing target, and the historical sensing target is a human.
[0014] Preferably, determining the overlapping relationship between the feature description of any one sensing element in each column of sensing vectors and the feature descriptions of each of the remaining sensing elements includes:
[0015] For the sensing result of each sensing element, retrieve the result analysis model consistent with the sensing type from the type-model comparison table according to the sensing type of the corresponding sensor;
[0016] Perform feature analysis on the corresponding sensing result based on the result analysis model to obtain a feature description;
[0017] Map the sensing results of each sensing element involved in each column of sensing vectors into the simulation space in chronological order, determine the time series of each sensing result, the sensing information at each time series, and the first quantity of the sensing types involved in the sensing information;
[0018] When the first quantity is 1, use the type threshold as the overlapping relationship of the sensing elements with the sensing types existing in the corresponding column of sensing vectors;
[0019] When the first quantity is 2, respectively obtain the first spatial increment of the sensing information of the two sensing elements with the sensing types existing in the corresponding column of sensing vectors in the simulation space, and obtain the first assignment value according to the first scenario intersection of the feature descriptions of the sensing information of the two sensing elements as the overlapping relationship;
[0020] When the first quantity is 3, respectively obtain the second spatial increment of the sensing information of each sensing element with the sensing types existing in the corresponding column of sensing vectors in the simulation space. At the same time, determine the corresponding second scenario intersection according to the feature descriptions of the sensing information of any two sensing elements;
[0021] For the second spatial increment, second scene intersection, and type threshold under the same sensing element, obtain a second assignment value as the overlapping relationship of the corresponding sensing element.
[0022] Preferably, in combination with the acquisition association relationship of the multi-source sensors, assign an association array to each sensing element, including:
[0023] Based on the acquisition association relationship, expand each overlapping relationship into two coefficients;
[0024] Take the two obtained coefficients as the association array of the corresponding sensing element.
[0025] Preferably, according to the shortest distance of all association clusters under the same sensor, determine whether the corresponding sensor needs program update, including:
[0026] Obtain the first distance between each association cluster and each remaining cluster respectively, and according to the principle of the smallest distance, obtain the shortest distance of all association clusters, where the shortest distance is the distance after connecting all association clusters according to the principle of the smallest distance;
[0027] Count the number of clusters of all association clusters under the same sensor;
[0028] Calculate the first association variance between all coefficients involved in one of the remaining sensors and all coefficients involved in another of the remaining sensors in all association arrays under the corresponding sensor, and determine the sum of the first association variance and the second association variance;
[0029] Round up the ratio of the sum to the variance threshold. If the rounded-up value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated;
[0030] If the rounded-up value is less than or equal to the number of clusters, at this time, if the product of the second ratio of the shortest distance to the distance threshold and the rounded-up value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated;
[0031] Otherwise, it is determined that the program of the corresponding sensor does not need to be updated.
[0032] Preferably, determine the update type of the corresponding sensor and perform program update, including:
[0033] According to all association clusters under the same sensor, sort them in sequence according to the shortest distance with the cluster with the smallest sum of coefficients among all association clusters as the starting point, and obtain a coefficient difference vector for each coefficient;
[0034] At the same time, determine the reference scene under each association cluster, and sort the reference scenes according to the sorting of all association clusters under the same sensor to obtain a scene difference vector;
[0035] Input the coefficient difference vector and the scenario difference vector into a vector analysis model to determine the update type of the sensor;
[0036] Match the update code from the type - program database to update the set program.
[0037] Preferably, controlling the robot to perform emotional interaction with the target to be sensed includes:
[0038] Performing emotional processing based on the perception result of the robot on the target to be sensed;
[0039] Determine the emotional interaction parameters according to the emotional processing result and send them to the facial unit of the robot for emotional control.
[0040] The present invention provides a multi - source sensor autonomous intelligent perception system in a complex dynamic environment, including:
[0041] A matrix construction module for constructing a perception matrix based on historical perception results;
[0042] An array construction module for assigning an association array to each perception element according to the overlap relationship between the feature description of any one perception element and the feature descriptions of each of the remaining perception elements in each column of perception vectors in the perception matrix, and in combination with the acquisition association relationship of the multi - source sensors, where the overlap relationship is related to the cross - probability of the perception scenario;
[0043] An update judgment module for performing clustering analysis on all association arrays under each sensor to obtain association clusters, and judging whether the corresponding sensor needs program update according to the shortest distance of all association clusters under the same sensor;
[0044] If no update is required, keep the set working parameters of the corresponding sensor unchanged;
[0045] If an update is required, determine the update type of the corresponding sensor and perform program update;
[0046] An interaction control module for perceiving the target to be sensed based on the updated multi - source sensors and controlling the robot to perform emotional interaction with the target to be sensed.
[0047] Compared with the prior art, the beneficial effects of the present application are as follows:
[0048] Obtain historical perception results to construct a perception matrix to assign association arrays to perception elements, and then subsequently determine whether the sensor needs to be updated by performing clustering analysis on all association arrays under the sensor, so as to avoid the inaccurate perception result caused by the low perception accuracy of the sensor, improve the robot's autonomous adaptation ability to environmental changes, and improve the interaction efficiency and experience effect.
[0049] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention may be realized and attained by the structure particularly pointed out in the written description and the drawings.
[0050] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0051] The drawings are used to provide a further understanding of the present invention, and constitute a part of the description. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0052] Figure 1 is a flowchart of a multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to an embodiment of the present invention;
[0053] Figure 2 is a structural diagram of a multi-source sensor autonomous intelligent perception system in a complex dynamic environment according to an embodiment of the present invention;
[0054] Figure 3 is a structural diagram of a walking robot according to an embodiment of the present invention. Detailed Embodiments
[0055] The preferred embodiments of the present invention will be described below with reference to the drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not used to limit the present invention.
[0056] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment, as Figure 1 shown, including:
[0057] Step 1: Construct a perception matrix based on historical perception results;
[0058] Step 2: According to the overlap relationship between the feature description of any one perception element in each column of perception vectors in the perception matrix and the feature descriptions of each remaining perception element, and in combination with the acquisition correlation relationship of the multi-source sensors, assign an association array to each perception element, wherein the overlap relationship is related to the scene intersection situation of the perception scene;
[0059] Step 3: Perform clustering analysis on all association arrays under each sensor to obtain association clusters, and judge whether the corresponding sensor needs program update according to the shortest distance of all association clusters under the same sensor;
[0060] If no update is required, keep the set program of the corresponding sensor unchanged;
[0061] If an update is required, determine the update type of the corresponding sensor and perform a program update;
[0062] Step 4: Sense the target to be sensed based on the updated multi-source sensors, and control the robot to perform emotional interaction with the target to be sensed.
[0063] Preferably, each row in the sensing matrix is the sensing result of the same sensor for different historical sensing targets, and each column in the sensing matrix is the sensing result of different sensors for the same historical sensing target, and the historical sensing targets are people.
[0064] In this embodiment, a complex dynamic environment refers to an environment with various changing factors, such as the movement of objects, changes in environmental conditions (temperature, light, etc.), multiple interference sources, etc. For example, in a crowded shopping mall with frequent pedestrian movement, the light changes with time and the light switch, and there are various sound interferences at the same time. This is a complex dynamic environment.
[0065] In this embodiment, the multi-source sensors are a set of sensors composed of various different types of sensors, which are used to obtain information from different angles, and these sensors are installed on the robot. They can be a high-definition camera for visual sensing, a microphone for sound sensing, and a radar for behavior sensing.
[0066] In this embodiment, it is assumed that a high-definition camera, a microphone, and a radar are set on the robot. At this time, the number of columns of the sensing matrix is 3, and the number of rows is the same as the number of historical sensing times involved in the historical situation. For the convenience of description, the sensing result under each device in the sensing matrix is regarded as a sensing element.
[0067] In this embodiment, the feature description refers to the information describing the characteristics of the sensing elements. For example, the feature description of the visual sensing element of a person by the camera may include the person's height, body shape, wearing color, etc.; the feature description of the sound sensing element of a person's speech by the microphone may include the frequency, volume, intonation, etc. of the sound, and there is such a situation in the subsequent determination of the overlapping relationship. For example, the camera senses that a person is wearing red clothes (visual sensing element), and the microphone senses that this person mentions red when speaking (sound sensing element). Then there is a certain overlapping relationship in the feature descriptions of these two sensing elements.
[0068] In this embodiment, the acquisition correlation relationship refers to the mutual correlation relationship existing between multi-source sensors when collecting information. For example, when the camera and the microphone collect information, they may simultaneously collect the visual and sound information of a person, and there is an acquisition correlation relationship between them. If they perform parallel acquisition at the same moment, the acquisition effect between the two devices is not as good as that of a single device for acquisition, etc., which may cause the generation or increase of noise.
[0069] The associated array contains two coefficient relationships, and these two coefficient relationships are realized based on the overlapping relationship. Since there are 3 devices, at this time, if it is the associated array of a determined microphone, then the associated array: {the relationship between the microphone and the high-definition camera, the relationship between the microphone and the radar}.
[0070] In this embodiment, the clustering analysis is implemented by the DBSCAN algorithm to achieve the clustering analysis of all associated arrays under the same sensor. And the clustering analysis belongs to the prior art means, so the associated clusters of the associated arrays involved under the same sensor can be directly obtained.
[0071] In this embodiment, the target to be sensed is the object that currently requires multi-source sensors for sensing. Here, a person is used as the target to be sensed. In the scene where the robot is located, as long as someone enters the scene, it is regarded as the target to be sensed.
[0072] In this embodiment, emotional interaction refers to the robot's response to the human emotional state through means such as voice and expression. For example, when a person shows happiness, the robot also expresses a happy emotion through voice and makes corresponding facial expression actions.
[0073] In this embodiment, the historical perception result depends on the robot's perception results (high-definition camera, microphone, radar) of the interactive user in the historical situation (within 3 months before the current time point). And the robot provides two types: a walking operation model and an alternative chassis operation model. The appearances of both models of robots incorporate Chinese elements, such as blue and white porcelain, auspicious clouds, Shijingshan, etc. Among them, for the chassis cheongsam model, only the appearance of the chassis needs to be designed, and the head refers to the proportion of the IP image of Huanhuan prototype without the need to redesign the appearance, as Figure 3 shown, is the walking robot, and this robot can perform the following operations:
[0074] Facial expression function: The robot can randomly make expressions such as blinking and eye rotation. Using the remote control, it can also be controlled to make rich expressions such as smiling, yawning, and frowning, vividly showing emotional changes.
[0075] Upper body limb movement function: Through the remote control, actions such as waving to say hello, shaking hands, randomly swinging the arms, and turning the head can be realized, meeting the needs of various social interaction scenarios.
[0076] Voice chat function: Based on the voice dialogue system of the large model, it supports setting a specified persona. The robot will chat and answer questions in the tone of the set persona, with specific voice wake-up and interruption functions, and can also play pre-recorded audio.
[0077] Visual perception function: With the camera built into the eyeball, the robot can observe the environmental scene and chat with people based on this, but the dialogue delay is about 3 - 5 seconds.
[0078] Biped walking function: The robot's legs can be conveniently controlled by a remote control to move forward, backward, and turn, enabling flexible walking. Both the biped walking of the walking operation model and the chassis walking of the chassis operation model have good maneuverability. The terrain adaptability of the chassis operation model further ensures stable movement in complex scenarios.
[0079] In this embodiment, the perception results are the activities of the user captured by the camera in different scenarios, the interactive behaviors such as the user's movement and gestures detected by the radar in the space, and the speech content recorded by the speech acquisition device. Suppose, in a home scenario, for 50 scenarios of user interaction with the intelligent device, the camera records visual information such as the user's actions, positions, and expressions; the radar obtains data such as the user's movement trajectory and the distance to the intelligent device; the speech acquisition device saves the command speech issued by the user. These data are organized into a 3×50 perception matrix. The first row is the perception data of the camera for 50 scenarios, the second row is the radar data, and the third row is the speech acquisition data; each column represents the perception results of the three sensors for the same scenario. By observing a certain column in the matrix, it is possible to intuitively understand that in a specific interaction scenario, for example, the camera captures the user's waving action, the radar detects that the user is 1 meter away from the intelligent device, and the speech acquisition device captures the user saying "turn on the TV". Specifically, the image recognition algorithm extracts the object positions and action features in the camera image, the radar data processing algorithm analyzes features such as the object's movement direction and distance, and the speech recognition and semantic analysis algorithm extracts the keywords and semantic information in the speech.
[0080] In this embodiment, when no update is required, taking the microphone as an example, its speech recognition, semantic analysis, and other programs maintain the current settings. If the camera needs to be updated, the update type is determined. At this time, the camera program is updated by downloading new model parameters and replacing the original model. The update program downloads the update package through the network and performs program replacement and configuration update according to the update interface specifications of the sensor device. Reasonable update decisions and operations enable the sensor to better adapt to the new environment and user behaviors, and its recognition accuracy has increased from 60% to 85%, improving the overall perception performance.
[0081] The beneficial effects of the above technical solution are: obtaining historical perception results to construct a perception matrix to assign associated arrays to perception elements, and then subsequently determining whether the sensor needs to be updated by performing clustering analysis on all associated arrays under the sensor, so as to avoid inaccurate perception results caused by low perception accuracy of the sensor, improve the robot's autonomous adaptation ability to environmental changes, and improve the interaction efficiency and experience effect.
[0082] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment, which determines the overlapping relationship between the feature description of any one sensing element in each column of sensing vectors and the feature descriptions of each of the remaining sensing elements, including:
[0083] For the sensing results of each sensing element, according to the sensing type of the corresponding sensor, retrieve the result analysis model consistent with the sensing type from the type-model look-up table;
[0084] Based on the result analysis model, perform feature analysis on the corresponding sensing results to obtain feature descriptions;
[0085] Map the sensing results of each sensing element involved in each column of sensing vectors to the simulation space in chronological order, determine the time series of each sensing result, the sensing information at each time sequence, and the first quantity of the sensing types involved in the sensing information;
[0086] When the first quantity is 1, use the type threshold as the overlapping relationship of the sensing elements with the sensing types existing in the corresponding column of sensing vectors;
[0087] When the first quantity is 2, respectively obtain the first spatial increment of the sensing information of the two sensing elements with the sensing types existing in the corresponding column of sensing vectors in the simulation space, and obtain the first assignment value according to the first scenario intersection of the feature descriptions of the sensing information of the two sensing elements as the overlapping relationship;
[0088] When the first quantity is 3, respectively obtain the second spatial increment of the sensing information of each sensing element with the sensing types existing in the corresponding column of sensing vectors in the simulation space. At the same time, determine the corresponding second scenario intersection according to the feature descriptions of the sensing information of any two sensing elements;
[0089] For the second spatial increment, the second scenario intersection, and the type threshold under the same sensing element, obtain the second assignment value as the overlapping relationship of the corresponding sensing element.
[0090] In this embodiment, during the system initialization phase, a "type - model correspondence table" is pre - constructed. This table stores the mapping relationships between different sensor perception types and the corresponding result analysis models. After obtaining the perception results of each perception element, first determine the type of sensor that produced the perception result. For example, if the perception result comes from a camera, then its perception type is determined as visual perception; if it comes from a radar, it is distance and motion perception; if it comes from a voice acquisition device, it is voice perception. Then, by querying the correspondence table, find the result analysis model that matches the perception type. Use the dictionary data structure in Python to store the correspondence table, and obtain the corresponding result analysis model through the sensor perception type as the key value. Specifically, in this table, the visual perception type corresponds to the image recognition analysis model, the voice perception type corresponds to the speech recognition and semantic analysis model, and the distance and motion perception type corresponds to the distance and motion analysis model. And the models included in this table are all pre - trained to facilitate direct use in this step to improve the efficiency of feature analysis.
[0091] For visual perception results (such as images or video clips captured by a camera), the image recognition analysis model will extract features in the image, such as the shape, color, edge information of objects, and the posture of the human body. For the distance and motion perception results of the radar, the analysis model will calculate features such as the movement speed, acceleration, and movement direction of the target object. For voice perception results, the speech recognition and semantic analysis model will extract frequency features, intonation features, semantic keywords, etc. of the voice. The specific implementation of these models may be based on various algorithms and technologies, such as convolutional neural networks in deep learning (for vision), hidden Markov models (for voice), etc. By calling the interface functions of the corresponding models, the perception results are passed to the models as input parameters. After the models run, they output feature description results. Specifically, the feature description results are: for the user's waving action captured by the camera, the feature description may be "the arm swings up and down at a frequency of 2 times per second, and the arm extension angle is between 120 degrees and 180 degrees"; for the user approaching the device detected by the radar, the feature description may be "the user moves towards the device at a speed of 0.5 meters per second, and the current distance from the device is 2 meters"; for the sentence "turn on the TV" collected by the voice, the feature description may be "including the keywords 'turn on' and 'TV', the speech intonation is declarative, and the frequency is concentrated in 200 - 400Hz".
[0092] In this embodiment, the simulation space is an abstract mathematical space used to map the perception results. For each column of perception vectors (representing the perception results of different sensors for the same historical perception target at different times), according to the timestamp information of each perception element, its perception results are mapped to different positions in the simulation space in chronological order. For example, time can be used as a dimension of the simulation space, and other dimensions can be defined according to the characteristics of the perception results. For visual perception results, the coordinate positions in the image can be used as additional dimensions; for radar perception results, distance and angle can be used as dimensions. During the mapping process, the time series corresponding to each perception result, as well as the perception information (i.e., the specific content of the perception element) and the perception type involved in the perception information at that moment are recorded. The number of different perception types in each time sequence is counted, which is the first quantity. Then, by using a data structure to store this information, a structure of nested dictionaries in Python lists is used. The list stores the information in each time sequence in chronological order, and the dictionary stores the perception information, perception type, and the first quantity, etc. It should be noted that the value of the first quantity is 1, 2, or 3. For example, in a certain time sequence, only the camera has perception information, then the first quantity is 1; if the camera and radar have perception information at the same time, the first quantity is 2; if the camera, radar, and voice collection device all have perception information, the first quantity is 3.
[0093] In this embodiment, the perception vector is a column in the perception matrix.
[0094] In this embodiment, the system has preset type thresholds for different perception types. When, in a certain time sequence, the first quantity obtained by statistics is 1, it means that only one perception type has perception information. At this time, directly use the type threshold corresponding to this perception type as the overlapping relationship of the perception elements with the perception type existing in this column of perception vectors. In terms of code implementation, a conditional judgment statement (such as an if statement) is used to detect whether the first quantity is 1. If so, obtain the threshold corresponding to the perception type from a variable or file that stores the type thresholds in advance, and assign it to the variable representing the overlapping relationship. Among them, the type thresholds for visual perception type, voice perception type, distance, and motion perception type are 0.8, 0.5, and 0.7 respectively.
[0095] In this embodiment, since there will be new perception results between the previous time sequence and a time sequence, new space points will be occupied at this time. By statistically analyzing the data of the new space points, the space increment is obtained, which is also the acquisition method of the first space increment and the second space increment.
[0096] In this embodiment, the first scenario intersection = a1×sim (feature description under one perception element, feature description under another perception element) + a2×the first spatial increment / the occupancy of the total spatial points under all perception results involved in the corresponding instance, where a1 and a2 are weights with values of 0.4 and 0.6 respectively.
[0097] Among them, the first spatial increment = the number of new spatial points occupied at the corresponding time sequence.
[0098] In this embodiment, the first assignment value = the type threshold of the corresponding perception type × (1 + the first scenario intersection). Among them, if it is the visual type and the voice type, at this time, it is calculated based on the visual type. Then, the type threshold is that of the visual type. At this time, the overlapping relationship is: the coefficient of the overlapping relationship between the visual type and the voice type is the first assignment value. If it is calculated based on the voice type, at this time, the calculated type threshold is that of the voice type. At this time, the overlapping relationship is: the coefficient of the overlapping relationship between the voice type and the visual type is the first assignment value.
[0099] In this embodiment, the second spatial increment is the number of new spatial points occupied by each perception element at the specified time sequence based on the simulated space.
[0100] The second scenario intersection is the similarity value of the feature descriptions of the perception information of any two perception elements, which is calculated based on the sim() similarity function.
[0101] If it is based on the voice type, at this time, the calculated type threshold is that of the voice type. At this time, the overlapping relationship is: the coefficient of the overlapping relationship between the voice type and the visual type is the corresponding second assignment value, and the coefficient of the overlapping relationship between the voice type and the distance and motion types is the corresponding second assignment value.
[0102] The second assignment value = the type threshold of the corresponding perception element × (1 + the second scenario intersection of the corresponding placement element that needs to perform overlapping analysis).
[0103] The second scenario intersection = a1×sim (feature description of the corresponding perception element, feature description of the corresponding perception element that needs to perform overlapping analysis) + a2×the second spatial increment of the corresponding perception element at the corresponding time sequence / the sum of the second spatial increments of the three types at the corresponding time sequence).
[0104] It should be noted that since there are 3 perception types, assumed to be type A1, type A2, and type A3 respectively. Here, it is necessary to calculate the second scenario intersection and the second assigned value between type A1 and type A2, and also calculate the second scenario intersection and the second assigned value between type A1 and type A3. At this time, the overlapping relationship of the perception elements of type A1 is: the second assigned value between type A1 and type A2, and the second assigned value between type A1 and type A3.
[0105] The beneficial effects of the above technical solution are as follows: The result analysis model is separately invoked based on different perception types to ensure the direct acquisition of the feature description of the perception result. The number of types at different time sequences is determined through spatial mapping, and then the overlapping relationship under different numbers of types is classified and discussed, which is convenient for targeted analysis, effectively establishing the association between different sensor types, providing a reasonable analysis basis for whether the program of the sensor is updated, and ensuring the effectiveness of the sensor update.
[0106] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment. Combining the acquisition association relationship of the multi-source sensors, an association array is assigned to each perception element, including:
[0107] Based on the acquisition association relationship, each overlapping relationship is extended into two coefficients;
[0108] The two obtained coefficients are used as the association array of the corresponding perception elements.
[0109] In this embodiment, for the case where the number of types is 3, there are already two coefficients. For the cases where the number of types is 1 and 2, the following method is used for extension:
[0110] Acquisition association relationship: the influence of the sensor of type A1 working synchronously with the sensor of type A2, the mutual influence of the sensor of type A1 working synchronously with the sensor of type A3, the mutual influence of the sensors of type A1, type A2, and type A3 working synchronously, the influence of the sensor of type A2 working synchronously with the sensor of type A3, and the influence of different types of sensors working synchronously is set in advance, and the value range of the influence is (0, 0.1), which is pre-stored in the sensor influence table. Just directly retrieve the result of the influence of different sensors working synchronously.
[0111] In this embodiment, the extension method for the case where the number of types is 1: Based on type A1:
[0112] The type threshold of type A1 × (1 - the influence of the sensor of type A1 working synchronously with the sensor of type A2); the type threshold of type A1 × (1 - the influence of the sensor of type A1 working synchronously with the sensor of type A3). These two results are the two extended coefficients.
[0113] In this embodiment, for the expansion method with the number of types being 1: Based on type A1, at this time, there is a first assignment value between type A1 and type A2:
[0114] The type threshold of type A1 × (1 - the influence of the sensors of type A1 and type A3 working synchronously), and take this calculation result and the first assignment value as the two coefficients for expansion.
[0115] For the case where the number of types is 3, there are already two coefficients itself: Based on type A1:
[0116] Take the second assignment value between type A1 and type A2 and the second assignment value between type A1 and type A3 directly as the two coefficients for expansion.
[0117] The associated array contains the two expanded coefficients, and there are two coefficients corresponding to type A1 at each time sequence.
[0118] The beneficial effect of the above technical solution is: Based on the association relationship, expand the overlap relationship with two coefficients to determine the unity of the array, ensure the accuracy of subsequent clustering analysis, and indirectly improve the perception accuracy.
[0119] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment. According to the shortest distance of all associated clusters under the same sensor, it is judged whether the corresponding sensor needs program update, including:
[0120] Obtain the first distance between each associated cluster and each remaining cluster respectively, and according to the principle of the smallest distance, obtain the shortest distance of all associated clusters, where the shortest distance is the distance after connecting all associated clusters according to the principle of the smallest distance;
[0121] Count the number of clusters of all associated clusters under the same sensor;
[0122] Calculate the first correlation variance between all coefficients involved in all associated arrays under the corresponding sensor and those of one of the other sensors, and the second correlation variance between all coefficients involved in all associated arrays under the corresponding sensor and those of the other remaining sensor, and determine the sum of the first correlation variance and the second correlation variance;
[0123] Round up the ratio of the sum to the variance threshold. If the rounded-up value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated;
[0124] If the rounded-up value is less than or equal to the number of clusters, at this time, if the product of the second ratio of the shortest distance to the distance threshold and the rounded-up value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated;
[0125] Otherwise, it is determined that the program of the corresponding sensor does not need to be updated.
[0126] In this embodiment, for each associated cluster, the minimum value is found from the distance values between it and other associated clusters, and these minimum values form the set of shortest distances for all associated clusters. To connect all associated clusters according to the principle of minimum distance, the minimum spanning tree algorithm (Prim algorithm or Kruskal algorithm) in graph theory can be used. Regarding the associated clusters as the nodes of a graph and the distance between clusters as the weight of an edge, the minimum weight path connecting all nodes (associated clusters) is obtained through the minimum spanning tree algorithm, and the total weight of this path is the distance after connecting all associated clusters according to the principle of minimum distance, that is, the final shortest distance.
[0127] The first distance refers to the distance value calculated between one associated cluster and another using a selected distance metric method. It is used to measure the proximity of two associated clusters in the data feature space.
[0128] The distance between cluster A and cluster B is 5, the distance between cluster A and cluster C is 3, and the distance between cluster B and cluster C is 4. For cluster A, its shortest distance to other clusters is 3 (the distance to cluster C); for cluster B, the shortest distance is 4 (the distance to cluster C); for cluster C, the shortest distance is 3 (the distance to cluster A). Connecting all associated clusters according to the principle of minimum distance and using the Prim algorithm, starting from cluster A, first connect to the nearest cluster C, and then connect cluster B. The final shortest distance obtained is 3 + 4 = 7.
[0129] After completing the clustering analysis to obtain the associated clusters, the clustering algorithm usually returns a structure containing all the associated clusters, such as a list, where each element represents an associated cluster. By simply obtaining the length of this list, the number of clusters of all associated clusters under the same sensor can be obtained.
[0130] There are two coefficients in each associated array. Therefore, for the same sensor: the first associated variance is obtained by calculating the variance of all coefficients of this sensor and the remaining one sensor, and the second associated variance is obtained by calculating the variance of all coefficients of this sensor and the remaining other sensor under the same sensor.
[0131] For example, all the coefficients of this sensor and the remaining one sensor are: [0.8, 0.6, 0.7, 0.9, 0.8, 0.75, 0.82, 0.78, 0.85, 0.72]. Calculating its average value is 0.78, and the first associated variance calculated according to the variance formula is approximately 0.0066.
[0132] The value of the variance threshold is 0.005. Assuming that the sum of the first correlation variance and the second correlation variance is 0.01, the ratio at this time is: 0.01 / 0.005 = 2. After rounding up, it is still 2. If the number of clusters of all associated clusters under this sensor is 1, since 2 is greater than 1, it is determined that the program of this sensor (such as a camera sensor) needs to be updated.
[0133] In this embodiment, assuming that the rounded value is 2, the shortest distance is 8, and the distance threshold is 4, then the second ratio of the shortest distance to the distance threshold is 8 / 4 = 2. Their product is 2 2 = 4. If the number of clusters of all associated clusters under this sensor is 3, since 4 is greater than 3, it is determined that the program of this sensor (such as a radar sensor) needs to be updated.
[0134] The beneficial effect of the above technical solution is that the shortest distance calculated can intuitively understand the tightness between associated clusters. A smaller shortest distance indicates that the difference between associated clusters is smaller and the data distribution is relatively concentrated; on the contrary, if the shortest distance is larger, it indicates that the difference between associated clusters is larger and the data distribution is more dispersed, providing a basis for the degree of data dispersion for subsequent judgment on whether the sensor program needs to be updated. A larger number of cluster quantities may mean a greater probability that the sensor data needs to be updated; a smaller number of cluster quantities indicates a smaller probability that the sensor needs to be updated. Based on the degree of change in the sensor data association relationship and the clustering situation of the data to judge whether the sensor program needs to be updated. If the rounded value is greater than the number of clusters, it means that the change in the sensor data association is relatively large, exceeding the range expected based on the number of clusters, and the program may need to be updated to adapt to this change to ensure the accuracy and stability of sensor data processing. When the condition for not needing to update is met, it means that the change in the sensor's data association relationship is within an acceptable range, and the sensor program can normally process the current data without the need for update operations, which helps to maintain the stability of the system and avoid the system overhead and potential risks brought by unnecessary program updates.
[0135] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment, determines the update type of the corresponding sensor, and performs program update, including:
[0136] According to all associated clusters under the same sensor, sort them in sequence according to the shortest distance with the cluster having the smallest coefficient sum among all associated clusters as the starting point to obtain the coefficient difference vector for each coefficient;
[0137] At the same time, determine the reference scenario for each associated cluster, and sort the reference scenarios according to the sorting of all associated clusters under the same sensor to obtain the scenario difference vector;
[0138] Input the coefficient difference vector and the scenario difference vector into a vector analysis model to determine the update type of the sensor;
[0139] Match the update code from the type - program database to update the set program.
[0140] In this embodiment, the associated cluster with the smallest sum of coefficients is used as the starting point. Starting from this starting point, the sorting sequence of the associated clusters is constructed according to the shortest distance. The idea of the greedy algorithm can be adopted. Starting from the starting point, each time the next cluster with the shortest distance to the current cluster is selected until all associated clusters are sorted. In the sorted sequence of associated clusters, calculate the difference of this coefficient value between adjacent clusters to form a coefficient difference vector, and there are two coefficient difference vectors. Each type of coefficient, for example, refers to that between type A1 and type A2, and between type A1 and type A3. And the difference calculation is the difference between two adjacent coefficients under the corresponding type of coefficient. For example, for the coefficient [u1, u2, u3, u4] between A1 and type A2, the obtained coefficient difference vector is: [u1 - u2, u2 - u3, u3 - u4].
[0141] In this embodiment, the reference scenario can be determined based on the semantic understanding of sensor data, the classification of application scenarios, etc. For example, for the associated cluster of a camera sensor, if an associated cluster mainly contains feature data of human walking, its reference scenario may be defined as "person movement scenario". The corresponding relationship between the associated cluster and the reference scenario is realized by using a dictionary in Python, where the key is the associated cluster identifier and the value is the corresponding reference scenario description.
[0142] In the sorted reference scenario sequence, analyze the difference between adjacent scenarios to construct a scenario difference vector, and the way to obtain the difference is similar to that of the coefficient difference vector. Specifically: if the previous scenario is "person static scenario" and the next scenario is "person movement scenario", then the scenario difference at this time is: from person static scenario to person movement scenario.
[0143] The vector analysis model is pre-trained and is a deep learning model such as a multi-layer perceptron (MLP). When building the model, a large amount of historical data is required for training. This historical data includes coefficient difference vectors and scenario difference vectors of different sensors under various working conditions, as well as the corresponding known update type labels. Deep learning libraries in Python (such as TensorFlow, PyTorch) are used to implement the construction, training, and prediction of the model. The obtained coefficient difference vectors and scenario difference vectors are used as input data and input into the trained vector analysis model. The model outputs the corresponding sensor update type through feature analysis and pattern recognition of the input vectors. It should be noted that the update types are updates for specific coefficient optimizations, scenario adaptability updates, algorithm upgrades, etc. Different update types correspond to different program modification and optimization directions. Suppose the trained vector analysis model is an SVM-based classifier, the input coefficient difference vector is 0.1, 0.2, 0.1, and the scenario difference vectors are from the static person scenario to the moving person scenario, from the moving person scenario to the squatting person scenario, and from the squatting person scenario to the running person scenario. After internal calculation and judgment by the model, the output update type is "scenario adaptability update". This indicates that the situation reflected by the current sensor data shows that the sensor program needs to be optimized for scenario adaptability.
[0144] The type-program database is pre-established, and this database stores the mapping relationship between different update types and the corresponding update codes. The update code can be a program script, a set of parameter adjustment values, a new algorithm module, etc., depending on the architecture and update method of the sensor program. A database management system (such as MySQL, SQLite, etc.) is used to store and manage this database. After the update type output by the vector analysis model, a query is made in the type-program database to find the matching update code. For example, an SQL query statement is used to search the database for records with the same value in the update type field as the model output and obtain the corresponding update code field value.
[0145] The obtained update code is applied to the set program of the sensor to achieve program update. This may involve operations such as code replacement, parameter modification, module loading, etc., and the specific operation method depends on the specific implementation of the sensor program.
[0146] Suppose there is a record in the type-program database with the update type "scenario adaptability update", and the corresponding update code is a Python script for adjusting the scenario classification threshold in the image recognition algorithm. After the vector analysis model determines that the update type of the sensor is "scenario adaptability update", the script code is queried from the database, and then this script code is applied to the image recognition program of the camera sensor to complete the update of the sensor program by modifying the corresponding threshold parameters.
[0147] The beneficial effects of the above technical solution are as follows: The scenario difference vector can intuitively show the change of the reference scenario with the sorting of the associated clusters, providing a basis for judging the sensor update type based on the scenario change. Through the vector analysis model, the update type of the sensor can be accurately determined according to the coefficient difference vector and the scenario difference vector. If the model training effect is good, the predicted update type can accurately reflect the actual problems of the sensor and the direction that needs to be optimized. This provides key information for subsequently matching the correct update code from the database, which helps to achieve precise sensor program updates. By matching the update code from the type-program database and applying it to the set program, the targeted update of the sensor program is realized. If the update code is correct and the application is error-free, the sensor should be able to better adapt to data changes and improve its perception and processing capabilities in subsequent work.
[0148] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment, which controls the robot to perform emotional interaction with the target to be perceived, including:
[0149] Performing emotional processing based on the perception result of the robot on the target to be perceived;
[0150] Determining emotional interaction parameters according to the emotional processing result and sending them to the facial unit of the robot for emotional control.
[0151] Controlling the robot to perform emotional interaction with the target to be perceived: For example, when a new user enters the room and becomes the target to be perceived, the updated camera recognizes that the user frowns and lowers the head (which may represent an unhappy expression), the radar detects that the user moves slowly (which may reflect the behavior when in a bad mood), and the voice collects that the user sighs and says "Today is really bad". Combining these perception information, the robot is controlled to comfort the user through voice "It sounds like you're not having a good day. Is there anything I can do to help?" and make a concerned expression (such as lowering the screen brightness and displaying a smiling icon).
[0152] Analyzing the user's emotional state through a pre-trained emotion recognition model (combining visual, motion, and speech features), and this model is based on a CNN neural network and is trained on the CNN neural network with samples of visual, motion, speech features and the corresponding emotional states, which belongs to the prior art. Then, according to the emotion recognition result, the voice synthesis and expression control programs of the robot are called to achieve emotional interaction. Experiments show that the user satisfaction with the robot has increased from 65% before the update to 80%, indicating that this perception and interaction method better meets the user's expectations.
[0153] The beneficial effects of the above technical solution are as follows: Through the collaborative work of the updated multi-source sensors, the robot can more accurately perceive the user's emotions and make appropriate interaction responses.
[0154] The present invention provides a multi-source sensor autonomous intelligent perception system in a complex dynamic environment, as Figure 2 shown, including:
[0155] A matrix construction module for constructing a perception matrix based on historical perception results;
[0156] An array construction module for assigning an association array to each perception element according to the overlap relationship between the feature descriptions of any one perception element and each remaining perception element in each column of perception vectors in the perception matrix, and in combination with the acquisition association relationship of the multi-source sensors, wherein the overlap relationship is related to the cross probability of the perception scenario;
[0157] An update judgment module for performing cluster analysis on all association arrays under each sensor to obtain association clusters, and judging whether the corresponding sensor needs program update according to the shortest distance of all association clusters under the same sensor;
[0158] If no update is required, keep the set working parameters of the corresponding sensor unchanged;
[0159] If an update is required, determine the update type of the corresponding sensor and perform program update;
[0160] An interaction control module for perceiving the target to be perceived based on the updated multi-source sensors and controlling the robot to perform emotional interaction with the target to be perceived.
[0161] The beneficial effects of the above technical solution are as follows: Obtain historical perception results to construct a perception matrix to assign association arrays to perception elements, and then subsequently determine whether the sensor needs to be updated by performing cluster analysis on all association arrays under the sensor, so as to avoid the situation that the perception result is inaccurate due to the low perception accuracy of the sensor, improve the robot's autonomous adaptation ability to environmental changes, and improve the interaction efficiency and experience effect.
[0162] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A multi-source sensor autonomous intelligent perception method in a complex dynamic environment, characterized in that, Including: Step 1: Construct a perception matrix based on historical perception results; Step 2: According to the overlapping relationship between the feature descriptions of any one perception element and the feature descriptions of each remaining perception element in each column perception vector of the perception matrix, and in combination with the acquisition association relationship of the multi-source sensors, assign an association array to each perception element, where the overlapping relationship is related to the scene intersection situation of the perception scene; Step 3: Perform clustering analysis on all association arrays under each sensor to obtain association clusters, and judge whether the corresponding sensor needs program update according to the shortest distance of all association clusters under the same sensor; If no update is required, keep the set program of the corresponding sensor unchanged; If an update is required, determine the update type of the corresponding sensor and perform program update; Step 4: Perceive the target to be perceived based on the updated multi-source sensors, and control the robot to perform emotional interaction with the target to be perceived; Among them, judging whether the corresponding sensor needs program update according to the shortest distance of all association clusters under the same sensor includes: Obtain the first distance between each association cluster and each remaining cluster respectively, and according to the principle of the smallest distance, obtain the shortest distance of all association clusters, where the shortest distance is the distance after connecting all association clusters according to the principle of the smallest distance; Count the number of clusters of all association clusters under the same sensor; Calculate the first association variance of all coefficients involved under one remaining sensor and the second association variance of all coefficients involved under another remaining sensor in all association arrays under the corresponding sensor, and determine the sum of the first association variance and the second association variance; Round up the ratio of the sum to the variance threshold. If the rounded-up value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated; If the rounded-up value is less than or equal to the number of clusters, at this time, if the product of the second ratio of the shortest distance to the distance threshold and the rounded-up value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated; Otherwise, it is determined that the program of the corresponding sensor does not need to be updated.
2. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1, wherein Each row in the perception matrix is the perception result of the same sensor for different historical perception targets, each column in the perception matrix is the perception result of different sensors for the same historical perception target, and the historical perception target is a person.
3. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1, wherein Determining the overlapping relationship between the feature descriptions of any one perception element and the feature descriptions of each remaining perception element in each column perception vector includes: For the perception results of each perception element, according to the perception type of the corresponding sensor, retrieve the result analysis model consistent with the perception type from the type-model comparison table; Based on the result analysis model, perform feature analysis on the corresponding perception results to obtain feature descriptions; Map the perception results of each perception element involved in each column perception vector to the simulation space in chronological order respectively, determine the time series of each perception result, the perception information at each time series, and the first quantity of the perception types involved in the perception information; When the first quantity is 1, use the type threshold as the overlapping relationship of the perception element with the existing perception type in the corresponding column perception vector; When the first quantity is 2, respectively obtain the first spatial increments of the sensing information of the two sensing elements with the sensing type in the corresponding column sensing vectors in the simulation space, and obtain a first assignment value according to the first scene intersection of the feature descriptions of the sensing information of the two sensing elements as the overlapping relationship; When the first quantity is 3, respectively obtain the second spatial increments of the sensing information of each sensing element with the sensing type in the corresponding column sensing vectors in the simulation space. At the same time, determine the corresponding second scene intersection according to the feature descriptions of the sensing information of any two sensing elements; For the second spatial increment, the second scene intersection and the type threshold under the same sensing element, obtain a second assignment value as the overlapping relationship of the corresponding sensing element.
4. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 3, characterized in that, Combined with the acquisition association relationship of the multi-source sensors, assign an association array to each sensing element, including: Based on the acquisition association relationship, expand each overlapping relationship into two coefficients; Use the two obtained coefficients as the association array of the corresponding sensing element.
5. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1, wherein Determine the update type of the corresponding sensor and perform program update, including: According to all the association clusters under the same sensor, sort them in sequence according to the shortest distance with the cluster with the smallest sum of coefficients in all the association clusters as the leading point to obtain the coefficient difference vector under each coefficient; At the same time, determine the reference scene under each association cluster, and sort the reference scenes according to the sorting of all the association clusters under the same sensor to obtain the scene difference vector; Input the coefficient difference vector and the scene difference vector into the vector analysis model to determine the update type of the sensor; Match the update code from the type-program database and update the set program.
6. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1, wherein Control the robot to perform emotional interaction with the target to be sensed, including: Perform emotional processing based on the sensing result of the robot on the target to be sensed; Determine the emotional interaction parameters according to the emotional processing result and send them to the facial unit of the robot for emotional control.
7. A multi-source sensor autonomous intelligent perception system in a complex dynamic environment, characterized in that, Including: A matrix construction module for constructing a sensing matrix based on historical sensing results; An array construction module for assigning an association array to each sensing element according to the overlapping relationship between the feature description of any one sensing element and the feature descriptions of each remaining sensing element in each column sensing vector in the sensing matrix, and combined with the acquisition association relationship of the multi-source sensors, wherein the overlapping relationship is related to the scene intersection of the sensing scene; An update judgment module for performing clustering analysis on all the association arrays under each sensor to obtain association clusters, and judging whether the corresponding sensor needs program update according to the shortest distance of all the association clusters under the same sensor; If no update is required, keep the set working parameters of the corresponding sensor unchanged; If an update is required, determine the update type of the corresponding sensor and perform program update; An interaction control module for sensing the target to be sensed based on the updated multi-source sensors and controlling the robot to perform emotional interaction with the target to be sensed; Among them, the update judgment module is used for: Obtain the first distance between each associated cluster and each remaining cluster respectively, and according to the principle of the smallest distance, obtain the shortest distance of all associated clusters, where the shortest distance is the distance after connecting all associated clusters according to the principle of the smallest distance; Count the number of clusters of all associated clusters under the same sensor; Calculate the first correlation variance between all associated arrays under the corresponding sensor and all coefficients involved under one of the other sensors, and the second correlation variance between all coefficients involved under another of the other sensors, and determine the sum of the first correlation variance and the second correlation variance; Round up the ratio of the sum to the variance threshold. If the rounded value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated; If the rounded value is less than or equal to the number of clusters, at this time, if the product of the second ratio of the shortest distance to the distance threshold and the rounded value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated; Otherwise, it is determined that the program of the corresponding sensor does not need to be updated.
Citation Information
Patent Citations
Road-end multi-source sensor fusion target sensing method and system for surface mine
CN114862901A
Environment sensing system based on multi-agent interaction
CN119202759A