Multi-source sensor autonomous intelligent sensing method and system in complex dynamic environment

By constructing a perception matrix and associative array, and combining cluster analysis to determine the sensor update situation, the problem of low sensor perception accuracy in complex dynamic environments is solved, and the efficiency of robots' autonomous adaptation and emotional interaction is improved.

CN120086618AActive Publication Date: 2025-06-03ZHIKAN SHENJIAN (BEIJING) TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510561529.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

In complex dynamic environments, a single sensor is difficult to meet the robot's comprehensive and accurate requirements for environmental perception, resulting in inaccurate perception results, affecting the robot's autonomous adaptability and emotional interaction efficiency.

Method used

By obtaining historical perception results, building a perception matrix, assigning associative arrays to perceptual elements, and determining whether the sensor needs to be updated through clustering analysis to improve perception accuracy and autonomous adaptability.

Benefits of technology

It effectively avoids the problem of inaccurate perception results caused by low sensor perception accuracy, and improves the robot's autonomous adaptability to environmental changes and emotional interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086618A_ABST
    Figure CN120086618A_ABST
Patent Text Reader

Abstract

The invention provides an autonomous intelligent sensing method and system for a multi-source sensor in a complex dynamic environment, and belongs to the technical field of autonomous sensing, and the method comprises the steps: constructing a sensing matrix based on a historical sensing result; according to the overlapping relationship between the feature description of any one sensing element in each column of sensing vectors in the sensing matrix and the feature description of each other sensing element, and in combination with the collection association relationship of the multi-source sensor, endowing each sensing element with an association array, carrying out clustering analysis on all the association arrays under each sensor to obtain an association cluster, and carrying out clustering analysis on all the association arrays under each sensor; according to the shortest distance of all the associated clusters under the same sensor, judging whether the corresponding sensor needs program updating or not; and sensing a to-be-sensed target based on the updated multi-source sensor, and controlling the robot to perform emotion interaction with the to-be-sensed target. The situation that the sensing result is inaccurate due to the fact that the sensing precision of the sensor is low is avoided, and the self-adaptive capacity of the robot to environment changes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous sensing, and particularly relates to a multi-source sensor autonomous intelligent sensing method and system in a complex dynamic environment. Background Art

[0002] With the continuous development of robot technology, robots are increasingly widely used in complex dynamic environments. In these environments, robots need to perceive surrounding environmental information in real time and accurately in order to make reasonable decisions and actions. However, complex dynamic environments have characteristics such as large amounts of information, rapid changes, and strong uncertainty, and a single sensor is difficult to meet the requirements of comprehensiveness and accuracy of environmental perception for robots.

[0003] Currently, multi-source sensor fusion technology is widely used to improve the environmental perception ability of robots. By fusing various types of sensor data, richer and more accurate environmental information can be obtained. However, existing multi-source sensor fusion methods still have some problems in complex dynamic environments. For example, during the operation of sensors, inaccurate perception results may occur due to low working accuracy, which in turn leads to poor autonomous adaptation ability of the robot to environmental changes, affects the efficiency of emotional interaction, and reduces the user experience effect.

[0004] Therefore, the present invention proposes a multi-source sensor autonomous intelligent sensing method and system in a complex dynamic environment. Summary of the Invention

[0005] The present invention provides a multi-source sensor autonomous intelligent sensing method and system in a complex dynamic environment, which is used to construct a perception matrix based on historical perception results to assign an association array to each perception element, and then subsequently determine whether the sensor needs to be updated by performing clustering analysis on all association arrays under the sensor, so as to avoid inaccurate perception results caused by low perception accuracy of the sensor, improve the robot's autonomous adaptation ability to environmental changes, and improve the interaction efficiency and experience effect.

[0006] The present invention provides a multi-source sensor autonomous intelligent sensing method in a complex dynamic environment, including: Step 1: Construct a perception matrix based on historical perception results; Step 2: According to the overlap relationship between the feature description of any one perception element and the feature descriptions of each of the remaining perception elements in each column perception vector in the perception matrix, and in combination with the acquisition association relationship of the multi-source sensors, assign an association array to each perception element, where the overlap relationship is related to the scene intersection situation of the perception scene; Step 3: Perform clustering analysis on all association arrays under each sensor to obtain association clusters, and determine whether the corresponding sensor needs program update according to the shortest distance of all association clusters under the same sensor; If no update is required, keep the setting program of the corresponding sensor unchanged; If an update is required, determine the update type of the corresponding sensor and perform a program update; Step 4: Sense the target to be sensed based on the updated multi-source sensors, and control the robot to perform emotional interaction with the target to be sensed.

[0007] Preferably, each row in the sensing matrix is the sensing result of the same sensor for different historical sensing targets, and each column in the sensing matrix is the sensing result of different sensors for the same historical sensing target, and the historical sensing target is a human.

[0008] Preferably, determining the overlapping relationship between the feature description of any one sensing element in each column of sensing vectors and the feature descriptions of each of the remaining sensing elements includes: For the sensing result of each sensing element, retrieve the result analysis model consistent with the sensing type from the type-model comparison table according to the sensing type of the corresponding sensor; Perform feature analysis on the corresponding sensing result based on the result analysis model to obtain a feature description; Map the sensing results of each sensing element involved in each column of sensing vectors to the simulation space in chronological order respectively, determine the time series of each sensing result, the sensing information at each time series, and the first quantity of the sensing types involved in the sensing information; When the first quantity is 1, use the type threshold as the overlapping relationship of the sensing elements with the sensing type in the corresponding column of sensing vectors; When the first quantity is 2, respectively obtain the first spatial increment of the sensing information of the two sensing elements with the sensing type in the corresponding column of sensing vectors in the simulation space, and obtain a first assignment value according to the first scene intersection of the feature descriptions of the sensing information of the two sensing elements as the overlapping relationship; When the first quantity is 3, respectively obtain the second spatial increment of the sensing information of each sensing element with the sensing type in the corresponding column of sensing vectors in the simulation space. At the same time, determine the corresponding second scene intersection according to the feature descriptions of the sensing information of any two sensing elements; For the second spatial increment, the second scene intersection, and the type threshold under the same sensing element, obtain a second assignment value as the overlapping relationship of the corresponding sensing element.

[0009] Preferably, combining the acquisition correlation relationship of the multi-source sensors, assign an association array to each sensing element, including: Based on the acquisition correlation relationship, expand each overlapping relationship into two coefficients; Use the two obtained coefficients as the association array of the corresponding sensing element.

[0010] Preferably, based on the shortest distances of all associated clusters under the same sensor, it is determined whether the corresponding sensor needs program update, including: Obtain the first distance between each associated cluster and each remaining cluster respectively, and according to the principle of the smallest distance, obtain the shortest distances of all associated clusters, where the shortest distance is the distance after connecting all associated clusters according to the principle of the smallest distance; Count the number of clusters of all associated clusters under the same sensor; Calculate the first correlation variance of all coefficients involved under one of the remaining sensors and the second correlation variance of all coefficients involved under the other remaining sensors in all associated arrays under the corresponding sensor, and determine the sum of the first correlation variance and the second correlation variance; Round up the ratio of the sum to the variance threshold. If the rounded-up value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated; If the rounded-up value is less than or equal to the number of clusters, at this time, if the product of the second ratio of the shortest distance to the distance threshold and the rounded-up value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated; Otherwise, it is determined that the program of the corresponding sensor does not need to be updated.

[0011] Preferably, determine the update type of the corresponding sensor and perform program update, including: According to all associated clusters under the same sensor, sort them in sequence according to the shortest distance with the cluster with the smallest sum of coefficients in all associated clusters as the starting point to obtain the coefficient difference vector for each coefficient; At the same time, determine the reference scenario for each associated cluster, and sort the reference scenarios according to the sorting of all associated clusters under the same sensor to obtain the scenario difference vector; Input the coefficient difference vector and the scenario difference vector into a vector analysis model to determine the update type of the sensor; Match the update code from the type - program database and update the set program.

[0012] Preferably, control the robot to perform emotional interaction with the target to be perceived, including: Perform emotional processing based on the perception result of the robot on the target to be perceived; Determine the emotional interaction parameters according to the emotional processing result and send them to the facial unit of the robot for emotional control.

[0013] The present invention provides a multi - source sensor autonomous intelligent perception system in a complex dynamic environment, including: A matrix construction module for constructing a perception matrix based on historical perception results; An array construction module, configured to assign an association array to each sensing element according to the overlapping relationship between the feature description of any one sensing element and the feature descriptions of each of the remaining sensing elements in each column of sensing vectors in the sensing matrix, and in combination with the acquisition association relationship of the multi-source sensors, where the overlapping relationship is related to the intersection probability of the sensing scenario; An update judgment module, configured to perform clustering analysis on all association arrays under each sensor to obtain association clusters, and determine whether the corresponding sensor needs program update according to the shortest distance of all association clusters under the same sensor; If no update is required, keep the set working parameters of the corresponding sensor unchanged; If an update is required, determine the update type of the corresponding sensor and perform program update; An interaction control module, configured to sense a target to be sensed based on the updated multi-source sensors and control the robot to perform emotional interaction with the target to be sensed.

[0014] Compared with the prior art, the beneficial effects of the present application are as follows: Obtain historical sensing results to construct a sensing matrix to assign an association array to the sensing elements, and then subsequently determine whether the sensor needs to be updated by performing clustering analysis on all association arrays under the sensor, so as to avoid inaccurate sensing results caused by low sensing accuracy of the sensor, improve the robot's autonomous adaptation ability to environmental changes, and improve the interaction efficiency and experience effect.

[0015] Other features and advantages of the present invention will be described in the subsequent description, and some of them will be obvious from the description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written description and the drawings.

[0016] The technical solutions of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0017] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of a multi-source sensor autonomous intelligent sensing method in a complex dynamic environment according to an embodiment of the present invention; Figure 2 is a structural diagram of a multi-source sensor autonomous intelligent sensing system in a complex dynamic environment according to an embodiment of the present invention; Figure 3 is a structural diagram of a walking robot according to an embodiment of the present invention. Detailed Embodiments

[0018] The preferred embodiments of the present invention will be described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0019] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment, as Figure 1 shown, including: Step 1: Construct a perception matrix based on historical perception results; Step 2: According to the overlap relationship between the feature description of any one perception element and the feature descriptions of each of the remaining perception elements in each column perception vector in the perception matrix, and in combination with the acquisition association relationship of the multi-source sensors, assign an association array to each perception element, where the overlap relationship is related to the scene intersection situation of the perception scene; Step 3: Perform clustering analysis on all the association arrays under each sensor to obtain association clusters, and judge whether the corresponding sensor needs program update according to the shortest distance of all the association clusters under the same sensor; If no update is required, keep the set program of the corresponding sensor unchanged; If an update is required, determine the update type of the corresponding sensor and perform program update; Step 4: Perceive the target to be perceived based on the updated multi-source sensors, and control the robot to perform emotional interaction with the target to be perceived.

[0020] Preferably, each row in the perception matrix is the perception result of the same sensor for different historical perception targets, each column in the perception matrix is the perception result of different sensors for the same historical perception target, and the historical perception target is a person.

[0021] In this embodiment, the complex dynamic environment refers to an environment with various changing factors, such as the movement of objects, the change of environmental conditions (temperature, light, etc.), multiple interference sources, etc. For example, in a crowded shopping mall with frequent pedestrian movement, the light will change with time and the turning on and off of lights, and there are also various sound interferences at the same time. This is a complex dynamic environment.

[0022] In this embodiment, the multi-source sensor is a sensor set composed of various different types of sensors, which is used to obtain information from different angles, and these sensors are installed on the robot. It can be a high-definition camera for visual perception, a microphone for sound perception, and a radar for behavior perception.

[0023] In this embodiment, it is assumed that a high-definition camera, a microphone, and a radar are set on the robot. At this time, the number of columns of the perception matrix is 3, and the number of rows is the same as the number of historical perception times involved in the historical situation. For the convenience of description, the perception result under each device in the perception matrix is regarded as a perception element.

[0024] In this embodiment, the feature description refers to the information describing the characteristics of the perception elements. For example, the feature description of the visual perception element of a person by a camera may include the person's height, body shape, wearing color, etc.; the feature description of the perception element of a person's speaking voice by a microphone may include the frequency, volume, intonation, etc. of the voice, and it exists in the process of determining the overlapping relationship subsequently. For example, the camera perceives that a person is wearing red clothes (visual perception element), and the microphone perceives that the person mentions red when speaking (voice perception element), then there is a certain overlapping relationship between the feature descriptions of these two perception elements.

[0025] In this embodiment, the acquisition correlation relationship refers to the mutual correlation relationship existing between multi-source sensors when collecting information. For example, when the camera and the microphone are collecting information, they may simultaneously collect the visual and voice information of a person, and there is an acquisition correlation relationship between them. If they are collected side by side at the same moment, the acquisition effect between the two devices is not as good as that of a single device for collection, etc., which may cause the generation or increase of noise.

[0026] The association array contains two coefficient relationships, and these two coefficient relationships are realized based on the overlapping relationship. Because there are 3 devices, at this time, if it is the association array of the microphone, at this time, the association array: {the relationship between the microphone and the high-definition camera, the relationship between the microphone and the radar}.

[0027] In this embodiment, the clustering analysis is implemented by the DBSCAN algorithm to achieve the clustering analysis of all association arrays under the same sensor. And the clustering analysis belongs to the prior art means, so the association clusters of the association arrays involved under the same sensor can be directly obtained.

[0028] In this embodiment, the target to be perceived is the object that currently requires multi-source sensors for perception. Here, a person is used as the target to be perceived. In the scene where the robot is located, as long as someone enters the scene, it is regarded as the target to be perceived.

[0029] In this embodiment, the emotional interaction refers to the robot's response to the emotional state of a person through means such as voice and expression. For example, when a person shows happiness, the robot also expresses the emotion of happiness through voice and makes corresponding facial expression actions.

[0030] In this embodiment, the historical perception result depends on the robot's perception of the interactive user in historical situations (within the past 3 months before the current moment) (using high-definition cameras, microphones, and radars). The robot is available in two types: a walking operation model and an alternative chassis operation model. The appearance of both models incorporates Chinese elements such as blue and white porcelain, auspicious clouds, and Shijingshan. For the chassis cheongsam model, only the appearance of the chassis needs to be designed, and the head refers to the proportion of the Huanhuan prototype IP image, without the need to redesign the appearance. As Figure 3 shown, it is a walking robot, and this robot can perform the following operations: Facial expression function: The robot can randomly make expressions such as blinking and eye rotation. Using the remote control, it can also be controlled to make rich expressions such as smiling, yawning, and frowning, vividly showing emotional changes.

[0031] Upper body limb movement function: Through the remote control, actions such as waving to say hello, shaking hands, randomly swinging the arms, and turning the head can be achieved, meeting the needs of various social interaction scenarios.

[0032] Voice chat function: Based on a large model voice dialogue system, it supports setting a specified persona. The robot will chat and answer questions in the tone of the set persona, with specific voice wake-up and interruption functions, and can also play pre-recorded audio.

[0033] Visual perception function: With the help of a camera built into the eyeball, the robot can observe the environmental scene and chat with people based on this, but the dialogue delay is about 3 - 5 seconds.

[0034] Bipedal walking function: Using the remote control, it is convenient to control the robot's legs to move forward, backward, and turn, achieving flexible walking. The bipedal walking of the walking operation model and the chassis walking of the chassis operation model both have good maneuverability. The terrain adaptability of the chassis operation model further ensures stable movement in complex scenarios.

[0035] In this embodiment, the perception results are the activities of the user captured by the camera in different scenarios. The radar detects the user's movement, gestures and other interaction behaviors in space, and the voice acquisition device records the user's speech content. Suppose that in a home scenario, for 50 scenarios of user interaction with intelligent devices, the camera records visual information such as the user's actions, positions, and expressions; the radar obtains data such as the user's movement trajectory and the distance to the intelligent device; the voice acquisition device saves the command voice issued by the user. These data are organized into a 3×50 perception matrix. The first row is the perception data of the camera for 50 scenarios, the second row is the radar data, and the third row is the voice acquisition data; each column represents the perception results of the three sensors for the same scenario. By observing a certain column in the matrix, it is possible to intuitively understand that in a specific interaction scenario, for example, the camera captures the user's waving action, the radar detects that the user is 1 meter away from the intelligent device, and the voice acquisition captures the user saying "turn on the TV". Specifically, the image recognition algorithm extracts the object position and action features in the camera image, the radar data processing algorithm analyzes features such as the object movement direction and distance, and the voice recognition and semantic analysis algorithm extracts the keywords and semantic information in the voice.

[0036] In this embodiment, when no update is required, taking the microphone as an example, its voice recognition, semantic analysis and other programs maintain the current settings. If the camera needs to be updated, the update type is determined. At this time, by downloading new model parameters and replacing the original ones, the camera program update is completed. The update program downloads the update package through the network and performs program replacement and configuration update according to the update interface specification of the sensor device. Reasonable update decisions and operations enable the sensor to better adapt to the new environment and user behaviors, and its recognition accuracy has increased from 60% to 85%, improving the overall perception performance.

[0037] The beneficial effects of the above technical solution are as follows: obtaining historical perception results to construct a perception matrix to assign associated arrays to perception elements, and then subsequently performing clustering analysis on all associated arrays under the sensor to determine whether the sensor needs to be updated, so as to avoid inaccurate perception results caused by low perception accuracy of the sensor, improve the robot's autonomous adaptation ability to environmental changes, and improve the interaction efficiency and experience effect.

[0038] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment, which determines the overlapping relationship between the feature descriptions of any one perception element and the feature descriptions of each of the remaining perception elements in each column of perception vectors, including: For the perception results of each perception element, according to the perception type of the corresponding sensor, retrieve the result analysis model consistent with the perception type from the type-model comparison table; Based on the result analysis model, perform feature analysis on the corresponding perception results to obtain feature descriptions; Map the perception results of each perception element involved in each column of perception vectors to the analog space in chronological order respectively, determine the time series of each perception result, the perception information at each time series, and the first quantity of the perception types involved in the perception information; When the first quantity is 1, use the type threshold as the overlapping relationship of the perception elements with the perception types existing in the corresponding column of perception vectors; When the first quantity is 2, respectively obtain the first spatial increment of the perception information of the two perception elements with the perception types existing in the corresponding column of perception vectors in the analog space, and obtain the first assignment value according to the first scenario intersection of the feature descriptions of the perception information of the two perception elements as the overlapping relationship; When the first quantity is 3, respectively obtain the second spatial increment of the perception information of each perception element with the perception types existing in the corresponding column of perception vectors in the analog space. At the same time, determine the corresponding second scenario intersection according to the feature descriptions of the perception information of any two perception elements; For the second spatial increment, the second scenario intersection, and the type threshold under the same perception element, obtain the second assignment value as the overlapping relationship of the corresponding perception element.

[0039] In this embodiment, in the system initialization stage, a "type-model look-up table" is constructed in advance. This table stores the mapping relationship between different sensor perception types and the corresponding result analysis models. After obtaining the perception results of each perception element, first determine the sensor type that generates the perception result. For example, if the perception result comes from a camera, determine its perception type as visual perception; if from a radar, it is distance and motion perception; if from a voice collection device, it is voice perception. Then, by querying the look-up table, find the result analysis model that matches the perception type. Use the dictionary data structure in Python to store the look-up table, and obtain the corresponding result analysis model through the sensor perception type as the key value. Specifically, in this table, the visual perception type corresponds to the image recognition analysis model, the voice perception type corresponds to the voice recognition and semantic analysis model, and the distance and motion perception type corresponds to the distance and motion analysis model. And the models included in this table are all pre-trained to facilitate direct use in this step to improve the efficiency of feature analysis.

[0040] For visual perception results (such as images or video clips captured by a camera), an image recognition and analysis model extracts features from the image, such as the shape, color, edge information of an object, and the pose of a human body. For the distance and motion perception results of a radar, the analysis model calculates features such as the motion speed, acceleration, and motion direction of the target object. For speech perception results, a speech recognition and semantic analysis model extracts features such as the frequency characteristics, intonation characteristics, and semantic keywords of the speech. The specific implementation of these models may be based on various algorithms and technologies, such as convolutional neural networks in deep learning (for vision), hidden Markov models (for speech), etc. By calling the interface functions of the corresponding models, the perception results are passed to the models as input parameters. After the models run, they output feature description results. The specific feature description results are as follows: for the user's waving action captured by the camera, the feature description may be "the arm swings up and down at a frequency of 2 times per second, and the extension angle of the arm is between 120 degrees and 180 degrees"; for the user approaching the device detected by the radar, the feature description may be "the user moves towards the device at a speed of 0.5 meters per second, and the current distance from the device is 2 meters"; for the sentence "turn on the TV" captured by the speech, the feature description may be "contains keywords 'turn on' and 'TV', the speech intonation is declarative, and the frequency is concentrated in 200 - 400 Hz".

[0041] In this embodiment, the simulation space is an abstract mathematical space used to map perception results. For each column of perception vectors (representing the perception results of different sensors for the same historical perception target at different times), according to the timestamp information of each perception element, its perception results are mapped to different positions in the simulation space in chronological order. For example, time can be used as a dimension of the simulation space, and other dimensions can be defined according to the characteristics of the perception results. For visual perception results, the coordinate positions in the image can be used as additional dimensions; for radar perception results, distance and angle can be used as dimensions. During the mapping process, the time series corresponding to each perception result, as well as the perception information at that moment (i.e., the specific content of the perception element) and the perception type involved in the perception information are recorded. The number of different perception types in each time sequence is counted, which is the first quantity. Then, by using a data structure to store this information, a structure of nested dictionaries in Python's list is used. The list stores the information in each time sequence in chronological order, and the dictionary stores the perception information, perception type, and the first quantity, etc. It should be noted that the value of the first quantity is 1, 2, or 3. For example, in a certain time sequence, only the camera has perception information, then the first quantity is 1; if the camera and the radar have perception information at the same time, the first quantity is 2; if the camera, the radar, and the speech acquisition device all have perception information, the first quantity is 3.

[0042] In this embodiment, the perception vector is a column in the perception matrix.

[0043] In this embodiment, the system pre-sets type thresholds for different perception types. When the first quantity statistically obtained at a certain time sequence is 1, it means that there is perception information for only one perception type. At this time, directly use the type threshold corresponding to this perception type as the overlapping relationship of the perception elements with the perception type existing in this column of perception vectors. In terms of code implementation, a conditional judgment statement (such as an if statement) is used to detect whether the first quantity is 1. If so, obtain the threshold corresponding to the perception type from the variable or file storing the type thresholds in advance, and assign it to the variable representing the overlapping relationship. Among them, the type thresholds for visual perception type, voice perception type, distance and motion perception type are 0.8, 0.5, and 0.7 respectively.

[0044] In this embodiment, since there will be new perception results between the previous time sequence and a time sequence, new space points will be occupied at this time. By statistically analyzing the data of the new space points, the space increment is obtained, that is, the acquisition methods of the first space increment and the second space increment.

[0045] In this embodiment, the first scenario intersection = a1×sim (feature descriptions under one perception element, feature descriptions under another perception element) + a2×the first space increment / the occupied amount of the total space points involved in the corresponding times for all perception results, where a1 and a2 are weights, and the values are 0.4 and 0.6 respectively.

[0046] Among them, the first space increment = the number of new space points occupied at the corresponding time sequence.

[0047] In this embodiment, the first assigned value = the type threshold corresponding to the perception type × (1 + the first scenario intersection). Among them, if it is the visual type and the voice type, at this time, it is calculated based on the visual type, then the type threshold is that of the visual type. At this time, the overlapping relationship is: the coefficient of the overlapping relationship between the visual type and the voice type is the first assigned value. If it is calculated based on the voice type, at this time, the calculated type threshold is that of the voice type. At this time, the overlapping relationship is: the coefficient of the overlapping relationship between the voice type and the visual type is the first assigned value.

[0048] In this embodiment, the second space increment is the number of new space points occupied by each perception element based on the simulated space at the specified time sequence.

[0049] The second scenario intersection is the similarity value of the feature descriptions of the perception information of any two perception elements, and is calculated based on the sim() similarity function.

[0050] If based on the voice type, at this time, the calculation type threshold is of the voice type, and the overlapping relationship at this time is: the coefficient of the overlapping relationship between the voice type and the visual type is the corresponding second assigned value, and the coefficient of the overlapping relationship between the voice type and the distance and motion types is the corresponding second assigned value.

[0051] The second assigned value = the type threshold of the corresponding perception element × (1 + the second scenario cross - property of the corresponding placement element that needs to perform overlapping analysis).

[0052] The second scenario cross - property = a1 × sim (the feature description of the corresponding perception element, the feature description of the corresponding perception element that needs to perform overlapping analysis) + a2 × (the second spatial increment of the corresponding perception element at the corresponding time sequence / the sum of the second spatial increments of the three types at the corresponding time sequence).

[0053] It should be noted that because there are 3 perception types, assumed to be type A1, type A2, and type A3 respectively. Here, it is necessary to calculate the second scenario cross - property and the second assigned value of type A1 and type A2, and also calculate the second scenario cross - property and the second assigned value of type A1 and type A3. At this time, the overlapping relationship of the perception elements of type A1 is: the second assigned value of type A1 and type A2, the second assigned value of type A1 and type A3.

[0054] The beneficial effects of the above - mentioned technical solution are: based on different perception types, the result analysis model is separately called to ensure the direct acquisition of the feature description of the perception result. The number of types at different time sequences is determined through spatial mapping, and then the overlapping relationship under different numbers of types is classified and discussed, which is convenient for targeted analysis, effectively establishing the association between different sensor types, providing a reasonable analysis basis for whether the program of the sensor is updated, and ensuring the effectiveness of sensor update.

[0055] The present invention provides a multi - source sensor autonomous intelligent perception method in a complex dynamic environment. Combining the acquisition association relationship of the multi - source sensors, an association array is assigned to each perception element, including: Based on the acquisition association relationship, each overlapping relationship is extended into two coefficients; The two obtained coefficients are used as the association array of the corresponding perception element.

[0056] In this embodiment, for the case where the number of types is 3, there are already two coefficients. For the cases where the number of types is 1 and 2, the following method is used for extension: Collection association relationship: The influence of sensors of type A1 working synchronously with sensors of type A2, the mutual influence of sensors of type A1 working synchronously with sensors of type A3, the mutual influence of sensors of type A1, type A2, and type A3 working synchronously, the influence of sensors of type A2 working synchronously with sensors of type A3, and the influence of different types of sensors working synchronously is set in advance, and the value range of the influence is (0, 0.1), which is pre-stored in the sensor influence table. Just directly retrieve the result of the influence of different sensors working synchronously.

[0057] In this embodiment, for the expansion method with a type quantity of 1: Based on type A1: The type threshold of type A1 × (1 - the influence of sensors of type A1 working synchronously with sensors of type A2); The type threshold of type A1 × (1 - the influence of sensors of type A1 working synchronously with sensors of type A3). These two results are the two expansion coefficients.

[0058] In this embodiment, for the expansion method with a type quantity of 1: Based on type A1, at this time, there is a first assignment value between type A1 and type A2: The type threshold of type A1 × (1 - the influence of sensors of type A1 working synchronously with sensors of type A3). Take this calculation result and the first assignment value as the two expansion coefficients.

[0059] For the case where the type quantity is 3, there are already two coefficients: Based on type A1: Take the second assignment value between type A1 and type A2 and the second assignment value between type A1 and type A3 directly as the two expansion coefficients.

[0060] The association array contains the two expanded coefficients, and there are two coefficients for type A1 at each time sequence.

[0061] The beneficial effect of the above technical solution is: Based on the association relationship, expand the overlap relationship with two coefficients to determine the unity of the array, ensure the accuracy of subsequent clustering analysis, and indirectly improve the perception accuracy.

[0062] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment. According to the shortest distance of all association clusters under the same sensor, it is judged whether the corresponding sensor needs program update, including: Obtain the first distance between each association cluster and each remaining cluster respectively, and according to the principle of the smallest distance, obtain the shortest distance of all association clusters, where the shortest distance is the distance after connecting all association clusters according to the principle of the smallest distance; Count the number of clusters of all association clusters under the same sensor; Calculate the first correlation variance of all coefficients involved under one of the remaining sensors and the second correlation variance of all coefficients involved under another of the remaining sensors in all associated arrays corresponding to the sensor, and determine the sum of the first correlation variance and the second correlation variance; Round up the ratio of the sum to the variance threshold. If the rounded value is greater than the number of clusters, it is determined that the program corresponding to the sensor needs to be updated; If the rounded value is less than or equal to the number of clusters, at this time, if the product of the second ratio of the shortest distance to the distance threshold and the rounded value is greater than the number of clusters, it is determined that the program corresponding to the sensor needs to be updated; Otherwise, it is determined that the program corresponding to the sensor does not need to be updated.

[0063] In this embodiment, for each associated cluster, find the minimum value from the distance values between it and other associated clusters. These minimum values form the set of shortest distances for all associated clusters. To connect all associated clusters according to the principle of the minimum distance, the minimum spanning tree algorithm (Prim algorithm or Kruskal algorithm) in graph theory can be used. Consider the associated clusters as the nodes of the graph and the distance between clusters as the weight of the edge. Through the minimum spanning tree algorithm, obtain the minimum weight path connecting all nodes (associated clusters). The total weight of this path is the distance after connecting all associated clusters according to the principle of the minimum distance, that is, the final shortest distance.

[0064] The first distance refers to the distance value between one associated cluster and another calculated using the selected distance metric method. It is used to measure the proximity of two associated clusters in the data feature space.

[0065] The distance between cluster A and cluster B is 5, the distance between cluster A and cluster C is 3, and the distance between cluster B and cluster C is 4. For cluster A, its shortest distance to other clusters is 3 (the distance to cluster C); for cluster B, the shortest distance is 4 (the distance to cluster C); for cluster C, the shortest distance is 3 (the distance to cluster A). Connect all associated clusters according to the principle of the minimum distance. Using the Prim algorithm, starting from cluster A, first connect to the closest cluster C, and then connect cluster B. The final shortest distance obtained is 3 + 4 = 7.

[0066] After completing the clustering analysis to obtain the associated clusters, the clustering algorithm usually returns a structure containing all the associated clusters, such as a list, where each element represents an associated cluster. By simply obtaining the length of this list, the number of clusters of all associated clusters under the same sensor can be obtained.

[0067] There are two coefficients in each associated array. Therefore, for the same sensor: the variances of all coefficients of this sensor and one of the other sensors are calculated to obtain the first associated variance, and the variances of all coefficients of this sensor and another one of the other sensors under the same sensor are calculated to obtain the second associated variance.

[0068] For example, all the coefficients of this sensor and one of the other sensors are: [0.8, 0.6, 0.7, 0.9, 0.8, 0.75, 0.82, 0.78, 0.85, 0.72]. The average value is calculated to be 0.78, and the first associated variance is calculated to be approximately 0.0066 according to the variance formula.

[0069] The value of the variance threshold is 0.005. Assuming that the sum of the first associated variance and the second associated variance is 0.01, the ratio at this time is: 0.01 / 0.005 = 2, and after rounding up, it is still 2. If the number of clusters of all associated clusters under this sensor is 1, since 2 is greater than 1, it is determined that the program of this sensor (such as a camera sensor) needs to be updated.

[0070] In this embodiment, assuming that the rounded value is 2, the shortest distance is 8, and the distance threshold is 4, then the second ratio of the shortest distance to the distance threshold is 8 / 4 = 2. Their product is 2 2 = 4. If the number of clusters of all associated clusters under this sensor is 3, since 4 is greater than 3, it is determined that the program of this sensor (such as a radar sensor) needs to be updated.

[0071] The beneficial effects of the above technical solution are as follows: The shortest distance calculated can intuitively understand the tightness between associated clusters. A smaller shortest distance indicates that the differences between associated clusters are smaller and the data distribution is relatively concentrated; conversely, if the shortest distance is larger, it indicates that the differences between associated clusters are larger and the data distribution is more dispersed, providing a basis for the degree of data dispersion for subsequent judgment on whether the sensor program needs to be updated. A larger number of cluster quantities may mean a greater probability that the sensor data needs to be updated; a smaller number of cluster quantities indicates a smaller probability that the sensor needs to be updated. Based on the degree of change in the sensor data association relationship and the clustering situation of the data, it is judged whether the sensor program needs to be updated. If the rounded value is greater than the number of clusters, it means that the change in the sensor data association is relatively large, exceeding the range expected based on the number of clusters, and the program may need to be updated to adapt to this change to ensure the accuracy and stability of sensor data processing. When the condition for not needing to update is met, it means that the change in the data association relationship of the sensor is within an acceptable range, and the sensor program can normally process the current data without the need for update operations, which helps to maintain the stability of the system and avoid the system overhead and potential risks brought by unnecessary program updates.

[0072] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment, determines the update type of the corresponding sensor, and updates the program, including: According to all associated clusters under the same sensor, sort them in sequence according to the shortest distance with the cluster having the smallest sum of coefficients in all associated clusters as the starting point, and obtain a coefficient difference vector for each coefficient; At the same time, determine the reference scenario for each associated cluster, and sort the reference scenarios according to the sorting of all associated clusters under the same sensor to obtain a scenario difference vector; Input the coefficient difference vector and the scenario difference vector into a vector analysis model to determine the update type of the sensor; Match the update code from the type-program database to update the set program.

[0073] In this embodiment, the associated cluster with the smallest sum of coefficients is used as the starting point. Starting from this starting point, construct a sorting sequence of associated clusters according to the shortest distance. The greedy algorithm idea can be adopted. Starting from the starting point, each time select the next cluster with the shortest distance to the current cluster until all associated clusters are sorted. In the sorted sequence of associated clusters, calculate the difference of this coefficient value between adjacent clusters to form a coefficient difference vector, and there are two coefficient difference vectors. Each coefficient, for example, refers to that of type A1 and type A2, type A1 and type A3, and the difference calculation is the difference between two adjacent coefficients corresponding to each coefficient. For example, for the coefficient [u1, u2, u3, u4] of A1 and type A2, the obtained coefficient difference vector is: [u1 - u2, u2 - u3, u3 - u4].

[0074] In this embodiment, the reference scenario can be determined based on the semantic understanding of sensor data, the classification of application scenarios, etc. For example, for the associated cluster of a camera sensor, if an associated cluster mainly contains feature data of human walking, its reference scenario may be defined as "person moving scenario", and the corresponding relationship between the associated cluster and the reference scenario is implemented by using a dictionary in Python, where the key is the associated cluster identifier and the value is the corresponding reference scenario description.

[0075] In the sorted reference scenario sequence, analyze the difference between adjacent scenarios to construct a scenario difference vector, and the way of obtaining the difference is similar to that of the coefficient difference vector. Specifically, if the previous scenario is "person static scenario" and the next scenario is "person moving scenario", the scenario difference at this time is: from person static scenario to person moving scenario.

[0076] The vector analysis model is pre-trained and is a deep learning model, such as a multi-layer perceptron (MLP). When building the model, a large amount of historical data is required for training. These historical data include coefficient difference vectors and scenario difference vectors of different sensors under various working conditions, as well as the corresponding known update type labels. Use deep learning libraries in Python (such as TensorFlow, PyTorch) to implement the construction, training, and prediction of the model. The obtained coefficient difference vectors and scenario difference vectors are used as input data and input into the trained vector analysis model. The model outputs the corresponding sensor update type through feature analysis and pattern recognition of the input vectors. It should be noted that the update type is an update for specific coefficient optimization, scenario adaptation, algorithm upgrade, etc. Different update types correspond to different program modification and optimization directions. Assume that the trained vector analysis model is an SVM-based classifier, the input coefficient difference vector is 0.1, 0.2, 0.1, and the scenario difference vectors are from the static person scenario to the moving person scenario, from the moving person scenario to the squatting person scenario, and from the squatting person scenario to the running person scenario. After internal calculation and judgment by the model, the output update type is "scenario adaptation update". This indicates that the situation reflected by the current sensor data shows that the sensor program needs to be optimized for scenario adaptation.

[0077] The type-program database is pre-established, and this database stores the mapping relationship between different update types and the corresponding update codes. The update code can be a program script, a set of parameter adjustment values, a new algorithm module, etc., depending on the architecture and update method of the sensor program. Use a database management system (such as MySQL, SQLite, etc.) to store and manage this database. After the update type output by the vector analysis model, query in the type-program database to find the matching update code. For example, use an SQL query statement to find the record in the database where the value of the update type field is the same as the model output and obtain the corresponding value of the update code field.

[0078] Apply the obtained update code to the set program of the sensor to achieve program update. This may involve operations such as code replacement, parameter modification, module loading, etc., and the specific operation method depends on the specific implementation of the sensor program.

[0079] Assume that there is a record in the type-program database, the update type is "scenario adaptation update", and the corresponding update code is a Python script for adjusting the scenario classification threshold in the image recognition algorithm. After the vector analysis model determines that the update type of the sensor is "scenario adaptation update", query the script code from the database, and then apply this script code to the image recognition program of the camera sensor to complete the update of the sensor program by modifying the corresponding threshold parameters.

[0080] The beneficial effects of the above technical solution are as follows: The scenario difference vector can intuitively show the changes in the reference scenario as the associated cluster is sorted, providing a basis for judging the sensor update type based on scenario changes. Through the vector analysis model, the update type of the sensor can be accurately determined according to the coefficient difference vector and the scenario difference vector. If the model training effect is good, the predicted update type can accurately reflect the actual problems of the sensor and the direction that needs to be optimized. This provides key information for subsequently matching the correct update code from the database, which helps to achieve accurate sensor program updates. By matching the update code from the type-program database and applying it to the set program, targeted updates of the sensor program are realized. If the update code is correct and applied correctly, the sensor should be able to better adapt to data changes and improve its perception and processing capabilities in subsequent operations.

[0081] The present invention provides a multi-source sensor autonomous intelligent perception method in a complex dynamic environment, which controls the robot to perform emotional interaction with the target to be perceived, including: Performing emotional processing based on the perception result of the robot on the target to be perceived; Determining emotional interaction parameters according to the emotional processing result and sending them to the facial unit of the robot for emotional control.

[0082] Controlling the robot to perform emotional interaction with the target to be perceived: For example, when a new user enters the room and becomes the target to be perceived, the updated camera recognizes that the user frowns and lowers the head (which may represent an unhappy expression), the radar detects that the user moves slowly (which may reflect the behavior when in a bad mood), and the voice collects that the user sighs and says "What a bad day today". Combining these perception information, the robot is controlled to comfort the user through voice "It sounds like you're having a rough day. Is there anything I can do to help?" and make a concerned expression (such as reducing the screen brightness and displaying a smiling icon).

[0083] Analyzing the user's emotional state through a pre-trained emotional recognition model (combining visual, motion, and voice features), and this model is based on a CNN neural network and is trained by using visual, motion, voice features and the corresponding emotional states as samples for the CNN neural network, which belongs to the prior art. Then, according to the emotional recognition result, the voice synthesis and expression control programs of the robot are called to achieve emotional interaction. Experiments show that the user satisfaction with the robot has increased from 65% before the update to 80%, indicating that this perception and interaction method better meets the user's expectations.

[0084] The beneficial effects of the above technical solution are as follows: Through the coordinated work of the updated multi-source sensors, the robot can more accurately perceive the user's emotions and make appropriate interaction responses.

[0085] The present invention provides a multi-source sensor autonomous intelligent perception system in a complex dynamic environment, as Figure 2 shown, including: A matrix construction module for constructing a perception matrix based on historical perception results; An array construction module for assigning an association array to each perception element according to the overlap relationship between the feature descriptions of any one perception element and each of the remaining perception elements in each column of perception vectors in the perception matrix, and in combination with the acquisition association relationship of the multi-source sensors, wherein the overlap relationship is related to the cross probability of the perception scenario; An update judgment module for performing clustering analysis on all association arrays under each sensor to obtain association clusters, and judging whether the corresponding sensor needs program update according to the shortest distance of all association clusters under the same sensor; If no update is required, keep the set working parameters of the corresponding sensor unchanged; If an update is required, determine the update type of the corresponding sensor and perform program update; An interaction control module for perceiving the target to be perceived based on the updated multi-source sensors and controlling the robot to perform emotional interaction with the target to be perceived.

[0086] The beneficial effects of the above technical solution are: obtaining historical perception results to construct a perception matrix to assign association arrays to perception elements, and then subsequently determining whether the sensor needs to be updated by performing clustering analysis on all association arrays under the sensor, so as to avoid the situation that the perception result is inaccurate due to the low perception accuracy of the sensor, improve the robot's autonomous adaptation ability to environmental changes, and improve the interaction efficiency and experience effect.

[0087] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A multi-source sensor autonomous intelligent perception method in a complex dynamic environment, characterized in that: include: Step 1: Construct a perception matrix based on historical perception results; Step 2: assigning an associative array to each sensing element according to the overlapping relationship between the feature description of any sensing element in each column of the sensing vector in the sensing matrix and the feature description of each remaining sensing element, and combining the acquisition association relationship of the multi-source sensor, wherein the overlapping relationship is related to the scene intersection of the sensing scene; Step 3: Perform cluster analysis on all association arrays under each sensor to obtain association clusters, and determine whether the corresponding sensor needs program update based on the shortest distance of all association clusters under the same sensor; If no update is required, keep the setting procedure of the corresponding sensor unchanged; If an update is required, determine the update type of the corresponding sensor and perform a program update; Step 4: Perceive the target to be perceived based on the updated multi-source sensor, and control the robot to interact emotionally with the target to be perceived.

2. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1 is characterized in that: Each row in the perception matrix is ​​the perception result of the same sensor on different historical perception targets, and each column in the perception matrix is ​​the perception result of different sensors on the same historical perception target, and the historical perception target is a person.

3. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1 is characterized in that: Determining the overlap between the feature description of any perceptual element in each column of the perceptual vector and the feature description of each remaining perceptual element includes: For each sensing result of the sensing element, according to the sensing type of the corresponding sensor, a result analysis model consistent with the sensing type is retrieved from the type-model comparison table; Performing feature analysis on the corresponding perception results based on the result analysis model to obtain feature description; Mapping the perception results of each perception element involved in each column of the perception vector to the simulation space in a time sequence order, determining the time sequence of each perception result and the perception information in each time sequence and the first number of perception types involved in the perception information; When the first number is 1, the type threshold is used as the overlapping relationship of the perception elements of the corresponding column perception vector having the perception type; When the first number is 2, first spatial increments of the perception information of two perception elements with the perception type in the corresponding column perception vector in the simulation space are respectively obtained, and a first assigned value is obtained as an overlapping relationship according to a first scene intersection of feature descriptions of the perception information of the two perception elements; When the first number is 3, respectively obtaining a second spatial increment of the perception information of each perception element having a perception type in the corresponding column perception vector in the simulation space, and determining the corresponding second scene intersection according to a feature description of the perception information of any two perception elements; For the second spatial increment, the second scene intersection and the type threshold under the same perception element, a second assigned value is obtained as the overlapping relationship of the corresponding perception element.

4. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 3 is characterized in that: In combination with the acquisition association relationship of the multi-source sensors, an association array is assigned to each sensing element, including: Expanding each overlapping relationship into two coefficients based on the acquisition association relationship; The two coefficients obtained are used as associative arrays of corresponding perceptual elements.

5. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1 is characterized in that: Based on the shortest distance of all associated clusters under the same sensor, determine whether the corresponding sensor needs a program update, including: Obtaining the first distance between each associated cluster and each remaining cluster, and obtaining the shortest distance of all associated clusters according to the minimum distance principle, wherein the shortest distance is the distance after connecting all associated clusters according to the minimum distance principle; Count the number of clusters of all associated clusters under the same sensor; Calculate the first correlation variance of all coefficients involved in one of the remaining sensors in all the correlation arrays under the corresponding sensor and the second correlation variance of all coefficients involved in another of the remaining sensors, and determine the sum of the first correlation variance and the second correlation variance; The ratio of the sum to the variance threshold is rounded up, and if the rounded value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated; If the rounded value is less than or equal to the number of clusters, at this time, if the product of the second ratio of the shortest distance to the distance threshold and the rounded value is greater than the number of clusters, it is determined that the program of the corresponding sensor needs to be updated; Otherwise, it is determined that the program corresponding to the sensor does not need to be updated.

6. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1 is characterized in that: Determine the update type for the corresponding sensor and perform program updates, including: All associated clusters under the same sensor are sorted in sequence according to the shortest distance and with the cluster with the smallest coefficient among all associated clusters as the first point, to obtain a coefficient difference vector under each coefficient; At the same time, the reference scene under each associated cluster is determined, and the reference scenes are sorted according to the sorting of all associated clusters under the same sensor to obtain the scene difference vector; Inputting the coefficient difference vector and the scene difference vector into a vector analysis model to determine an update type of the sensor; The setting program is updated by matching the update code from the type-program database.

7. The multi-source sensor autonomous intelligent perception method in a complex dynamic environment according to claim 1 is characterized in that: Controlling the robot to perform emotional interaction with a target to be sensed, including: Performing emotion processing based on the robot's perception result of the target to be perceived; Emotional interaction parameters are determined according to the emotion processing results and sent to the facial unit of the robot for emotion control.

8. A multi-source sensor autonomous intelligent perception system in a complex dynamic environment, characterized by: include: A matrix building module, used to build a perception matrix based on historical perception results; An array construction module, configured to assign an associative array to each sensing element according to an overlapping relationship between a feature description of any sensing element in each column of the sensing vector in the sensing matrix and a feature description of each remaining sensing element, and in combination with an acquisition association relationship of the multi-source sensor, wherein the overlapping relationship is related to a scene intersection condition of the sensing scene; The update judgment module is used to perform cluster analysis on all association arrays under each sensor to obtain association clusters, and judge whether the corresponding sensor needs program update based on the shortest distance of all association clusters under the same sensor; If no update is required, keep the set working parameters of the corresponding sensor unchanged; If an update is required, determine the update type of the corresponding sensor and perform a program update; The interactive control module is used to perceive the target to be perceived based on the updated multi-source sensor, and control the robot to interact emotionally with the target to be perceived.

Citation Information

Patent Citations

  • Road-end multi-source sensor fusion target sensing method and system for surface mine

    CN114862901A

  • Multi-target fusion tracking method and device based on sensing area

    CN116820087A

  • Environment sensing system based on multi-agent interaction

    CN119202759A

  • Methods and systems for secure and reliable identity-based computing

    US20150033305A1