A multi-dimensional pixel fusion method and system based on multi-source perception

Through multi-source perception technology, the integration of visible light, infrared and lidar data, combined with user tone feedback, accurate behavioral operation instructions are generated, which solves the problem of low interaction accuracy in the existing technology and improves the interaction efficiency and user experience of humanoid robots.

CN119942291BActive Publication Date: 2025-08-15ZHIKAN SHENJIAN (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510423254.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-15
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

When interacting with users, existing humanoid robots rely on a single interaction means such as voice or text feedback, and cannot provide multi-directional feedback in combination with the user's facial expressions and tone, resulting in low interaction accuracy and unable to satisfy the user's interactive experience.

Method used

Using a multi-source perception method, data is collected through visible light cameras, infrared cameras and lidars, coordinates are unified and timing aligned, features are extracted, and multi-dimensional pixel fusion is combined with user tone attributes to generate behavioral operation instructions.

Benefits of technology

It improves interaction efficiency, ensures the accuracy of behavioral operation instructions, and improves the user's interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942291B_ABST
    Figure CN119942291B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-dimensional pixel fusion method and system based on multi-source perception, belonging to the field of pixel fusion technology. The method includes: controlling the various sensors provided on the robot to collect data, unifying the coordinates and aligning the collected data of each sensor in time sequence to obtain a collection array for each collection point; arranging and processing the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and extracting features from the collected image according to a feature extraction method matching the corresponding sensor type; aligning the extracted features of each sensor type, and performing multi-dimensional pixel fusion based on the collection array of the collection point and the tone attributes of the target user to obtain a behavioral operation instruction, which is output to the robot's control unit for execution. This ensures the reliability of the multi-dimensional pixel fusion, thereby making the behavioral operation instruction more accurate, improving interaction efficiency, and satisfying the user's interactive experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of pixel fusion technology, and in particular to a multi-dimensional pixel fusion method and system based on multi-source perception. Background Art

[0002] With the increasing promotion of humanoid robots and the increasing maturity of technology, more and more users are being attracted to actively chat and interact with humanoid robots. Common humanoid robots provide interactive feedback by collecting user interaction information, and generally rely on a single interactive method during the collection process. For example, simply outputting text or voice feedback to the user does not provide multi-faceted feedback based on the user's facial expressions or tone, resulting in low interaction accuracy. In other words, the user's interactive experience cannot be satisfied due to low interaction efficiency.

[0003] Therefore, the present invention proposes a multi-dimensional pixel fusion method and system based on multi-source perception. Summary of the Invention

[0004] The present invention provides a multi-dimensional pixel fusion method and system based on multi-source perception, which is used to determine multi-dimensional pixel information from multi-source perception, and combines the influence of tone feedback during the user's voice process to ensure the reliability of multi-dimensional pixel fusion, thereby making behavioral operation instructions more accurate, improving interaction efficiency and satisfying the user's interactive experience.

[0005] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, comprising:

[0006] Step 1: Control the sensors installed on the robot to collect data, wherein the sensors include: a visible light camera, an infrared camera, and a lidar;

[0007] Step 2: Unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point;

[0008] Step 3: Arrange the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to the feature extraction method matching the corresponding type of sensor;

[0009] Step 4: Align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion based on the acquisition array of the acquisition points and the tone attributes of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.

[0010] Preferably, before controlling the sensors provided on the robot to collect data, the following steps are included:

[0011] Capturing a trigger instruction to the robot from external information, and determining the instruction type of the trigger instruction;

[0012] A start-up unit consistent with the instruction type is matched from a type-startup database, and the robot is started in a power-on state based on the start-up unit.

[0013] Preferably, controlling the sensors provided on the robot itself to collect data includes:

[0014] When the robot is in a powered-on state, the robot collects an interactive input set of a target user currently received by the robot, and interactively analyzes the interactive input set to set a collection period for different types of sensors;

[0015] The sensors of the corresponding type are controlled to collect data according to the collection period to obtain an initial set of each sensor.

[0016] Preferably, setting acquisition cycles for different types of sensors includes:

[0017] Determining an input type of the interactive input set, and when the input type is voice input, respectively obtaining a response curve of each microphone collecting the user's voice;

[0018] Based on the current distance between the robot's microphone location and the target user's sound source, the response curves are sequentially sorted and aligned to obtain the response difference at each time point;

[0019] Based on the response difference and time delay, a compensation adjustment is performed on the second curve that is the second closest, and waveform superposition processing is performed on the first curve that is the closest;

[0020] Performing speech noise reduction processing on the superimposed speech to obtain the converted text;

[0021] Performing keyword extraction on the converted text;

[0022] Determining the duration of the process of the robot receiving the user's voice and performing semantic analysis on the user's voice, and at the same time, determining the voice duration of the user's voice;

[0023] Obtaining a tone complexity coefficient according to the speech duration, the process duration, the tone information of the speech after the noise reduction process, and the number of words in the converted text;

[0024] Retrieving a tone feature model that matches the tone complexity coefficient from a coefficient-feature comparison table to analyze the speech after noise reduction processing to determine a tone annotation for each keyword;

[0025] Merging the tone annotation with the extracted keywords to obtain a first annotation;

[0026] When the input type is text input, extract keywords from the input text to obtain a second annotation;

[0027] The target user's interaction intention is determined according to the annotation, and the acquisition cycles of different types of sensors are matched from the intention comparison table.

[0028] Preferably, the extracted features of each type of sensor are aligned, and multi-dimensional pixel fusion is performed in combination with the acquisition array of the acquisition points and the tone attributes of the target user, including:

[0029] Constructing a gradient vector for each pixel in the acquisition array at each acquisition point, wherein the gradient vector includes a brightness gradient component of the corresponding pixel under different types of sensors;

[0030] Determine the information saturation coefficient of each pixel according to the gradient vector;

[0031] When the information saturation coefficient is greater than a preset coefficient, the corresponding collection point is retained; otherwise, the corresponding collection point is eliminated;

[0032] Perform spatial domain conversion on the extracted features of each type of sensor to obtain a spatial array for each spatial point. When there are remaining points that are consistent with the corresponding spatial points, perform pixel fusion on the spatial array and the acquisition array at the consistent points according to the tone attributes of the target user to obtain and retain the fused array.

[0033] When there is no corresponding spatial point consistent with the remaining points, the acquisition array corresponding to the remaining points is retained;

[0034] Analyze the retained array of each point to obtain behavioral operation instructions.

[0035] Preferably, pixel fusion is performed on the spatial array and the acquisition array at the consistent point according to the tone attribute of the target user to obtain a fused array, including:

[0036] constructing an initial matrix based on the spatial array and the acquisition array;

[0037] When the tone attribute indicates that the input type of the target user's interaction with the robot is unrelated to speech, an average value of each column element in the initial matrix is calculated to obtain a first array;

[0038] When the tone attribute is an input type related to speech for the target user's interaction with the robot, determining the facial region of the target user on which the consistent point acts based on a comparison relationship between the consistent point and the target user's face, and setting simulation parameters for the consistent point according to an expression control database for the facial region;

[0039] Extracting the current tone of the target user at each time point during the interaction with the robot, and constructing facial change information for each consistent point based on the target user's tone portrait and simulation parameters, thereby constructing a facial array based on the tone;

[0040] Appending the face array to the first row of the initial matrix and combining the weights assigned to each row vector in the initial matrix to obtain a second array;

[0041] Among them, the first array and the second array are corresponding fusion arrays.

[0042] Preferably, the retained array of each point is analyzed to obtain behavioral operation instructions, including:

[0043] Divide the retained array of each point based on the user's face block criteria to obtain the current representation of each block;

[0044] Based on all current representations, the block weight of each divided block, and the user intention determined by the user voice of the target user, a behavioral operation instruction is obtained from the intention-combination-instruction comparison table.

[0045] The present invention provides a multi-dimensional pixel fusion system based on multi-source perception, comprising:

[0046] A data acquisition module is used to control the various sensors provided on the robot to collect data, wherein the sensors include: a visible light camera, an infrared camera, and a lidar;

[0047] The alignment processing module is used to unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point;

[0048] A feature extraction module is used to arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and to extract features from the collected image according to a feature extraction method that matches the corresponding type of sensor;

[0049] The multi-dimensional pixel fusion module is used to align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion based on the acquisition array of the acquisition points and the tone attributes of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.

[0050] Compared with the prior art, the present invention has the following advantages:

[0051] Starting from multi-source perception to determine multi-dimensional pixel information, and combining the influence of tone feedback during the user's voice process to ensure the reliability of multi-dimensional pixel fusion, thereby making behavioral operation instructions more accurate, improving interaction efficiency and satisfying the user's interactive experience.

[0052] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0053] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0055] Figure 1 This is a flow chart of a multi-dimensional pixel fusion method based on multi-source perception in an embodiment of the present invention;

[0056] Figure 2 This is a structural diagram of a multi-dimensional pixel fusion system based on multi-source perception in an embodiment of the present invention;

[0057] Figure 3 This is an external view of the robot in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0059] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, such as Figure 1 Shown, including:

[0060] Step 1: Control the various sensors set on the robot to collect data, wherein the sensors include: visible light camera, infrared camera and laser radar;

[0061] Step 2: Unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point;

[0062] Step 3: Arrange the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to the feature extraction method matching the corresponding type of sensor;

[0063] Step 4: Align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion based on the acquisition array of the acquisition points and the tone attributes of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.

[0064] In this embodiment, the robot is an existing robot that can mainly interact with humans. The robot is designed to be a female image, 168 cm tall, with black hair and black eyes. It is available in two types: a walking operation model and an optional chassis operation model. The appearance incorporates Chinese elements to bring a unique experience to users.

[0065] The appearance of the robots are all integrated with Chinese elements, such as blue and white porcelain, auspicious clouds, Shijingshan, etc. Among them, the cheongsam model only needs to design the chassis appearance, and the head refers to the proportion of the prototype IP image of Huanhuan, without redesigning the appearance. Figure 3 shown.

[0066] In this embodiment, the functions of the robot are as follows, including:

[0067] Facial Expression: The robot can randomly blink and roll its eyes. Using a remote control, it can also be controlled to smile, yawn, frown, and other expressions, vividly displaying its emotional changes.

[0068] Upper body movement function: Through the remote control, you can perform actions such as waving, shaking hands, swinging your arms randomly, turning your head, etc., to meet the needs of various social interaction scenarios.

[0069] Voice chat: A large-scale voice dialogue system supports the ability to set a specific persona. The robot will chat and answer questions in the voice of the set persona, with specific voice wake-up and interruption functions, and can also play pre-recorded audio.

[0070] Visual perception function: With the help of the camera built into the eyeball, the robot can observe the environment and chat with people based on it, but the conversation delay is about 3-5 seconds.

[0071] Bipedal walking function: The remote control can be used to easily control the robot's legs to move forward, backward, and turn, achieving flexible walking. Both the bipedal walking of the walking operation model and the chassis walking of the chassis operation model have good controllability. The terrain adaptability of the chassis operation model further ensures stable movement in complex scenes.

[0072] In this embodiment, the behavioral operation instructions are facial emotion interaction instructions, voice chat interaction instructions, body movement interaction instructions, walking control instructions, etc. between the robot and the user.

[0073] In this embodiment, multi-dimensional pixel fusion for multi-source sensing refers to the technology that fuses pixel data containing information of different dimensions acquired by multiple sensors (such as cameras and radar) to generate a richer, more accurate, and more comprehensive image or data representation. For example, pixel information from a visible light image can be fused with pixel information from an infrared image, so that the fused image contains both details from visible light and characteristic information such as temperature from infrared light.

[0074] In this embodiment, coordinate unification and temporal alignment are performed to ensure that pixels of different source images accurately correspond in space. For example, corresponding points at the same physical position in different images are found through methods such as feature point matching, and the images are transformed by translation, rotation, scaling, etc. to align them in space.

[0075] In this embodiment, representative features are extracted from pixels of each source image, such as extracting features such as edges and textures from visible light images, and extracting geometric features such as distances to objects from depth images.

[0076] In this embodiment, pixel fusion algorithms include weighted average fusion, wavelet transform-based fusion, sparse representation-based fusion, etc. For example, weighted average fusion assigns a weight to each pixel based on factors such as the reliability of different source images, and then performs weighted summation to obtain the fused pixel value.

[0077] In this embodiment, the collection point is the point where the face of the target user is collected, and the collected data is pixel data related to the user's face.

[0078] In this embodiment, the acquisition array = {pixel information acquired by the sensor of each type at the corresponding acquisition point}.

[0079] In this embodiment, because there is at least one sensor of each type and the sensors are installed at different positions, a complete captured image can be obtained during the process of arranging the sensors in spatial order, that is, each type of sensor corresponds to a captured image at the corresponding moment, which can be a facial image of the user, and the target user refers to the user who interacts with the robot.

[0080] In this embodiment, PyTorch can be used to extract features of visible light images, local invariant features can be used to extract features of infrared images, and the DBSCAN algorithm can be used to extract features of lidar images.

[0081] In this embodiment, feature alignment processing refers to position alignment processing of the features extracted at each time point. This is because only valuable points are retained during the feature extraction process. Therefore, feature information of related spatial points can be obtained after alignment processing.

[0082] In this embodiment, the tone attribute of the target user includes: whether there is voice interaction between the user and the robot, or whether there is no voice interaction between the user and the robot.

[0083] In this embodiment, multi-dimensional pixel fusion refers to fusing pixels at corresponding locations to facilitate subsequent analysis and obtain behavioral operation instructions.

[0084] The beneficial effects of the above technical solution are: starting from multi-source perception to determine multi-dimensional pixel information, and combining the influence of tone feedback during the user's voice process to ensure the reliability of multi-dimensional pixel fusion, thereby making behavioral operation instructions more accurate, thereby improving interaction efficiency and satisfying the user's interactive experience.

[0085] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which includes:

[0086] Capturing a trigger instruction to the robot from external information, and determining the instruction type of the trigger instruction;

[0087] A start-up unit consistent with the instruction type is matched from a type-startup database, and the robot is started in a power-on state based on the start-up unit.

[0088] In this embodiment, the trigger instruction refers to whether the target user has a behavioral instruction to interact with the user, and the instruction type is a combination of any one or more of the voice interaction type, behavioral interaction type, and emotional interaction type.

[0089] In this embodiment, the type-startup database contains different instruction types and start-up units in the robot that match the type, that is, the corresponding function of the robot can be triggered to start through the start-up unit, and then the robot is in the power-on state after startup.

[0090] The beneficial effect of the above technical solution is: by determining the instruction type, the starting unit is determined, and then the robot is started, and it is determined that subsequent data can be collected normally.

[0091] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which controls the sensors set on the robot itself to collect data, including:

[0092] When the robot is in a powered-on state, the robot collects an interactive input set of a target user currently received by the robot, and interactively analyzes the interactive input set to set a collection period for different types of sensors;

[0093] The sensors of the corresponding type are controlled to collect data according to the collection period to obtain an initial set of each sensor.

[0094] In this embodiment, the interactive input set may be user's voice input information or text input information.

[0095] In this embodiment, the purpose of interactive analysis is to determine the user's intention, and then set the collection period for the sensor to ensure the comprehensiveness and accuracy of the collection.

[0096] In this embodiment, the initial set includes collection results of the corresponding sensor in different collection cycles.

[0097] The beneficial effect of the above technical solution is: by analyzing the interactive input set to set the acquisition cycle, to determine the comprehensiveness and accuracy of the acquisition, and to provide a basis for subsequent pixel fusion.

[0098] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which sets acquisition cycles for different types of sensors, including:

[0099] Determining an input type of the interactive input set, and when the input type is voice input, respectively obtaining a response curve of each microphone collecting the user's voice;

[0100] Based on the current distance between the robot's microphone location and the target user's sound source, the response curves are sequentially sorted and aligned to obtain the response difference at each time point;

[0101] Based on the response difference and time delay, a compensation adjustment is performed on the second curve that is the second closest, and waveform superposition processing is performed on the first curve that is the closest;

[0102] Performing speech noise reduction processing on the superimposed speech to obtain the converted text;

[0103] Performing keyword extraction on the converted text;

[0104] Determining the duration of the process of the robot receiving the user's voice and performing semantic analysis on the user's voice, and at the same time, determining the voice duration of the user's voice;

[0105] Obtaining a tone complexity coefficient according to the speech duration, the process duration, the tone information of the speech after the noise reduction processing, and the number of words in the converted text;

[0106] Retrieving a tone feature model that matches the tone complexity coefficient from a coefficient-feature comparison table to analyze the speech after noise reduction processing to determine a tone annotation for each keyword;

[0107] Merging the tone annotation with the extracted keywords to obtain a first annotation;

[0108] When the input type is text input, extract keywords from the input text to obtain a second annotation;

[0109] The target user's interaction intention is determined according to the annotation, and the acquisition cycles of different types of sensors are matched from the intention comparison table.

[0110] In this embodiment, keyword extraction from text is achieved based on TF-IDF keyword extraction.

[0111] In this embodiment, the input type can only be voice type or text type.

[0112] In this embodiment, the robot itself is equipped with multiple microphones to ensure the reliability of the collected voice. However, due to the different positions of the microphones, in order to ensure the reliability of the voice, it is necessary to analyze the noise and other conditions in the voice collection process to make the voice more realistic.

[0113] In this embodiment, the response curve refers to a waveform curve of the user's voice collected by the corresponding microphone.

[0114] In this embodiment, since the positions of the microphones are slightly different, each acquired response curve is sorted in order according to the distance between the sound source and the microphone to determine the response difference at each time point: ,in, represents the energy difference between the first microphone and the second microphone at time point t after sequential sorting; represents the energy difference between the n-1th microphone and the nth microphone at time point t based on the sequential sorting; They represent the audio energy of the first microphone, the second microphone, the n-1th microphone, and the nth microphone at the time point t after sequential sorting, respectively.

[0115] In this embodiment, the time delay is obtained by matching from the position-delay comparison table, and different position distances correspond to different delay times. Therefore, the time delay caused by the distance delay between the first microphone closest to the microphone and the second microphone closest to the microphone can be obtained, which facilitates the alignment of the two curves.

[0116] In this embodiment, the second curve that is closest to the first curve is aligned with the first curve after time shift according to the time delay to obtain a third curve that matches the second curve.

[0117] At this point, the delay factor for each time point is determined based on the response difference:

[0118] Each difference in the response difference at the corresponding time point is within the standard energy attenuation range at the corresponding distance between the two microphone positions. In this case, the attenuation factor is determined to be 0;

[0119] Otherwise, the attenuation factor is determined to be not 0, and the attenuation factor is calculated according to the following formula, which is:

[0120] ;

[0121] Where SJ represents the attenuation factor at the corresponding time point t; Nc represents the number of response differences at the corresponding time point t that are not within the standard energy attenuation range; M represents the total number of differences, and M=n-1; Indicates the absolute value of the difference of the i1th value that is not within the standard energy attenuation range; Indicates the maximum value in the i1th standard energy attenuation range;

[0122] At this time, the energy value at the corresponding time point t in the third curve is adjusted according to SJ, that is: the energy at the time point t in the third curve is A×ln(2+SJ), thereby achieving compensation adjustment of the third curve.

[0123] In this embodiment, the speech noise reduction process may be implemented by using a filter.

[0124] In this embodiment, the semantic parsing model is obtained by training a neural network model based on different speech and semantic analysis results of the speech as samples. At this time, the duration of the semantic parsing process can be obtained by inputting the speech into the model, that is, the duration of the model analyzing the speech.

[0125] In this embodiment, the tone complexity coefficient = (process duration / speech duration) × (storage capacity of tone information / storage capacity of text words).

[0126] In this embodiment, the coefficient-feature comparison table includes different tone complexity coefficients and the corresponding tone feature models, and the larger the coefficient, the higher the accuracy of the corresponding tone feature model. It is mainly used to determine the tone annotation of keywords, such as sigh, surprise, astonishment, etc.

[0127] In this embodiment, the tone feature model is obtained by training a neural network model based on samples of different keywords and the emotional tone results expressed by the keywords. Therefore, tone analysis and tone annotation can be directly implemented through the model, for example, keyword 1 - surprise, keyword 2 - surprise + anger.

[0128] In this embodiment, the first annotation includes all keywords and tone annotations for each keyword.

[0129] In this embodiment, when there are modal particles in the text, the corresponding keywords are directly annotated; if not, no annotation is performed.

[0130] In this embodiment, the tone of voice can be used to assist in analyzing the user's true intention, thereby obtaining the true interaction intention, such as interacting with the robot through facial emotions, such as having the robot imitate the user's facial expressions.

[0131] In this embodiment, the intent comparison table contains different interaction intentions and collection periods that match the intentions. For example, if it is an emotion imitation intention, the collection period is 0.1s. If it is a voice interaction intention, the collection period is 1s. It should be noted that the collection periods of different types of sensors can be the same or different.

[0132] The beneficial effect of the above technical solution is: the curves of the voice collected by the microphone are sorted in order to determine the corresponding differences at each time point, and then the curve closest to the second one is compensated and adjusted to ensure the reliability of the superposition processing with the second curve waveform, and then the text is obtained for keyword extraction and tone annotation, to ensure the authenticity of the interaction intention, provide an accurate basis for subsequent data collection, and ensure interaction efficiency.

[0133] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which aligns the extracted features of each type of sensor and combines the acquisition array of the acquisition points and the tone attributes of the target user to perform multi-dimensional pixel fusion, including:

[0134] Constructing a gradient vector for each pixel in the acquisition array at each acquisition point, wherein the gradient vector includes a brightness gradient component of the corresponding pixel under different types of sensors;

[0135] Determine the information saturation coefficient of each pixel according to the gradient vector;

[0136] When the information saturation coefficient is greater than a preset coefficient, the corresponding collection point is retained; otherwise, the corresponding collection point is eliminated;

[0137] Perform spatial domain conversion on the extracted features of each type of sensor to obtain a spatial array for each spatial point. When there are remaining points that are consistent with the corresponding spatial points, perform pixel fusion on the spatial array and the acquisition array at the consistent points according to the tone attributes of the target user to obtain and retain the fused array.

[0138] When there is no corresponding spatial point consistent with the remaining points, the acquisition array corresponding to the remaining points is retained;

[0139] Analyze the retained array of each point to obtain behavioral operation instructions.

[0140] In this embodiment, the brightness gradient component is the image gradient, and the image gradient refers to the rate of change of a certain pixel in the image in the x and y directions (compared with adjacent pixels), that is, the gradient vector is a two-dimensional vector.

[0141] In this embodiment, the information saturation coefficient ,in, 、 Represent the rate of change based on the x and y directions respectively; r1, g1, l1 represent the actual pixel component values based on the red channel, green channel and blue channel respectively; r0, g0, l0 represent the maximum pixel component values based on the red channel, green channel and blue channel respectively; min represents the minimum value sign; ln represents the logarithmic function sign.

[0142] In this embodiment, the preset coefficient is set to 0.4.

[0143] In this embodiment, spatial domain conversion refers to performing feature-position mapping conversion on the extracted features, that is, comparing the features with the corresponding relevant position points one by one to match the features with the position points. For example, feature A1 corresponds to position 1 and position 2, that is, feature A1 is determined based on position 1 and position 2. It should be noted that the spatial points are related position points, and the number of spatial points is less than the number of position points.

[0144] In this embodiment, for example, the spatial points are u1, u2, u3, u4, and u5, and the remaining points are u1, u2, u3, u4, u5, c1, c2, c3, c4, and c5. At this time, the consistent points are: u1, u2, u3, u4, and u5.

[0145] In this embodiment, the spatial array is the feature information of the corresponding spatial points under different types of sensors, and then the pixel information corresponding to the feature information is obtained from the feature-pixel comparison table, which contains different features and pre-set pixel conditions that match the features.

[0146] In this embodiment, the tone attribute is used to further optimize the array results and facilitate obtaining a fused array.

[0147] The beneficial effect of the above technical solution is: the information saturation coefficient of each pixel point is calculated based on the gradient vector to preliminarily determine whether the collection point is retained, and then the spatial array is obtained by performing spatial domain conversion on the features and combining the collection array and tone attributes to obtain a fusion array to ensure the accuracy of the acquisition of behavioral operation instructions.

[0148] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which performs pixel fusion on a spatial array and a collection array at a consistent point according to the tone attribute of the target user to obtain a fused array, including:

[0149] constructing an initial matrix based on the spatial array and the acquisition array;

[0150] When the tone attribute indicates that the input type of the target user's interaction with the robot is unrelated to speech, an average value of each column element in the initial matrix is calculated to obtain a first array;

[0151] When the tone attribute is an input type related to speech for the target user's interaction with the robot, determining the facial region of the target user on which the consistent point acts based on a comparison relationship between the consistent point and the target user's face, and setting simulation parameters for the consistent point according to an expression control database for the facial region;

[0152] Extracting the current tone of the target user at each time point during the interaction with the robot, and constructing facial change information for each consistent point based on the target user's tone portrait and simulation parameters, thereby constructing a facial array based on the tone;

[0153] Appending the face array to the first row of the initial matrix and combining the weights assigned to each row vector in the initial matrix to obtain a second array;

[0154] Among them, the first array and the second array are corresponding fusion arrays.

[0155] Since the acquisition array = {pixel information collected by each type of sensor at the corresponding acquisition point}, the spatial array is the feature information of the corresponding spatial point under different types of sensors, and then the pixel information corresponding to the feature information is obtained from the feature-pixel comparison table.

[0156] Therefore, the initial matrix is = .

[0157] In this embodiment, the first array = [the average value of all pixel values in each column].

[0158] In this embodiment, the correspondence relationship between the consistent point and the target user's face refers to the specific area of the point on the user's face, for example, the consistent point a1 is in the apple muscle area of the left cheek of the user's face.

[0159] In this embodiment, the expression control database stores expression emotion data related to different facial area blocks and simulation parameters for different expression emotion data, and the setting of the simulation parameters is to determine the parameters of the robot's facial emotion simulation, and the simulation parameters corresponding to different facial expressions are different. For example, the simulation parameters corresponding to the facial expression of laughing are: a1, a2, a4, a6, and laughter is achieved through simulation.

[0160] In this embodiment, the user tone portrait is pre-set, and the user tone portrait refers to the facial expression expressed by the user's tone, and the user tone portrait can be obtained by performing a portrait analysis of the user's facial expressions under different tones, which belongs to the existing technology.

[0161] In this embodiment, the current tone, user tone portrait, and simulation parameters are input into the facial analysis model to obtain corresponding facial change information. The facial analysis model is based on different tones, user tone portraits, and simulation parameters, and the facial change results for these parameters are obtained by training the neural network model with samples. Therefore, facial change information can be directly obtained, such as the facial changes of the consistent point under the corresponding tone, such as the cheek position C1 changes from height h1 to height h2.

[0162] Face array = {pixel information of each consistent point based on facial change information under different sensor types}. It should be noted that the pixel information of facial change information under different sensor types is obtained by matching from the transformation-pixel comparison table. The comparison table contains the pixel conversion of the final change result of the facial change information under different sensor types, and thus the corresponding pixel information can be directly obtained.

[0163] In this embodiment, the weights assigned are pre-set. Generally, the weights corresponding to the facial array, acquisition array, and spatial array are 0.3, 0.4, and 0.3, respectively. Then, the pixel information of the corresponding column is: weight × the sum of the element values of the corresponding column, and the second array = {the sum of the element values of each column}.

[0164] The beneficial effects of the above technical solution are: constructing an initial matrix based on the spatial array and the acquisition array, and then analyzing the arrays in different situations based on whether the voice is relevant. When it is relevant to the voice, based on the tone and portrait, and combined with the expression control database to set simulation parameters, a facial array is constructed to optimize the initial matrix and achieve accurate acquisition of the fusion array.

[0165] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which analyzes the retained array of each point to obtain behavioral operation instructions, including:

[0166] Divide the retained array of each point based on the user's face block criteria to obtain the current representation of each block;

[0167] Based on all current representations, the block weight of each divided block, and the user intention determined by the user voice of the target user, a behavioral operation instruction is obtained from the intention-combination-instruction comparison table.

[0168] In this embodiment, the user's face is divided into blocks based on the following criteria: nose, eyebrows, eyes, forehead, chin, cheeks, etc.

[0169] In this embodiment, all the retained arrays of each divided block are assigned to the corresponding blocks, and all the retained arrays under each divided block are sequentially input into the corresponding facial area model to obtain the current representation. The facial area model is obtained by training the neural network model based on the pixel arrays under different expressions of the area and the behavioral representation of the user's real emotions as samples. Therefore, the current representation can be directly obtained.

[0170] In this embodiment, the weight of each divided block is preset, and the sum of the weights is 1. It should be noted that the weights corresponding to different facial regions may be different.

[0171] In this embodiment, the intention-combination-instruction comparison table includes the intention, the current representation of different combinations, the block weights and the corresponding behavioral operation instructions, so that the behavioral operation instructions can be directly obtained. It should be noted that the operation instructions can be directly determined based on the intention, but through further analysis of the face and tone, the accuracy of the instruction determination can be further guaranteed.

[0172] The beneficial effect of the above technical solution is: by determining the current representation of each facial area and combining the user intention and block weight, the behavioral operation instructions are directly obtained, ensuring the accuracy of instruction acquisition and thus improving the interaction efficiency.

[0173] The present invention provides a multi-dimensional pixel fusion system based on multi-source perception, such as Figure 2 Shown, including:

[0174] A data acquisition module is used to control the various sensors provided on the robot to collect data, wherein the sensors include: a visible light camera, an infrared camera, and a lidar;

[0175] The alignment processing module is used to unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point;

[0176] A feature extraction module is used to arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and to extract features from the collected image according to a feature extraction method that matches the corresponding type of sensor;

[0177] The multi-dimensional pixel fusion module is used to align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion based on the acquisition array of the acquisition points and the tone attributes of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.

[0178] The beneficial effects of the above technical solution are: starting from multi-source perception to determine multi-dimensional pixel information, and combining the influence of tone feedback during the user's voice process to ensure the reliability of multi-dimensional pixel fusion, thereby making behavioral operation instructions more accurate, thereby improving interaction efficiency and satisfying the user's interactive experience.

[0179] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A multi-dimensional pixel fusion method based on multi-source perception, characterized in that: include: Step 1: Control the sensors installed on the robot to collect data, wherein the sensors include: a visible light camera, an infrared camera, and a lidar; Step 2: Unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point; Step 3: Arrange the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to the feature extraction method matching the corresponding type of sensor; Step 4: Align the extracted features of each sensor type, and perform multi-dimensional pixel fusion based on the acquisition array of the acquisition points and the tone attributes of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to execute the corresponding operation; The process of controlling the sensors set on the robot to collect data includes: Determine the input type of the target user's interactive input set currently received by the robot. When the input type is voice input, obtain the response curve of each microphone collecting the user's voice; Based on the current distance between the robot's microphone location and the target user's sound source, the response curves are sequentially sorted and aligned to obtain the response difference at each time point; Based on the response difference and time delay, a compensation adjustment is performed on the second curve that is the second closest, and waveform superposition processing is performed on the first curve that is the closest; Performing speech noise reduction processing on the superimposed speech to obtain the converted text; Performing keyword extraction on the converted text; Determining the duration of the process of the robot receiving the user's voice and performing semantic analysis on the user's voice, and at the same time, determining the voice duration of the user's voice; Obtaining a tone complexity coefficient according to the speech duration, the process duration, the tone information of the speech after the noise reduction process, and the number of words in the converted text; Retrieving a tone feature model that matches the tone complexity coefficient from a coefficient-feature comparison table to analyze the speech after noise reduction processing to determine a tone annotation for each keyword; Merging the tone annotation with the extracted keywords to obtain a first annotation; When the input type is text input, extract keywords from the input text to obtain a second annotation; Determine the target user's interaction intention based on the annotation, and match the acquisition cycles of different types of sensors from the intention comparison table; The method further includes: performing a time shift and alignment process on a second curve that is the second closest to the first curve according to the time delay, to obtain a third curve that matches the second curve; Each difference in the response difference at the corresponding time point is within the standard energy attenuation range at the corresponding distance between the two microphone positions. In this case, the attenuation factor is determined to be 0; Otherwise, the attenuation factor is determined to be not 0, and the attenuation factor is calculated according to the following formula, which is: ; Where SJ represents the attenuation factor at the corresponding time point t; Nc represents the number of response differences at the corresponding time point t that are not within the standard energy attenuation range; M represents the total number of differences, and M=n-1, where n represents the total number of microphones equipped on the robot; represents the absolute value of the difference of the i1th value that is not within the standard energy attenuation range; B represents the maximum value in the standard energy attenuation range; The energy value at the corresponding time point t in the third curve is adjusted according to SJ, that is: the energy at the time point t in the third curve A×ln(2+SJ), to achieve compensation adjustment of the third curve; Among them, the tone annotation refers to the emotional tone results.

2. The multi-dimensional pixel fusion method based on multi-source perception according to claim 1, characterized in that: Before controlling the sensors on the robot to collect data, the following steps must be performed: Capturing a trigger instruction to the robot from external information, and determining the instruction type of the trigger instruction; A start-up unit consistent with the instruction type is matched from a type-startup database, and the robot is started in a power-on state based on the start-up unit.

3. The multi-dimensional pixel fusion method based on multi-source perception according to claim 2, characterized in that: Control the robot's own sensors to collect data, including: When the robot is in a powered-on state, the robot collects an interactive input set of a target user currently received by the robot, and interactively analyzes the interactive input set to set a collection period for different types of sensors; The sensors of the corresponding type are controlled to collect data according to the collection period to obtain an initial set of each sensor.

4. The multi-dimensional pixel fusion method based on multi-source perception according to claim 1, characterized in that: Align the extracted features of each sensor type, and perform multi-dimensional pixel fusion based on the acquisition array of the acquisition points and the tone attributes of the target user, including: Constructing a gradient vector for each pixel in the acquisition array at each acquisition point, wherein the gradient vector includes a brightness gradient component of the corresponding pixel under different types of sensors; Determine the information saturation coefficient of each pixel according to the gradient vector; When the information saturation coefficient is greater than a preset coefficient, the corresponding collection point is retained; otherwise, the corresponding collection point is eliminated; Perform spatial domain conversion on the extracted features of each type of sensor to obtain a spatial array for each spatial point. When there are remaining points that are consistent with the corresponding spatial points, perform pixel fusion on the spatial array and the acquisition array at the consistent points according to the tone attributes of the target user to obtain and retain the fused array. When there is no corresponding spatial point consistent with the remaining points, the acquisition array corresponding to the remaining points is retained; Analyze the retained array of each point to obtain behavioral operation instructions.

5. The multi-dimensional pixel fusion method based on multi-source perception according to claim 4, characterized in that: Perform pixel fusion on the spatial array and the collection array at the consistent point according to the tone attribute of the target user to obtain a fused array, including: constructing an initial matrix based on the spatial array and the acquisition array; When the tone attribute indicates that the input type of the target user's interaction with the robot is unrelated to speech, an average value of each column element in the initial matrix is calculated to obtain a first array; When the tone attribute is an input type related to speech for the target user's interaction with the robot, determining the facial region of the target user on which the consistent point acts based on a comparison relationship between the consistent point and the target user's face, and setting simulation parameters for the consistent point according to an expression control database for the facial region; Extracting the current tone of the target user at each time point during the interaction with the robot, and constructing facial change information for each consistent point based on the target user's tone portrait and simulation parameters, thereby constructing a facial array based on the tone; Appending the face array to the first row of the initial matrix and combining the weights assigned to each row vector in the initial matrix to obtain a second array; Among them, the first array and the second array are corresponding fusion arrays.

6. The multi-dimensional pixel fusion method based on multi-source perception according to claim 5, characterized in that: Analyze the reserved array of each point to obtain behavioral operation instructions, including: Divide the retained array of each point based on the user's face block criteria to obtain the current representation of each block; Based on all current representations, the block weight of each divided block, and the user intention determined by the user voice of the target user, a behavioral operation instruction is obtained from the intention-combination-instruction comparison table.

7. A multi-dimensional pixel fusion system based on multi-source perception, characterized in that: include: A data acquisition module is used to control the various sensors provided on the robot to collect data, wherein the sensors include: a visible light camera, an infrared camera, and a lidar; The alignment processing module is used to unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point; A feature extraction module is used to arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and to extract features from the collected image according to a feature extraction method that matches the corresponding type of sensor; A multi-dimensional pixel fusion module is used to align the extracted features of each type of sensor, and perform multi-dimensional pixel fusion based on the acquisition array of the acquisition points and the tone attributes of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to execute the corresponding operation; The process of controlling the sensors set on the robot to collect data includes: Determine the input type of the target user's interactive input set currently received by the robot. When the input type is voice input, obtain the response curve of each microphone collecting the user's voice; Based on the current distance between the robot's microphone location and the target user's sound source, the response curves are sequentially sorted and aligned to obtain the response difference at each time point; Based on the response difference and time delay, a compensation adjustment is performed on the second curve that is the second closest, and waveform superposition processing is performed on the first curve that is the closest; Performing speech noise reduction processing on the superimposed speech to obtain the converted text; Performing keyword extraction on the converted text; Determining the duration of the process of the robot receiving the user's voice and performing semantic analysis on the user's voice, and at the same time, determining the voice duration of the user's voice; Obtaining a tone complexity coefficient according to the speech duration, the process duration, the tone information of the speech after the noise reduction processing, and the number of words in the converted text; Retrieving a tone feature model that matches the tone complexity coefficient from a coefficient-feature comparison table to analyze the speech after noise reduction processing to determine a tone annotation for each keyword; Merging the tone annotation with the extracted keywords to obtain a first annotation; When the input type is text input, extract keywords from the input text to obtain a second annotation; Determine the target user's interaction intention based on the annotation, and match the acquisition cycles of different types of sensors from the intention comparison table; The method further includes: performing a time shift and alignment process on a second curve that is the second closest to the first curve according to the time delay, to obtain a third curve that matches the second curve; Each difference in the response difference at the corresponding time point is within the standard energy attenuation range at the corresponding distance between the two microphone positions. In this case, the attenuation factor is determined to be 0; Otherwise, the attenuation factor is determined to be not 0, and the attenuation factor is calculated according to the following formula, which is: ; Where SJ represents the attenuation factor at the corresponding time point t; Nc represents the number of response differences at the corresponding time point t that are not within the standard energy attenuation range; M represents the total number of differences, and M=n-1, where n represents the total number of microphones equipped on the robot; represents the absolute value of the difference of the i1th value that is not within the standard energy attenuation range; B represents the maximum value in the standard energy attenuation range; The energy value at the corresponding time point t in the third curve is adjusted according to SJ, that is: the energy at the time point t in the third curve A×ln(2+SJ), to achieve compensation adjustment of the third curve; Among them, the tone annotation refers to the emotional tone results.

Citation Information

Patent Citations

  • Control system and method of artificial intelligence drive robot

    CN108297098A

  • Robot-oriented multimodal fusion emotion calculation method and system

    CN108960191A